A code-level tour of how the vLLM AFD Plugin separates Attention and FFN: plugin bootstrap, role-specific model construction, per-layer transfers, connector control and data planes, GPU fan-in/fan-out, ubatching, and the evidence-driven roadmap.
🔗 Key References:
Architecture, connector support, and early performance results
Proposed workstreams, evidence policy, and feedback topics