Doctoral Speaking Skills Talk - Zikun Li
October 8, 2026 3:00PM—4:00PM
Location:
8102
-
Gates and Hillman Centers
Speaker:
ZIKUN LI,
Ph.D. Student, Computer Science Department, Carnegie Mellon University
https://zikun-li.github.io/
LLM serving systems increasingly disaggregate inference into finer-grained stages, with recent approaches separating attention from FFN or MoE execution during decode. This operator-level disaggregated serving (ODS) can improve hardware matching and enable independent scaling, particularly across heterogeneous devices. However, existing systems fix operator boundaries and lack a unified characterization of when disaggregation reduces serving cost.
In this talk we present OpWeave, an end-to-end framework for heterogeneous ODS. OpWeave provides an analytical cost model that bounds the gains of homogeneous and heterogeneous ODS over colocated serving. It jointly optimizes operator partitioning and deployment configuration through a regularity-aware planner that keeps the search tractable even for hybrid-attention models. A vLLM-based runtime executes the synthesized plans with flexible operator stages across heterogeneous device groups. In our evaluation, OpWeave reduces serving cost by up to 1.78× on homogeneous and 1.89× on heterogeneous GPU clusters relative to the best feasible baseline, while meeting latency SLOs.
Presented in Partial Fulfillment of the CSD Speaking Skills Requirement
Contact
Matt Stewart