Skip to main content

One post tagged with "PD Disaggregation"

View All Tags

Are You Making Good Use of Your Compute? Three Stages of vLLM Inference Cluster Optimization

· 11 min read
HAMi Maintainer, Co-founder & CTO of Dynamia

Three Stages of vLLM Inference Cluster Optimization | Li Mengxuan

On July 16, 2026, Li Mengxuan, Co-founder & CTO of Dynamia and HAMi author, delivered a technical talk on vLLM deployment and compute optimization at vLLM Meetup. Built around one pointed question, "Are you making good use of your compute?", the talk laid out a complete evolution path for vLLM inference clusters, from "getting it to run" to "squeezing the hardware dry", broken down into three clear stages.

This recap walks through the talk slide by slide, combining the deck with the on-site Q&A notes.

CNCFHAMi is a CNCF Incubating project