
An open-source LLM inference and serving engine that maximizes throughput and minimizes cost through PagedAttention, continuous batching, quantization, disaggregated prefill/decode, and distributed inference across diverse hardware, with a drop-in OpenAI-compatible API.