vLLM is an open-source, high-throughput and memory-efficient inference and serving engine for large language models (LLMs), originally developed at UC Berkeley's Sky Computing Lab and now maintained by a large community of contributors from academia and industry. It provides a Python library and OpenAI-compatible API server that enables fast, cost-efficient LLM deployment across diverse hardware including NVIDIA and AMD GPUs, CPUs, and various accelerator ecosystems.