
A production-grade AI model serving platform featuring a fully disaggregated inference engine with custom kernels, speculative decoding, and KV caching, offering serverless, on-demand, and reserved capacity deployment options for open-source and custom models across text, vision, and audio modalities.