Auto Inference provides an agentic CLI for production LLM inference optimization that runs on user-owned infrastructure. It automatically selects the optimal inference engine for a given model and hardware configuration, tunes all parameters, synthesizes custom kernels, and stress-tests deployments to maximize GPU throughput across NVIDIA, AMD, Intel, Google TPU, and Apple Silicon hardware.