
An agentic CLI that automates production LLM inference optimization by profiling workloads, selecting the best inference engine (vLLM, SGLang, TensorRT-LLM, etc.), tuning configurations via Bayesian search, synthesizing custom CUDA/ROCm/Triton kernels, and validating deployments with functional, load, and security tests — maximizing GPU throughput and minimizing latency on user-owned hardware.