Luminal is a San Francisco-based AI inference compiler company (YC S25) that compiles and optimizes AI models for GPUs and ASICs, delivering high-throughput, low-latency inference in production. Their compiler-first approach enables single-line deployment to production and supports heterogeneous hardware backends including Nvidia GPUs and custom accelerators like Positron Atlas, using techniques such as megakernel fusion, disaggregated serving, and large-scale kernel search.