The Large Model Systems Organization develops large models and systems that are open, accessible, and scalable.
Latest Blog
See all postsAdvanced CUDA Graph Techniques in SGLang
CUDA Graphs promise to remove kernel-launch overhead, but getting close to that benefit in a real inference engine requires graphing as much of the workload as possible without sacrificing compatibili...

SGLang and Miles Add Day-0 Support for Qwen3.8
We are excited to announce Day-0 support for Qwen3.8-2.4T-A95B in SGLang and Miles. It is Qwen's largest open-source model, with 2.4T total parameters and 95B active per token, and its hybrid attentio...

SGLang Adds Day-0 Support for NVIDIA Nemotron 3.5 Lightning
SGLang is excited to announce Day-0 support for NVIDIA Nemotron 3.5 Lightning, a customizable open model built to power always-on agents across local systems, the edge, the datacenter, and the cloud. ...
Projects
View all projectsOur Sponsors & Partners
Backed by leading companies and institutions advancing AI research.
Voltage Park, NVIDIA, Nebius, Google Cloud, AtlasCloud, a16z, AMD, InnoMatrix, Laude Institute, Hyperbolic, NovitaAI, Verda Cloud, Sky9, Kaggle, MBZUAI, Together, RunPod, Anyscale, HuggingFace




