The Large Model Systems Organization develops large models and systems that are open, accessible, and scalable.
Latest Blog
See all posts
SGLang and Miles Add Day-0 Support for Qwen3.8
We are excited to announce Day-0 support for Qwen3.8-2.4T-A95B in SGLang and Miles. It is Qwen's largest open-source model, with 2.4T total parameters and 95B active per token, and its hybrid attentio...

SGLang Adds Day-0 Support for NVIDIA Nemotron 3.5 Lightning
SGLang is excited to announce Day-0 support for NVIDIA Nemotron 3.5 Lightning, a customizable open model built to power always-on agents across local systems, the edge, the datacenter, and the cloud. ...
Unified Radix Cache: One Tree for Hybrid Model Prefix Caching
Prefix caching reuses KV when requests share the same token prefix. Under full attention, once the KV for a shared prefix is computed, it remains valid as more tokens are appended. A later request wit...
Projects
View all projectsOur Sponsors & Partners
Backed by leading companies and institutions advancing AI research.
Voltage Park, NVIDIA, Nebius, Google Cloud, AtlasCloud, a16z, AMD, InnoMatrix, Laude Institute, Hyperbolic, NovitaAI, Verda Cloud, Sky9, Kaggle, MBZUAI, Together, RunPod, Anyscale, HuggingFace




