The Large Model Systems Organization develops large models and systems that are open, accessible, and scalable.
Latest Blog
See all postsAccelerating Long-Context and Agentic Inference with NVFP4 KV Cache
The KV cache is a fundamental building block of the modern LLM inference system. The context from multiple conversation rounds in agent sessions is cached as keys and values (KV) in GPU memory, allowi...

SGLang and Miles Add Day-0 Support for DeepSeek-V4.1
DeepSeek-V4.1 introduces several architecture choices that shape the serving stack. Low-ratio compression and sliding-window attention (SWA). Each layer maintains an fp8 sliding-window cache for its ...

Running DeepSeek-V4-Flash and Kimi-K3 on Consumer Hardware with SSD Expert Pack
SGLang brings the core idea of SSD-LLaMA to MoE inference: keep routed experts that do not fit in VRAM and host RAM on an NVMe SSD, load only the experts selected by the router, and use Expert Pack la...
Projects
View all projectsOur Sponsors & Partners
Backed by leading companies and institutions advancing AI research.
Voltage Park, NVIDIA, Nebius, Google Cloud, AtlasCloud, a16z, AMD, InnoMatrix, Laude Institute, Hyperbolic, NovitaAI, Verda Cloud, Sky9, Kaggle, MBZUAI, Together, RunPod, Anyscale, HuggingFace




