The Large Model Systems Organization develops large models and systems that are open, accessible, and scalable.
Latest Blog
See all posts
Infer-forge: Harness, Loop, and Graph Engineering Around SGLang
Inference optimization may look local in code, but its validity is global. A kernel, communication path, or scheduling change becomes meaningful only at a specific deployment point defined by the mode...

MiniMax-H3 on 8×H200: 1.95× Lossless, Up to 6.24× at 0.76–0.91 SSIM
We benchmarked MiniMax-H3 video generation on 8× NVIDIA H200 with SGLang Diffusion, holding prompts, seeds, resolution, frame rate, and denoising steps fixed across six workloads. - SGLang's dense, l...

Qwen3.8-Flash-Next: Day-0 Support in SGLang
Today, the Qwen team open-sourced Qwen3.8-Flash-Next, a multimodal MoE model and an early preview of the Qwen4 architecture. It plays the same role for Qwen4 that Qwen3-Next played for Qwen3.5. The Ga...
Projects
View all projectsOur Sponsors & Partners
Backed by leading companies and institutions advancing AI research.
Voltage Park, NVIDIA, Nebius, Google Cloud, AtlasCloud, a16z, AMD, InnoMatrix, Laude Institute, Hyperbolic, NovitaAI, Verda Cloud, Sky9, Kaggle, MBZUAI, Together, RunPod, Anyscale, HuggingFace




