This release fixes issues for SREs watching stability and regressions.
✓ No known CVEs patched in this version
Topics
+12 more
Summary
AI summaryFixed NaN outputs from flashinfer trtllm FP4 MoE kernels on long inputs.
Changes in this release
| Type | Severity | Summary | CVE |
|---|---|---|---|
| Bugfix | Medium |
Fixes NaN outputs from flashinfer trtllm FP4 MoE kernels on long inputs. Fixes NaN outputs from flashinfer trtllm FP4 MoE kernels on long inputs. Source: llm_adapter@2026-07-14 Confidence: high |
— |
| Bugfix | Medium |
Fixes GLM 5.2 IndexShare issues on PD disaggregation setting. Fixes GLM 5.2 IndexShare issues on PD disaggregation setting. Source: llm_adapter@2026-07-14 Confidence: high |
— |
| Bugfix | Medium |
Fixes GLM 5.2 IndexShare issues on Context Parallel setting. Fixes GLM 5.2 IndexShare issues on Context Parallel setting. Source: llm_adapter@2026-07-14 Confidence: high |
— |
| Bugfix | Medium |
Fixes DSA model launching on non Cuda/HIP devices. Fixes DSA model launching on non Cuda/HIP devices. Source: llm_adapter@2026-07-14 Confidence: high |
— |
| Bugfix | Medium |
Fixes flashinfer dependency on Cuda 12 images. Fixes flashinfer dependency on Cuda 12 images. Source: llm_adapter@2026-07-14 Confidence: high |
— |
Full changelog
v0.5.15.post1 includes a few patches, mostly for GLM 5.2
- #30454 #30627: Fix DSA model launching on non Cuda/HIP devices
- #30858: Fix flashinfer dependency on Cuda 12 images
- #31001: Fix NaN outputs caused by flashinfer trtllm FP4 MoE kernels on long input
- #30839: Fix GLM 5.2 IndexShare on PD disaggregation setting
- #30992: Fix GLM 5.2 IndexShare on Context Parallel setting
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
Share this release
About sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Beta — feedback welcome: [email protected]