Skip to content
BigMoeOnEdge
LLM Frameworks
Run massive Mixture‑of‑Experts models on edge devices with far less RAM by streaming only needed experts from flash storage
C++
·
Latest v0.15.1 · 4d ago
Security brief →
Features
-
Expert streaming: reads only the required experts directly from flash per token
-
Expert cache: keeps frequently used experts in RAM (manual or auto sizing)
-
Direct flash reads with configurable I/O threads to bypass OS page cache
-
Dense‑weight policy options (mmap, warm, anon) for always‑used model parts
-
I/O–compute overlap hides storage latency behind computation
Review required
v0.15.1
New feature
·
Expert‑dropping warnings
No immediate action
v0.15.0
New feature
·
Cache‑aware expert dropping
No immediate action
v0.14.0
New feature
·
Pinned dense weights via dma-buf
No immediate action
v0.13.2
Breaking risk
·
Flag precedence override
No immediate action
v0.13.1
Bug fix
·
`none` reporting fix
Weekly OSS security release digest.
The CVE patches and breaking changes that affected production tools this week. One email, every Sunday.
No spam, unsubscribe anytime.
About
Languages
C++
·
Python
·
PowerShell
View on GitHub
Search tools, categories, lists, and users
Use ↑↓ to navigate, Enter to open, Esc to close
No results for ""
⌘K to open
↑↓ navigate
⏎ open