Run frontier open weights locally with ds4.
DwarfStar 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.
SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE · C / METAL / CUDA / ROCM · QWEN ON 64GB
A 284-billion-parameter star
DeepSeek V4 Flash is a large mixture-of-experts model. The usual path is remote serving; ds4 starts from the opposite constraint.
Compressed, not lobotomized
Asymmetric quantization targets the routed experts while preserving critical paths. The model becomes practical on high-memory machines.
Dense, resident, yours
The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.
Asymmetric 2-bit quantization
Compress the routed experts, keep critical shared paths precise. That is how the supported routed-MoE builds fit their target machines.
KV cache as a disk citizen
Save long prefixes to SSD and resume by prompt hash. Restarts do not have to mean full re-prefill.
One engine, three interfaces
Use ./ds4 for chat, ./ds4-server for local APIs and ./ds4-agent for persistent coding sessions.
RUNTIME MAP · SIMPLIFIED. SEE ARCHITECTURE NOTES FOR THE FULL DRAWING.
V4 Flash Q2 is the baseline. At 128 GB, GLM 5.3 Q2 and Qwen Q4 also fit; V4.1 Q2 streams from SSD.
./download_model.sh ds4f-q2 && make
REF · M5 MAX 128GB · 32K CTX: 34.4 T/S GEN · 557 T/S PREFILL
Estimates from the ds4 benchmark table . Full guide in Hardware and Installation .
Own your local AI inference.
Start with the quickstart, check the hardware matrix, then connect your editor, agent or API client to the local server.