What Tablif read on ds4's site
ds4 describes itself as: “ds4 Local inference engine for DeepSeek V4, GLM, Qwen”
- Stage: Live, no waitlist
- How it differs: “Not a generic GGUF runner. ds4 follows a small, opportunistic set of model families and validates each supported layout end to end.”
What it is
Calls itself “inference engine”.
- Calls itself: inference engine
“…ds4 Local inference engine for DeepSeek V4, GLM, Qwen”
“DwarfStar 4 (ds4) is a narrow C inference engine from Salvatore Sanfilippo (antirez)”
Where it puts itself
Rules out “Generic GGUF files”.
- Says it will not: Generic GGUF files
“Not a generic GGUF runner. ds4 follows a small, opportunistic set of model families and validates each supported layout end to end.”
“Generic GGUF files are not the target.”
What it costs
No free plan.
- Free plan: no
“Engine MIT · antirez/ds4”
“DwarfStar 4 (ds4) is a narrow C inference engine from Salvatore Sanfilippo (antirez) for running DeepSeek V4/V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next locally on high-memory Mac, CUDA and ROCm machines.”
Who it sells to
Sells to developers. Made for coding agent.
- Sells to: developers
- Made for: coding agent
“…ds4-server speaks OpenAI and Anthropic-style APIs, so local coding agents can connect to your own machine with a base URL.”
“…and a native coding agent sharing the same model state.”
How it runs
A command-line tool for Linux. Runs locally and works offline. Offers an API. AI is the core of the product.
- Kind: cli
- Runs on: Linux
- Runs where: on the device, offline
- API: yes
- AI role: core to the product
“Ships a CLI, OpenAI/Anthropic-compatible local APIs, and a native coding agent sharing the same model state.”
“NVIDIA DGX Spark or a generic CUDA Linux box”
What it works with
Integrates with Claude Code, Codex and ds4-server. Names DeepSeek V4, GLM 5.x and 1 more on its pages.
- Integrates with: Claude Code · Codex · ds4-server
- Names: DeepSeek V4 · GLM 5.x · Qwen3.8 Flash Next
“Use ds4 from Codex, Claude Code and OpenCode”
“DeepSeek V4/V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next locally on high-memory Mac, CUDA and ROCm machines.”
What it offers as proof
Live, no waitlist. Instant signup. Publishes docs and a blog.
- Instant signup: yes
“Get started →”
“Docs -> https://dwarfstar.sh/docs/quickstart/”
Sources
Tablif read these pages, last on 7 Oct 2026. Everything above is what the product says about itself; Tablif does not verify every claim.
Similar products in Tablif Town
- DSH Desktop: Official app for DeepSeek’s open-source agent harness
- Prime Inference: Fast, reliable serving for frontier open models
- Neurogrid Community Cloud: Community Marketplace for LLM inference
- Netra Runtime: Fastest AI Inference at Any Scale
- Tensorant: Fine-tuning and LLM Inference Provider
- Lifeboat: Run 2–6x more agents on the GPU you already have
- ModelRunner: One API for every AI model: image, video, audio, 3D, llm
- IronMule: High-performance local LLM runtime for Apple Silicon