You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
Serving Qwen3.8-27B-FP8 on a single DGX Spark (GB10): 7.88 to 58.5 tok/s single-stream from decode strategy alone, weights untouched. Speculative decoding and prefix caching benchmarked, plus DFlash 2 — the only Qwen3.8-27B build that can serve it under vLLM.
Qwen 3.8 DFlash2: 27B AI Running at 236 TOK/S Locally! - High-speed local prose and code autocomplete powered by speculative decoding with lightweight draft models.
Native macOS control center and local AI agent gateway for Qwen3.8 on Apple Silicon — MLX, DFlash2, OpenAI Responses, Anthropic Messages, Claude Code, Codex, OpenCode and Grok Build.
Qwen3.8-27B on SGLang for Blackwell SM120: DFlash2 block-16 + triton. 268 tok/s on one RTX PRO 5000 (48GB), full 262K context. Measured, with honest sample sizes.