Rapid-MLX v0.16.0 Released: TensorFold Acceleration and System One Decision Models on Apple Silicon
Rapid-MLX v0.16.0 has officially launched for Apple Silicon Mac users, introducing TensorFold execution profiles that boost DeepSeek V4 Flash throughput up to 2
On October 10, 2026, Rapid-MLX officially released version 0.16.0, bringing major runtime updates for running local large language models (LLMs) on Apple Silicon Macs. The update introduces hardware-accelerated TensorFold profiles powered by MLX, native integration for System One decision models, and enhanced support for OpenCode 2.x coding agents.
![]()
Image source: @rapidmlx via X
Image Credit: @rapidmlx via X
Rapid-MLX v0.16.0 focuses on substantial kernel throughput optimizations while making local triage, ticket routing, and decision scoring immediately accessible directly on macOS without relying on cloud services.
TensorFold Profiles and Model Throughput Gains
The centerpiece of Rapid-MLX v0.16.0 is the official introduction of TensorFold profiles, integrating optimizations derived from Pierre's mlx2 research to maximize compute efficiency on Apple Silicon unified memory architectures.
- DeepSeek V4 Flash: Delivers up to 2.1× inference throughput compared to earlier baselines.
- Bonsai 2 27B: Gains a +42% throughput increase.
- Gemma 4 26B: Shows a +4% speedup.
While acceleration gains vary across model families, the TensorFold profile allows developers to extract significantly higher token generation rates from existing Mac hardware without sending prompts over the network.
Native System One Decision Models (Decider, Clef, and OpenJev)
Moving beyond standard generative chat, Rapid-MLX v0.16.0 adds native MLX support for fast System One decision workflows such as ticket routing, urgency scoring, and answer ranking.
Developers can launch decision models directly via the CLI:
rapid-mlx system-one decider-2b
The supported decision model roster includes:
- Decider 2B: A lightweight, high-speed local decision model.
- Clef Flash 9B and Clef 27B: Native multimodal input support alongside text-based judgment tasks.
- OpenJev 27B: A larger reasoning and ranking model (licensed under non-commercial terms).
OpenCode 2.x Support and Developer Workflow Improvements
Rapid-MLX v0.16.0 also strengthens its integrations for automated coding workflows and agent environments:
rapid-mlx agents opencode --setup --yes
- OpenCode 2.x Integration: Improved context reuse across conversation turns reduces memory overhead and latency in multi-turn coding sessions.
- Reliable Tool-Calling JSON: Enhanced JSON output reliability minimizes formatting errors during agent tool invocation.
- Concurrent Streaming and Resumable Downloads: Smoother multi-stream response handling and fault-tolerant resumable model downloads improve daily operations.
Sources
- Rapid-MLX Official X (@rapidmlx): Rapid-MLX v0.16.0 Release Announcement
- GitHub Release: raullenchai/Rapid-MLX v0.16.0