Rapid-MLX v0.15.4: Qwen3.8-27B TensorFold and DFlash2 Speculative Decoding on Apple Silicon
Apple Silicon local LLM engine Rapid-MLX v0.15.4 adds an experimental TensorFold profile with DFlash2 speculative decoding for Qwen3.8-27B alongside MTP acceler
Rapid-MLX, an open-source local LLM inference engine optimized for Apple Silicon, has officially rolled out version 0.15.4, introducing an experimental opt-in profile named qwen3.8-27b-tensorfold for the Qwen3.8-27B foundation model. By pairing the target checkpoint with a DFlash2 speculative drafter, this update mitigates the sequential decoding bottleneck of autoregressive generation on Apple Silicon's unified memory architecture without modifying underlying model weights or altering output distributions.
![]()
Image source: Rapid-MLX
DFlash2 Speculative Decoding and the TensorFold Acceleration Pipeline
Speculative decoding is an inference optimization technique where a compact draft model proposes multiple future token candidates in advance, allowing the primary target model to verify them simultaneously in a single forward pass.
In Rapid-MLX v0.15.4, the qwen3.8-27b-tensorfold profile pairs the Qwen3.8-27B checkpoint with a dedicated DFlash2 drafter:
- Parallel Token Verification: As proposed tokens from the drafter are accepted, the 27B target model generates multiple confirmed tokens per expensive forward pass, transforming sequential token generation into a significantly more parallelized computation.
- Exact Distribution Preservation: Because speculative decoding strictly verifies candidate tokens against the target model's probability distribution, generation output remains bitwise-exact and identical in quality without accuracy degradation or quantization loss.
- Hardware-Dependent Performance: Community testing on Apple Silicon M4 Max hardware reported generation speeds reaching 154 tok/s for numerical tasks and 114 tok/s for code generation. However, official release notes avoid quoting a fixed speedup multiplier, noting that acceleration depends heavily on hardware tier, memory bandwidth, and prompt acceptance rates.
TensorFold Constraints and the Feature-Complete 4-Bit Baseline
The newly added TensorFold execution path is tailored specifically for high-throughput single-stream text generation, establishing clear boundaries for production deployments:
- Single Stream Only: Designed to process exactly one text generation request at a time, making it unsuitable for concurrent multi-stream serving.
- Agent Feature Exclusions: Does not support tool calling, multimodal media inputs, or grammar constraints, explicitly rejecting requests that include these parameters.
- Default
qwen3.8-27b-4bitPath: Standard agent workflows and structured tool-use environments continue to rely on the feature-complete default path. Launched via$ rapid-mlx serve qwen3.8-27b-4bit, it natively supports the Hermes tool-calling envelope and<thought>chain-of-thought streaming across/v1/chat/completionsand/v1/responsesendpoints. - Baseline Performance Metrics: On a Mac mini M4 Pro (48GB), the standard
qwen3.8-27b-4bitprofile (16.3GB weight footprint) clocks 24.1 tok/s decode on a single stream, 24.0 tok/s across 4 concurrent streams, with a peak Metal memory footprint of 25.3GB.
GLM-5.3 Flash MTP Support and Local Apple Silicon Inference
Beyond Qwen3.8-27B, Rapid-MLX v0.15.4 broadens acceleration support for high-memory Mac workstations:
- GLM-5.3 Flash TensorFold Profile: Provides a dedicated TensorFold acceleration profile tailored for 256GB Apple Silicon Mac configurations.
- Out-of-the-Box MTP Head Acceleration: Natively accelerates models equipped with Multi-Token Prediction (MTP) heads, enhancing execution efficiency for next-generation open-weights models.
Rapid-MLX v0.15.4 illustrates how inference-level speculative scheduling can unlock substantial generation throughput on Apple Silicon unified memory, enabling practical local execution of high-capability 27B models without requiring destructive weight compression.
Sources
- Rapid-MLX Model Catalog (Qwen3.8 27B): https://rapidmlx.com/models/qwen3.8-27b
- Rapid-MLX Official Website: https://rapidmlx.com/
- Rapid-MLX Qwen Family Documentation: https://rapidmlx.com/docs/models/families/qwen
- FHILY X (@Oluwaphilemon1): Rapid-MLX v0.15.4 Qwen3.8-27B acceleration release
- GitHub Issue (raullenchai/Rapid-MLX #1941): DSpark speculative decoding for Qwen3.8-27B