Rapid-MLX v0.15.4: Qwen3.8-27B TensorFold and DFlash2 Speculative Decoding on Apple Silicon

Apple Silicon local LLM engine Rapid-MLX v0.15.4 adds an experimental TensorFold profile with DFlash2 speculative decoding for Qwen3.8-27B alongside MTP acceler

tau · October 3, 2026

#Rapid-MLX #AppleSilicon #Qwen3.8 #SpeculativeDecoding #LLM #MLX

Rapid-MLX v0.15.4: Qwen3.8-27B TensorFold and DFlash2 Speculative Decoding on Apple Silicon

Rapid-MLX, an open-source local LLM inference engine optimized for Apple Silicon, has officially rolled out version 0.15.4, introducing an experimental opt-in profile named qwen3.8-27b-tensorfold for the Qwen3.8-27B foundation model. By pairing the target checkpoint with a DFlash2 speculative drafter, this update mitigates the sequential decoding bottleneck of autoregressive generation on Apple Silicon's unified memory architecture without modifying underlying model weights or altering output distributions.

Architectural diagram of Rapid-MLX v0.15.4 Apple Silicon inference engine with Qwen3.8-27B TensorFold profile and DFlash2 speculative drafter pipeline

Image source: Rapid-MLX

DFlash2 Speculative Decoding and the TensorFold Acceleration Pipeline

Speculative decoding is an inference optimization technique where a compact draft model proposes multiple future token candidates in advance, allowing the primary target model to verify them simultaneously in a single forward pass.

In Rapid-MLX v0.15.4, the qwen3.8-27b-tensorfold profile pairs the Qwen3.8-27B checkpoint with a dedicated DFlash2 drafter:

  • Parallel Token Verification: As proposed tokens from the drafter are accepted, the 27B target model generates multiple confirmed tokens per expensive forward pass, transforming sequential token generation into a significantly more parallelized computation.
  • Exact Distribution Preservation: Because speculative decoding strictly verifies candidate tokens against the target model's probability distribution, generation output remains bitwise-exact and identical in quality without accuracy degradation or quantization loss.
  • Hardware-Dependent Performance: Community testing on Apple Silicon M4 Max hardware reported generation speeds reaching 154 tok/s for numerical tasks and 114 tok/s for code generation. However, official release notes avoid quoting a fixed speedup multiplier, noting that acceleration depends heavily on hardware tier, memory bandwidth, and prompt acceptance rates.

TensorFold Constraints and the Feature-Complete 4-Bit Baseline

The newly added TensorFold execution path is tailored specifically for high-throughput single-stream text generation, establishing clear boundaries for production deployments:

  • Single Stream Only: Designed to process exactly one text generation request at a time, making it unsuitable for concurrent multi-stream serving.
  • Agent Feature Exclusions: Does not support tool calling, multimodal media inputs, or grammar constraints, explicitly rejecting requests that include these parameters.
  • Default qwen3.8-27b-4bit Path: Standard agent workflows and structured tool-use environments continue to rely on the feature-complete default path. Launched via $ rapid-mlx serve qwen3.8-27b-4bit, it natively supports the Hermes tool-calling envelope and <thought> chain-of-thought streaming across /v1/chat/completions and /v1/responses endpoints.
  • Baseline Performance Metrics: On a Mac mini M4 Pro (48GB), the standard qwen3.8-27b-4bit profile (16.3GB weight footprint) clocks 24.1 tok/s decode on a single stream, 24.0 tok/s across 4 concurrent streams, with a peak Metal memory footprint of 25.3GB.

GLM-5.3 Flash MTP Support and Local Apple Silicon Inference

Beyond Qwen3.8-27B, Rapid-MLX v0.15.4 broadens acceleration support for high-memory Mac workstations:

  • GLM-5.3 Flash TensorFold Profile: Provides a dedicated TensorFold acceleration profile tailored for 256GB Apple Silicon Mac configurations.
  • Out-of-the-Box MTP Head Acceleration: Natively accelerates models equipped with Multi-Token Prediction (MTP) heads, enhancing execution efficiency for next-generation open-weights models.

Rapid-MLX v0.15.4 illustrates how inference-level speculative scheduling can unlock substantial generation throughput on Apple Silicon unified memory, enabling practical local execution of high-capability 27B models without requiring destructive weight compression.

Sources