Apple MLR Introduces 'LoopCD': Training-Free Contrastive Decoding for Looped Transformers

Apple Machine Learning Research (Apple MLR) has unveiled LoopCD, a training-free contrastive decoding technique for looped transformers that boosts AIME 2024 ac

tau · October 4, 2026

#Apple-MLR #LoopCD #Looped-Transformers #Contrastive-Decoding #Reasoning-Efficiency

Apple MLR Introduces 'LoopCD': Training-Free Contrastive Decoding for Looped Transformers

On October 2, 2026, researchers at Apple Machine Learning Research (Apple MLR) released a paper titled 'Decoding Looped Transformers Better for (Almost) Free' (arXiv:2610.02185), introducing LoopCD, a training-free contrastive decoding framework that significantly boosts reasoning performance and computational efficiency for looped transformers.

Architectural diagram of Apple MLR's LoopCD training-free contrastive decoding for looped transformers

Image source: Apple MLR / Ruixiang Zhang (@onloglogn)

Led by Ruixiang Zhang, Weihao Liu, and colleagues at Apple MLR, the research demonstrates how internal intermediate loop states—traditionally discarded during inference—can serve as built-in weak references to guide token selection without requiring auxiliary models or additional fine-tuning.

Repurposing Early Recurrent States into Contrastive Signals

Looped Transformers achieve remarkable parameter efficiency by repeatedly cycling through a shared layer block across recurrent iterations, increasing effective computational depth without expanding physical parameter counts.

Under standard autoregressive decoding, however, only the final loop's output is used to predict the next token, while intermediate representations generated in earlier passes are thrown away. The Apple MLR team observed that because earlier iterations are trained toward the same next-token objective with less computation, they naturally provide aligned weak predictions.

Unlike conventional contrastive decoding methods that require running a separate, smaller auxiliary model in parallel, LoopCD derives its contrastive signal directly from within the same model execution:

  • LoopCD-Logits: Operates in logit space by computing the probability distribution divergence between the final loop and an early loop, re-ranking candidate tokens with just a single extra output projection pass.
  • LoopCD-Hidden: Applies contrastive adjustments directly in the hidden-state space prior to the language model head, introducing zero output projection overhead.

73.3% on AIME 2024 and Up to 48.2% FLOPs Reduction

The researchers evaluated LoopCD across four representative open-source looped transformer families: Ouro, Huginn, Parcae, and Looped-Qwen3.

  • Mathematical Reasoning (AIME 2024): Applying LoopCD to Ouro-2.6B-Thinking increased pass@1 accuracy from 61.88% (61.9%) to 73.33% (73.3%), representing an 11.45 percentage point gain.
  • Code Generation (HumanEval): On the Huginn model, pass@1 accuracy rose from 22.56% (22.6%) to 31.71% (31.7%).

Beyond raw accuracy gains, LoopCD delivers substantial compute savings. When recurrent iterations were halved, models augmented with LoopCD maintained the average performance across seven standard benchmarks compared to full-depth execution, reducing forward computational cost (FLOPs) by 22.5% to 48.2%. The team verified that using the first recurrent iteration as the weak reference provided the strongest contrastive benefit.

Scope, Architectural Boundaries, and Practical Considerations

While LoopCD offers significant improvements without extra parameter training or model hosting overhead, its application is bound to specific architectural constraints:

  • Looped Architectures Only: LoopCD is designed explicitly for looped transformers utilizing weight-tied recurrence; it cannot be applied directly to conventional stacked, non-recurrent transformer architectures.
  • Frontier Model Context: While broader community discussions frequently cite rumors regarding looped transformer architectures in frontier models (such as GPT-6 Astra or Gemini 4), the experimental findings in this paper were conducted rigorously on open-source looped model checkpoints.

For engineers and researchers deploying compact recurrent reasoning models on resource-constrained or on-device environments, LoopCD presents a practical, drop-in decoding enhancement that extracts higher accuracy and lower latency from existing model weights.

Sources