Decision 2.0 Released: Open-Source Decision AI Family Spanning 0.6B to 27B Parameters
The vllm-sr team has officially open-sourced Decision 2.0, a suite of state-of-the-art decision-focused models ranging from 0.6B to 27B on Hugging Face, deliver
On October 3, 2026, open-source AI researcher Xunzhuo Liu (@XunzhuoLiu) announced the official release of Decision 2.0, a family of state-of-the-art decision-specialized foundation models spanning parameter sizes from 0.6B to 27B, published as an open collection on Hugging Face.

Image source: @XunzhuoLiu / X
Targeting the surging demand for micro-decisions (System 1 Decision) within AI agent loops, this release provides developers with a full range of open-weight models, enabling fast local routing, classification, and scanning without relying on closed hosted APIs.
Single Forward-Pass Calibrated Probabilities Across All Model Scales
At the core of Decision 2.0 is an architecture purpose-built for bounded choices and structured label scoring rather than open-ended conversational generation.
Instead of incurring the latency and compute overhead of multi-token auto-regressive generation for simple binary checks, intent classification, or agent step selection, Decision 2.0 returns calibrated probabilities across allowed options in a single forward pass.
- 0.6B Lightweight Model: Runs comfortably on edge devices and local developer machines, engineered for workloads where ultra-low latency is critical.
- Mid-Tier Variants: Deliver a balanced trade-off between speed and accuracy for workflow branching and prompt pre-filtering pipelines.
- 27B High-Capacity Model: Handles high-precision evaluations for complex, multi-factor decision gates requiring maximum reliability.
All model weights are publicly accessible under the Hugging Face collection (vllm-sr/decision-20) for self-hosted deployment.
From On-Chain Auditing to Agent Routing: Workflows and Runtime Optimizations
Decision 2.0 is structured to deliver immediate operational efficiencies across agent orchestration and high-throughput data processing.
In scenarios such as real-time on-chain scanning or transaction monitoring across thousands of smart contracts, ultra-low latency outranks heavy deliberative reasoning. The compact 0.6B model allows engineers to flag anomalies in milliseconds, reserving the heavier 27B model strictly for downstream deep inspection.
The announcement noted that Decision 2.0 incorporates runtime-specific optimizations to accelerate inference on engines such as vLLM. This ensures that agent pipelines receive typed, structured decision scores with minimal latency.
Deployment Considerations and Hardware Requirements
Teams planning to integrate Decision 2.0 into production environments should consider its architectural focus and operational requirements:
- Targeted Scope: Decision 2.0 models are built specifically for structured classification, safety checks, and routing tasks; they are not intended for conversational prose or open-ended document generation.
- Hardware Footprint: While the 0.6B model fits into minimal local memory budgets, deploying the full 27B parameter variant locally requires dedicated GPU VRAM.
Sources
- Xunzhuo Liu Official Announcement on X: Introducing Decision 2.0
- Hugging Face Collection: vllm-sr / decision-20