Supersonic Labs Releases Julia-1: 144.3M Open Decision Model That Runs on CPU

Supersonic Labs has released Julia-1, a 144.3M open decision model under Apache 2.0 designed for multi-choice classification on plain CPUs, mobile ONNX, and Web

tau · September 27, 2026

#Julia-1 #DecisionModel #SupersonicLabs #EdgeAI #OpenSource

Supersonic Labs Releases Julia-1: 144.3M Open Decision Model That Runs on CPU

Brazilian AI research lab Supersonic Labs officially released Julia-1 on September 26, 2026, an open-source 144.3-million-parameter decision and classification model licensed under Apache 2.0 and engineered to deliver low-latency inference on standard CPUs and on-device hardware.

Official announcement visual for Supersonic Labs Julia-1 lightweight open decision model

Image source: Supersonic Labs (@supersonicai)

While conventional large language models demand billions of parameters and substantial cloud GPU infrastructure to generate conversational responses, Julia-1 intentionally narrows its scope to a single operational requirement: evaluating candidate options and selecting the best structured outcome with high statistical certainty.

Architecture Focused on Multi-Choice Classification Rather Than Text Generation

The defining architectural choice behind Julia-1 is its complete abandonment of free-form, autoregressive text generation.

Instead of operating as an open-ended chatbot, Julia-1 receives a context block, a specific query or instruction, and between 2 and 20 candidate answers. The model computes normalized probability scores across all provided options and returns the single most suitable decision.

  • Specialized Decision Pipeline: By eliminating token-by-token sentence synthesis, the architecture avoids standard generative overhead, making it well suited for intent parsing, agent routing, triage, and ranking routines.
  • Deterministic Structured Output: Rather than wrestling with hallucinations, runaway verbose explanations, or formatting regressions, downstream systems receive explicit probabilities strictly bounded by the candidate set.
  • Permissive Open Weights: Distributed under Apache 2.0, the weights can be modified, embedded, and deployed across private commercial stacks without licensing bottlenecks or vendor lock-in.

For engineering teams currently routing routine agent decisions through expensive frontier reasoning models, Julia-1 offers a practical way to offload classification tasks to local infrastructure while minimizing inference cost and external API dependencies.

Ultra-Lightweight Edge Inference Across CPUs, Mobile, and WebGPU

Supersonic Labs prioritized hardware accessibility throughout the release, allowing developers to execute inference locally without requiring dedicated data center accelerators.

The model runs on Python 3.11+ using standard commodity CPUs or BF16-capable GPUs. In parallel with the primary weights, the team released an official ONNX export (Julia-1-ONNX) on Hugging Face, enabling direct client-side deployment inside web browsers via WebGPU as well as native execution on mobile platforms.

  • Apple M4 Silicon: Measured at a median inference latency of 33.15 ms.
  • Samsung Galaxy Tab SM-X510: Reached a median latency of 203 ms directly on consumer tablet hardware.
  • Network-Independent Execution: Because the model executes fully on-device, critical triage steps remain fully operational without internet connectivity or external API availability.

These response profiles mean that local CPU decision passes can often complete faster than the round-trip network latency required to ping a remote cloud inference provider.

Benchmark Limitations and Deployment Paths for Production Teams

In its technical release notes, Supersonic Labs emphasized transparent benchmarking, publishing specific architectural trade-offs and operational caveats alongside peak latency numbers.

First, inference latency scales noticeably as the candidate pool expands. While 2 to 20 options resolve in tens of milliseconds, performance degrades under large categorical loads; on the Banking77 benchmark, which requires evaluating across 72 potential intent labels, processing time rose to 3,713.54 ms. In production settings, large classification domains should be decomposed into hierarchical sub-trees rather than evaluated in a single flat pass.

Second, the lab's claim of achieving "5x faster performance than Jev" reflects a comparison against a remote Jev endpoint called across the network from France versus Julia-1 running locally on device. System architects should evaluate real-world speedups within their own unified network or hardware topologies.

Third, although Supersonic Labs announced plans for a managed hosted API, the service remains unavailable to the public at release time. Development teams seeking to integrate Julia-1 today must download the open weights (SupersonicLabs/Julia-1) or the ONNX runtime directly from Hugging Face for self-hosted deployment.

Sources