Ollama Officially Supports Cloudflare Open-Source Clef and Clef-Flash Decision Models

Ollama now officially supports local serving of Cloudflare's open-source 27B and 9B multimodal decision models, Clef and Clef-Flash. Powered by a Jev-compatible

tau · October 3, 2026

#ollama #cloudflare #clef #decision-model #local-ai #dev-tools

Ollama Officially Supports Cloudflare Open-Source Clef and Clef-Flash Decision Models

On October 3, 2026, local AI runtime platform Ollama announced official local serving support for Cloudflare's open-source multimodal decision models, Clef and Clef-Flash. Developers running Ollama 0.35.1 or later can now pull and serve both models locally with straightforward CLI commands (ollama pull clef and ollama pull clef-flash).

Ollama announcement graphic supporting Cloudflare Clef and Clef Flash open-source decision models

Image source: Ollama (@ollama) / X

Unlike conventional autoregressive large language models that generate text token-by-token, the Clef family belongs to the System One category of decision models. They score every option across structured questions simultaneously in a single non-autoregressive forward pass, returning typed probabilities without generation delay or schema parsing errors.

Clef 27B vs. Clef-Flash 9B: Architecture and Specifications

Open-sourced by Cloudflare's Workers AI team under the Apache 2.0 license, the Clef model family is divided into two distinct tiers optimized for precision and latency respectively:

  • Clef (27B, ollama pull clef): A flagship decision model fine-tuned from Qwen3.8-27B. It is designed for complex policy enforcement, multi-attribute document verification, and scenarios demanding highest-precision routing.
  • Clef-Flash (9B, ollama pull clef-flash): A lightweight, high-velocity model fine-tuned from Qwen3.5-9B. In Cloudflare's Decision Index benchmarks, Clef-Flash clocked a median latency of 38.8 ms—roughly five times faster than Clef 27B (209.3 ms)—making it ideal for latency-critical hot-path routing and gateway firewalls.
  • 64K Context Window: Clef provides a 64K token context window—doubling the 32K window of previous text decision models like Jev—allowing teams to pass extensive conversational histories or dense operational states in a single request.
  • Built-in Multimodal Vision Encoder: In addition to plain text and structured JSON objects, Clef includes a native vision encoder that evaluates screenshots, receipts, forms, and photos (PNG, JPEG, WebP) directly alongside the input state.

Single Forward Pass Architecture and the /v1/systemone REST API

Using general-purpose generative LLMs for routing often introduces brittle JSON parsing pipelines, schema hallucination, and unpredictable generation latency.

Clef solves this by evaluating the input state against typed questions in one unified pass. Because no free-form text tokens are generated, parsing failures are eliminated entirely, and applications receive exact probability distributions and confidence scores for every permitted choice.

Ollama integrates Clef using a dedicated POST /v1/systemone endpoint that remains 100% compatible with Typesafe Jev and System One API specifications.

# Example local request to Clef via Ollama
curl http://localhost:11434/v1/systemone -d '{
  "model": "clef",
  "state": {
    "ticket": "I was charged twice for my subscription. Please refund the extra amount.",
    "attached": "Attached screenshot of bank statement showing duplicate charge."
  },
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this support ticket?",
      "criteria": {
        "billing": "Payments, invoices, and refunds",
        "technical": "System outages, bugs, and integration issues",
        "other": "General inquiries"
      }
    },
    "refund": {
      "type": "noul",
      "instructions": "Does the customer explicitly request a refund?"
    },
    "urgency": {
      "type": "score",
      "instructions": "How urgent is this ticket?",
      "criteria": ["Routine", "Soon", "Urgent"]
    }
  }
}'

The API supports three question primitives: multi-choice selection (choice), boolean conditions (noul), and ordinal scale ratings (score), allowing up to 64 distinct questions evaluated in parallel within a single request.

Production Use Cases and Implementation Caveats

Running Clef locally via Ollama enables cost-effective, private automation across critical engineering and operational workflows:

  • Automated Support Triage and Routing: Classify ticket urgency and assign ownership to the right operational squad within tens of milliseconds without human-in-the-loop intervention.
  • AI Agent Guardrails and Tool Gating: Perform low-latency safety checks before an autonomous agent executes high-impact shell commands or database mutations.
  • Local Privacy and Sensitive Data Filtering: Evaluate whether desktop screen captures or session logs contain sensitive personal information before they enter persistent indexing pipelines.

Implementation Caveats:

  1. Minimum Version Requirement: Running Clef or Clef-Flash requires Ollama version 0.35.1 or later.
  2. Client Library Status: Decision endpoints are not currently exposed in the Ollama interactive CLI chat mode or official Python/JavaScript SDK wrappers. Developers should interact directly via the HTTP REST API or the TypeSafe SDK.
  3. Specialized Scope: Clef models are specialized decision engines and do not generate free-form text or conversational responses. They should be deployed exclusively for routing, classification, and gating tasks.

Sources