clefdecide中文

Run Clef and Clef-Flash locally with Ollama

Ollama now ships both of Cloudflare’s decision models. Running them on your own machine needs no Cloudflare account and costs nothing per call; you just need hardware that fits the model. For now, Ollama runs them locally only.

Set up in three steps

  1. Install or update to Ollama 0.35.1 or later (needed for Clef’s image input).
  2. Pull a model: ollama pull clef-flash (about 11 GB) or ollama pull clef (about 18 GB).
  3. Call the local endpoint:
curl http://localhost:11434/v1/systemone -d '{
  "model": "clef-flash",
  "state": "Checkout has been failing for every customer for the last hour.",
  "questions": {
    "urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this request?",
      "criteria": { "billing": "Payments and refunds", "technical": "Outages and errors", "sales": "Plans and upgrades" }
    }
  }
}'

The response has the same shape as Workers AI: answers.urgent.noul is the probability of yes, and answers.team has choice, probabilities and confidence.

Which tag to pull

TagSizeWorks with /v1/systemone
clef-flash
same file: clef-flash:9b, clef-flash:9b-q8_0
11 GBYes
clef-flash:9b-mxfp812 GBNo (MLX)
clef-flash:9b-mlx-bf1619 GBNo (MLX)
clef
same file: clef:27b, clef:27b-q4_k_m
18 GBYes
clef:27b-q8_030 GBYes
clef:27b-nvfp418 GBNo (MLX)
clef:27b-mxfp831 GBNo (MLX)
clef:27b-mlx-bf1655 GBNo (MLX)

Ollama’s System One API only supports GGUF weights, so the MLX tags don’t work with the decision endpoint yet. Plan for more free memory than the download size.

Differences from Cloudflare Workers AI

Ollama (local)Cloudflare Workers AI
EndpointPOST /v1/systemone on port 11434POST …/ai/run/@cf/cloudflare/clef-flash
AuthNone for local requestsCloudflare API token, or the Worker AI binding
ImagesRaw base64; no data: prefixData URLs, up to 4
Size limits64 KiB without images, 32 MiB with imagesWhole body under 13 MiB
Long inputNever truncated; the request fails insteadTruncated to fit the context window
choice options2–26Per the published schema
CostYour hardware and electricity$0.09 / $0.24 per million input tokens
HostedLocal only for nowFully hosted

Port a Workers AI request to Ollama

Your questions stay the same. Change the URL and strip the data URL prefix from images:

// Same questions, two backends. Ollama wants raw base64 images, Workers AI wants data URLs.
const toOllama = body => ({
  model: body.model,                       // "clef-flash" or "clef"
  state: body.state,
  questions: body.questions,
  ...(body.images ? { images: body.images.map(url => url.replace(/^data:[^,]+,/, '')) } : {})
});

const res = await fetch('http://localhost:11434/v1/systemone', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify(toOllama(workersAiBody))
});
const { answers } = await res.json();

Local or Cloudflare?

Want to see results before downloading 11 GB? Try it in the playground (3 free runs). Published benchmarks for both models are in Clef vs Clef-Flash; request details are in the Clef API guide.

FAQ

Is Clef available on Ollama’s cloud?

Not as of 2026-10-03. Ollama’s docs say decision models are currently available locally only, and the System One API doesn’t support cloud models. If you don’t want to run hardware, use the Cloudflare Workers AI API.

Why does the MLX tag fail on my Mac?

Ollama’s System One API needs GGUF weights; MLX and Safetensors models are not supported. Pull the default clef-flash or clef tag instead.

Will local results match Cloudflare’s?

We haven’t compared them answer by answer. The default local tags are quantized (q4_k_m, q8_0), so probabilities can differ a little from Workers AI. Run your own samples on both before you switch.

Can I use ollama run?

Not yet. Ollama says decision models aren’t in the CLI or the official Python and JavaScript libraries yet. Call the API directly, or use the TypeSafe SDK.

Sources: ollama.com/library/clef-flash, ollama.com/library/clef, Ollama’s decision model guide and System One API reference, checked 2026-10-03. We have not tested it locally.