Run Clef and Clef-Flash locally with Ollama
Ollama now ships both of Cloudflare’s decision models. Running them on your own machine needs no Cloudflare account and costs nothing per call; you just need hardware that fits the model. For now, Ollama runs them locally only.
Set up in three steps
- Install or update to Ollama 0.35.1 or later (needed for Clef’s image input).
- Pull a model:
ollama pull clef-flash(about 11 GB) orollama pull clef(about 18 GB). - Call the local endpoint:
curl http://localhost:11434/v1/systemone -d '{
"model": "clef-flash",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": { "billing": "Payments and refunds", "technical": "Outages and errors", "sales": "Plans and upgrades" }
}
}
}'
The response has the same shape as Workers AI: answers.urgent.noul is the probability of yes, and answers.team has choice, probabilities and confidence.
Which tag to pull
| Tag | Size | Works with /v1/systemone |
|---|---|---|
clef-flashsame file: clef-flash:9b, clef-flash:9b-q8_0 | 11 GB | Yes |
clef-flash:9b-mxfp8 | 12 GB | No (MLX) |
clef-flash:9b-mlx-bf16 | 19 GB | No (MLX) |
clefsame file: clef:27b, clef:27b-q4_k_m | 18 GB | Yes |
clef:27b-q8_0 | 30 GB | Yes |
clef:27b-nvfp4 | 18 GB | No (MLX) |
clef:27b-mxfp8 | 31 GB | No (MLX) |
clef:27b-mlx-bf16 | 55 GB | No (MLX) |
Ollama’s System One API only supports GGUF weights, so the MLX tags don’t work with the decision endpoint yet. Plan for more free memory than the download size.
Differences from Cloudflare Workers AI
| Ollama (local) | Cloudflare Workers AI | |
|---|---|---|
| Endpoint | POST /v1/systemone on port 11434 | POST …/ai/run/@cf/cloudflare/clef-flash |
| Auth | None for local requests | Cloudflare API token, or the Worker AI binding |
| Images | Raw base64; no data: prefix | Data URLs, up to 4 |
| Size limits | 64 KiB without images, 32 MiB with images | Whole body under 13 MiB |
| Long input | Never truncated; the request fails instead | Truncated to fit the context window |
| choice options | 2–26 | Per the published schema |
| Cost | Your hardware and electricity | $0.09 / $0.24 per million input tokens |
| Hosted | Local only for now | Fully hosted |
Port a Workers AI request to Ollama
Your questions stay the same. Change the URL and strip the data URL prefix from images:
// Same questions, two backends. Ollama wants raw base64 images, Workers AI wants data URLs.
const toOllama = body => ({
model: body.model, // "clef-flash" or "clef"
state: body.state,
questions: body.questions,
...(body.images ? { images: body.images.map(url => url.replace(/^data:[^,]+,/, '')) } : {})
});
const res = await fetch('http://localhost:11434/v1/systemone', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(toOllama(workersAiBody))
});
const { answers } = await res.json();
Local or Cloudflare?
- Ollama on your machine: data never leaves the box, works offline, no per-call cost. You need a GPU or a Mac with plenty of memory, and you run it.
- Cloudflare Workers AI: nothing to run, pay per use, about $0.000017 for a short ticket on Clef-Flash. Better for production traffic that comes and goes.
Want to see results before downloading 11 GB? Try it in the playground (3 free runs). Published benchmarks for both models are in Clef vs Clef-Flash; request details are in the Clef API guide.
FAQ
Is Clef available on Ollama’s cloud?
Not as of 2026-10-03. Ollama’s docs say decision models are currently available locally only, and the System One API doesn’t support cloud models. If you don’t want to run hardware, use the Cloudflare Workers AI API.
Why does the MLX tag fail on my Mac?
Ollama’s System One API needs GGUF weights; MLX and Safetensors models are not supported. Pull the default clef-flash or clef tag instead.
Will local results match Cloudflare’s?
We haven’t compared them answer by answer. The default local tags are quantized (q4_k_m, q8_0), so probabilities can differ a little from Workers AI. Run your own samples on both before you switch.
Can I use ollama run?
Not yet. Ollama says decision models aren’t in the CLI or the official Python and JavaScript libraries yet. Call the API directly, or use the TypeSafe SDK.
Sources: ollama.com/library/clef-flash, ollama.com/library/clef, Ollama’s decision model guide and System One API reference, checked 2026-10-03. We have not tested it locally.