clefdecideEnglish

用 Ollama 本地运行 Clef 和 Clef-Flash

Ollama 已经上架 Cloudflare 的两个决策模型。在自己的机器上运行不需要 Cloudflare 账号,也不按 token 计费;代价是你要准备能装下模型的硬件。目前 Ollama 只支持本地运行。

三步跑起来

  1. 安装或升级到 Ollama 0.35.1 或更高版本(Clef 读图需要这个版本)。
  2. 下载模型:ollama pull clef-flash(约 11 GB)或 ollama pull clef(约 18 GB)。
  3. 调用本机接口:
curl http://localhost:11434/v1/systemone -d '{
  "model": "clef-flash",
  "state": "Checkout has been failing for every customer for the last hour.",
  "questions": {
    "urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this request?",
      "criteria": { "billing": "Payments and refunds", "technical": "Outages and errors", "sales": "Plans and upgrades" }
    }
  }
}'

返回格式和 Workers AI 基本一致:answers.urgent.noul 是“是”的概率,answers.team 里有 choice、probabilities 和 confidence。

选哪个版本

版本大小能用 /v1/systemone
clef-flash
同一文件:clef-flash:9b, clef-flash:9b-q8_0
11 GB可以
clef-flash:9b-mxfp812 GB不行(MLX)
clef-flash:9b-mlx-bf1619 GB不行(MLX)
clef
同一文件:clef:27b, clef:27b-q4_k_m
18 GB可以
clef:27b-q8_030 GB可以
clef:27b-nvfp418 GB不行(MLX)
clef:27b-mxfp831 GB不行(MLX)
clef:27b-mlx-bf1655 GB不行(MLX)

Ollama 的 System One 接口只支持 GGUF 权重,MLX 版本暂时不能用于决策接口。内存最好比文件大小多留一些余量。

和 Cloudflare Workers AI 的区别

Ollama(本地)Cloudflare Workers AI
接口POST /v1/systemone(本机 11434 端口)POST …/ai/run/@cf/cloudflare/clef-flash
认证本地请求不需要 API KeyCloudflare API Token 或 Worker 的 AI 绑定
图片纯 base64,不能带 data: 前缀data URL,最多 4 张
大小限制无图 64 KiB,有图 32 MiB整个请求小于 13 MiB
超长输入不截断,直接报错截断到上下文长度
choice 选项数2–26 个按官方 schema
费用自己的硬件和电费每百万输入 token $0.09 / $0.24
云端目前只能本地运行托管服务,不用管机器

把 Workers AI 请求改成 Ollama 请求

问题定义不用改,只需换地址,并把图片的 data URL 前缀去掉:

// Same questions, two backends. Ollama wants raw base64 images, Workers AI wants data URLs.
const toOllama = body => ({
  model: body.model,                       // "clef-flash" or "clef"
  state: body.state,
  questions: body.questions,
  ...(body.images ? { images: body.images.map(url => url.replace(/^data:[^,]+,/, '')) } : {})
});

const res = await fetch('http://localhost:11434/v1/systemone', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify(toOllama(workersAiBody))
});
const { answers } = await res.json();

本地还是 Cloudflare?

下载 11 GB 之前想先看看效果?在在线试用里直接跑,免费 3 次。两个模型的公开评测见 Clef 与 Clef-Flash 对比,接口细节见中文 API 教程。

常见问题

Ollama 有云端版的 Clef 吗?

截至 2026-10-03 没有。Ollama 文档写明决策模型“目前只能在本地使用”,System One 接口也不支持云端模型。不想自己部署,就用 Cloudflare Workers AI 的 API。

为什么我的 Mac 上 MLX 版本报错?

Ollama 的 System One 接口要求 GGUF 权重,MLX/Safetensors 版本不支持。请拉取 clef-flash 或 clef 默认版本。

本地结果和 Cloudflare 上的一样吗?

我们没有逐条对比过。本地默认是量化版本(例如 q4_k_m、q8_0),概率可能和 Workers AI 上略有不同。上线前用你自己的样本两边都跑一遍。

可以用 ollama run 吗?

还不行。Ollama 说明决策模型暂时不在 CLI 和官方 Python/JS 库里,请直接调 API,或用 TypeSafe SDK。

资料来源:ollama.com/library/clef-flash、ollama.com/library/clef、Ollama 决策模型文档与 System One API,核对于 2026-10-03。我们没有在本地实测。