Particle.news

Cloudflare Releases Clef Open Decision Models

The launch gives developers a local, open-weights option for structured decision calls that reduces cost and preserves on-site data.

Overview

  • Cloudflare announced Clef and Clef-flash on October 1, offering open-weights decision models built on Qwen backbones that return calibrated probabilities for closed-choice and binary questions.
  • An independent 42-case benchmark run by a DEV Community author compared Clef-Flash 9B running locally on an RTX 3090 against the hosted TypeSafe Jev API and found overall accuracy of 66.7% for Clef-Flash and 71.4% for Jev.
  • The per-task breakdown showed parity on GUI action choices (10/10) and shared supervision misses (8/12 each) while Jev led on message triage (12/20 versus Clef 10/20), highlighting where model size and nuance matter.
  • Latency and deployment trade-offs differed: Jev’s median latency was 225 ms over the network versus Clef-Flash’s 315 ms locally, Clef-27B requires roughly 54 GB of VRAM, and the decision head currently needs the PyTorch transformers path rather than GGUF/llama.cpp runtimes.
  • Practical impact for teams is clear: running decision models locally can cut API cost and keep sensitive agent state on-premise, but teams must empirically validate calibration per workload and choose which calls stay local and which use hosted, conservative fallbacks such as shadow mode or human vetoes.