Overview
- Cloudflare released two decision models, Clef (27B) and Clef‑flash (9B), and published the open weights on Hugging Face after announcing them on October 1, 2026.
- Clef is a 27‑billion‑parameter model with a 64k‑token context and multi‑input support while Clef‑flash is a 9‑billion‑parameter, latency‑focused variant with a vision encoder and lower VRAM needs.
- Cloudflare says both models are compatible with the Jev API, lists per‑input pricing for hosted edge use, and is offering edge GPU hosting plus a reinforcement‑learning fine‑tuning platform.
- An independent local benchmark running Clef‑flash on a 24GB RTX 3090 found the 9B model agreed with the hosted Jev API on 83% of calls but trailed on nuanced triage cases, recording overall accuracy of 66.7% versus Jev’s 71.4%.
- Teams picking between hosted Jev and Clef must weigh tradeoffs in accuracy, latency, cost, privacy, and hardware: the full Clef‑27B needs roughly 54GB of bf16 VRAM while the 9B Clef‑flash runs on a single 24GB consumer GPU and keeps data on‑premise.