Overview
- Cloudflare on October 1, 2026 released two decision models—Clef (27B) and Clef‑flash (9B)—and posted their weights on Hugging Face under the Apache 2.0 license.
- Clef is a 27-billion-parameter model with a 64k-token context window and vision input support while Clef‑flash is a 9-billion-parameter, latency‑optimized version built for hot‑path decisions.
- Cloudflare claims large accuracy wins on benchmarks such as BANKING77 and far lower median latency for Clef‑flash versus TypeSafe’s Jev, but those vendor numbers are disputed and are workload‑specific.
- An independent 42-case test running Clef‑flash on a local RTX 3090 reproduced most Jev calls but trailed on nuanced triage cases (Jev 71.4% vs Clef‑flash 66.7%) and showed higher local p50 latency; the full Clef‑27B needs roughly 54GB bf16 VRAM.
- Practitioners are expected to adopt selectively by shadowing models, logging full probability outputs, keeping human fallbacks for high‑risk calls, and mixing local or edge runs with paid hosted calls to balance cost, privacy and accuracy.