Cloudflare Launches Clef and Clef-flash
Cloud infrastructure firm Cloudflare wants AI agents to react faster during their tasks. To that end, it has launched two open-source decision models that it trained itself: Clef and Clef-flash.
Both models are fully integrated into Workers AI, the company’s serverless GPU platform. Cloudflare also offers the model weights on Hugging Face under the permissive Apache 2.0 license. With this move, it challenges the Jev decision model from TypeSafe head-on.
Cutting Out Text Generation
Clef differs sharply from familiar large language models. It is not built for chat, long reasoning, or free-form text. Instead, it is a decision model focused on judgment and classification.
The Old Way
In a traditional AI agent workflow, the system may need to decide whether a support email is urgent. Usually, it must wait while a large language model generates text token by token. Then the program parses that text for a specific JSON format or keyword.
The Clef Approach
Clef skips that slow autoregressive generation step entirely. Developers define the options and scoring criteria in advance. After reading the input state, Clef uses cross-attention inside its neural network. It computes all options in parallel and returns a probability score for each one.
Striking Speed Gains
This change in architecture brings a remarkable boost in speed. Cloudflare published benchmarks across 43 evaluation tasks:
- The smaller Clef-flash showed a median latency of about 38.8 milliseconds.
- The larger Clef reached 209.3 milliseconds.
- By comparison, the rival Jev took 524.1 milliseconds.
Built on Qwen With Vision Support
Under the hood, Cloudflare trained the models on Alibaba Cloud’s open-source Qwen series. Clef builds on the 27-billion-parameter Qwen3.8-27B. Clef-flash, tuned for top speed, builds on the 9-billion-parameter Qwen3.5-9B.
Notably, the models keep Qwen’s vision encoder. Therefore, Clef goes beyond text-only decisions. It can take text, JSON, and up to four images as its input state.
Drop-In Compatibility With Jev
To ease the switch for developers, the Clef API is fully compatible with Jev’s System One API specification. Developers only need to change the endpoint name. They can then move existing Jev apps to Clef for testing without friction.
Reinforcement Learning Fine-Tuning
Cloudflare is also launching a reinforcement learning (RL) fine-tuning service. At first, its in-house forward deployed engineer (FDE) team will help enterprise customers tune models for specific workflows. Later, the company plans a fully self-service platform for data preparation and model deployment.
Lightning-Fast Decisions Unlock Real Automated Routing
In real software engineering, about 80% of AI agent tasks need no long “thinking” from a large language model. They simply need a clear Boolean value, a yes or a no. Or they need a classification label, such as “send to sales” or “mark as spam.”
Clef targets this pain point precisely. Decision latency near 40 milliseconds means AI judgment can sit directly on a system’s hot path. In other words, the system can finish AI routing the moment a user sends a request, with no noticeable lag.
Combined with Workers AI compute across Cloudflare’s global edge nodes, this kind of specialized model gives answers without filler. It may well become a new infrastructure standard for large-scale enterprise AI workflows. In that role, it could replace general-purpose large language models at the 10-billion-parameter scale.
Support Our Threat Intelligence
Find our tech and OS security coverage helpful? Support our work today and unlock a 100% ad-free reading experience!