Cloudflare Clef Goes Open Weight: 64K Context, Multimodal AI and Ultra-Fast Decisions

The next major shift in artificial intelligence may not come from making large language models larger. It may come from giving AI systems a more specialized layer for making fast, bounded decisions.
Cloudflare’s Clef and Clef-flash decision models represent an important development in that direction. Designed specifically for structured decision-making, the models are intended to sit inside AI workflows and determine what should happen next, rather than generate long-form language or perform open-ended reasoning.
That distinction addresses a growing problem in agentic AI. Large language models are remarkably flexible, but using a general-purpose model for every decision can introduce unnecessary latency, cost and unpredictability. A specialized decision model can instead evaluate a defined question and return structured probabilities that software can immediately act upon.
Cloudflare is positioning Clef as an open-weight alternative in this emerging category, with models available through Workers AI and Hugging Face. The company also introduced a reinforcement learning fine-tuning service designed to let organizations adapt Clef to highly specific business workflows.
Together, these developments point toward a more modular architecture for AI agents, where large language models handle broad reasoning and generation while smaller decision models operate in the critical execution path.
What Is a Decision Model?
A decision model is designed to answer bounded questions rather than generate unrestricted text.
Consider an enterprise support system receiving thousands of customer requests. Instead of asking a large language model to compose a complete response, an application could ask a decision model several structured questions:
Is the request urgent?
Which department should receive it?
How severe is the issue?
Should the matter be escalated?
Is human review required?
The model can return typed outputs and probabilities that software can consume directly.
This architecture creates a division between reasoning and execution.
An LLM might understand a complex customer message and generate a plan. A decision model can then determine which predefined path the workflow should take.
That distinction becomes increasingly important as AI agents move from conversational interfaces into systems capable of taking actions.
An agent does not always need another paragraph of generated text. Often it needs a reliable answer to a narrow question.
Why Cloudflare Built Clef
Cloudflare's motivation is closely connected to the emerging agentic AI architecture.
Modern agents frequently depend on multiple models and tools. A general-purpose LLM can interpret information, call APIs, write code and reason through complicated tasks. But using such a model for every small classification introduces computational overhead.
Clef attempts to provide a specialized control layer.
For example, Cloudflare tested Clef within its threat intelligence workflows to classify website domains. The model can assess multiple characteristics simultaneously, producing probabilities for categories such as fashion, ecommerce or phishing.
In one cited workflow, Clef took 2.2 seconds to fetch, render and classify a website. Cloudflare reported that its fastest general LLM in the same workflow, gpt-oss-120b, required 4.7 seconds and produced only two classifications.
The significance is not simply that one model is faster than another. It demonstrates how a purpose-built decision layer can reduce the computational burden associated with repetitive classification inside an agentic system.
At scale, such latency differences can affect user experience, infrastructure requirements and operating costs.
Clef Versus Traditional Large Language Models
Large language models are optimized for breadth.
They can generate natural language, reason over diverse information, interpret instructions and interact with tools. Their flexibility is their greatest advantage.
It is also their weakness when a workflow requires predictable structured decisions.
A decision model narrows the output space. Instead of generating arbitrary language, Clef evaluates valid choices defined by a schema and assigns probabilities to those options.
This produces several architectural advantages:
Capability | General-purpose LLM | Decision model |
Open-ended generation | Strong | Not the primary purpose |
Structured classification | Capable | Core function |
Output predictability | Variable | Strongly constrained |
Decision latency | Can be relatively high | Optimized for fast decisions |
Tool-oriented reasoning | Strong | Designed to support routing |
Specialized workflow decisions | Requires prompting | Native design objective |
Multimodal input | Model dependent | Clef supports visual inputs |
Fine-tuning for specific decisions | Possible | Central use case |
The most effective architecture may therefore not be LLM versus decision model.
It may be LLM plus decision model.
The larger model can reason broadly while Clef handles frequent, tightly defined decisions.
Cloudflare's Technical Approach
Clef is built around Qwen-based backbones, with Cloudflare using a frozen Qwen3.8-27B model for Clef and Qwen3.5-9B for Clef-flash.
The important architectural difference is what happens during inference.
Clef uses the backbone for a prefill pass and then evaluates valid schema choices in parallel. The decision process is non-autoregressive, meaning the system does not need to generate a sequence of intermediate tokens one after another before reaching its structured output.
Instead, the architecture derives scores directly from representations produced by the underlying model.
Cloudflare describes a specialized attention-routing mechanism in which different fields can extract relevant information from the input and interact with other fields before the final schema choices are scored.
This approach is significant because autoregressive generation is naturally sequential. Even when a model ultimately needs to return only a few structured values, conventional generation can require token-by-token computation.
A decision model can eliminate much of that unnecessary work when the output space is known in advance.
Clef's Benchmark Position
Cloudflare reports strong results across a range of decision-oriented evaluations.
On the Jev Decision Index, Clef scored 98.47 on BFCL case exact, 69.19 on ToolRet nDCG@10, 91.93 on API-Bank accuracy, 94.20 macro-F1 on BANKING77, and 97.43 macro-F1 on CLINC150+OOS.
Clef-flash performed particularly strongly on several individual benchmarks, including BFCL, API-Bank and home appliance decision tasks.
Cloudflare also tested its models against TypeSafe's workflow evaluations.
Workflow | Clef | Clef-flash | Jev |
Invoice processing | 64.7 | 57.1 | 61.8 |
Customer service | 76.3 | 77.0 | 76.0 |
Security incidents | 62.9 | 61.7 | 61.7 |
Agent trace observability | 68.5 | 69.8 | 71.6 |
Cloudflare says its models surpassed Jev in three of the four workflow categories.
However, benchmark results should be interpreted in context. The reported comparisons are primarily vendor-reported evaluations, and independent reproduction is important before treating benchmark leadership as definitive. Decision-model performance can also vary substantially according to schema design, input distribution and deployment conditions.
The more durable question is whether specialized decision models deliver measurable improvements in real production workloads.
Speed Is the Strategic Advantage
Cloudflare reports a median latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash across its latency evaluation.
For comparison, the reported median for Jev was 524.1 milliseconds.
At the 95th percentile, Clef-flash reached 122.4 milliseconds, while Clef recorded 238.6 milliseconds and Jev 536.0 milliseconds.
The difference between median and tail latency is particularly important for enterprise agents.
A workflow that makes thousands or millions of decisions can be affected by the slowest portion of requests. Lower latency makes it easier to put a decision model directly into an agent's critical path.
Cloudflare's edge infrastructure provides another potential advantage. Hosting Clef on Workers AI allows inference to occur on Cloudflare's distributed computing infrastructure, reducing network distance in applications already operating on the platform.
This creates a potentially powerful combination: specialized inference plus edge deployment.
Open Weight Changes the Economics of Decision Models
One of Clef's most consequential characteristics is its availability as an open-weight model.
Cloudflare makes Clef available through Hugging Face under Apache 2.0 terms, allowing organizations to download and experiment with the model rather than relying exclusively on hosted inference.
This is particularly important for enterprises with privacy, latency or infrastructure requirements.
But open weight does not mean effortless local deployment.
According to Cloudflare's information reported in the supplied research, Clef-flash requires approximately 41 GB of GPU VRAM under the stated single-concurrency, 64K-context assumptions, while Clef requires approximately 85 GB.
That makes the models substantially more accessible than some frontier-scale systems, but still places them beyond ordinary consumer hardware.
Cloudflare's hosted option therefore remains important for organizations that want the decision-model architecture without managing large GPUs themselves.
Clef Adds Vision to the Decision Model Category
Another differentiator is multimodal capability.
Cloudflare says Clef includes a vision encoder, allowing it to classify visual information in addition to text.
This expands the potential range of decision workflows.
An enterprise system could potentially evaluate:
Screenshots
Documents
Product images
Visual safety submissions
Web pages
Security evidence
Images associated with customer requests
The practical advantage is that classification can occur closer to the point where information enters a workflow.
Instead of converting every visual input into text before making a decision, a multimodal decision model can potentially reason directly over the relevant visual content.
Cloudflare also describes Clef as supporting a 64K context window, providing considerable space for contextual information and decision criteria.

From Generic Models to Specialized AI Infrastructure
The rise of decision models reflects a broader maturation of AI infrastructure.
Early generative AI deployments often centered around one large model. As applications become more complex, developers increasingly need multiple specialized components.
A production agent may eventually resemble a software stack containing:
A general-purpose reasoning model
Specialized decision models
Retrieval systems
Tool-use interfaces
Memory systems
Security and policy layers
Monitoring and evaluation infrastructure
In this architecture, Clef is not necessarily intended to replace the LLM. It can act as a control mechanism surrounding it.
That makes decision models potentially valuable infrastructure rather than standalone AI products.
Reinforcement Learning Could Make Clef More Valuable
Cloudflare's second major announcement is its reinforcement learning fine-tuning platform.
The motivation is straightforward: general decision models are useful, but organizations often have proprietary decision patterns.
Cloudflare itself has more than 15 years of network data across different domains and says its internal teams have accumulated extensive labeled decisions.
Potential applications include Trust & Safety classification, support triage and determining whether web crawlers represent desirable or undesirable bots.
A generic model can provide broad capability. Fine-tuning can specialize that capability for a particular environment.
This introduces an important trade-off. More specialization can improve accuracy and efficiency within a defined workload, but excessive specialization can reduce general-purpose performance.
The ideal model is therefore not necessarily the most capable one in the abstract. It is the model that performs reliably against the specific distribution of decisions an organization actually needs to make.
Cloudflare's Emerging RL Pipeline
Cloudflare is assembling its reinforcement learning service around several existing platform components.
The proposed workflow connects:
AI Gateway for capturing AI traffic and creating datasets
Workers AI for generating rollouts
Cloudflare Containers for reinforcement learning sandboxes
A new Trainer component for updating model weights
Workers AI and Bring Your Own Model capabilities for deployment
This is strategically significant because fine-tuning is often fragmented across data collection, training, evaluation and deployment systems.
By combining these steps into one platform, Cloudflare is attempting to make custom decision models part of the application development lifecycle.
The longer-term objective is a feedback loop.
Applications generate decisions. Those decisions create data. The data becomes training material. Fine-tuning improves the model. The updated model returns to production, where new outcomes generate additional feedback.
That is the foundation of continuously improving enterprise AI systems.
The Business Case for Decision Models
For enterprises, the value proposition is ultimately economic.
Suppose an organization uses a large model for every classification task. The organization is paying for capabilities it may not need on every request.
A specialized decision model can potentially reduce:
Inference latency
Compute consumption
Token usage
Network overhead
Operational complexity
Human escalation volume
The strongest applications will be workflows where decisions occur frequently and the acceptable outputs are well defined.
Security classification, customer routing, fraud detection, moderation, workflow prioritization and automated quality checks are natural candidates.
The economics become particularly compelling when decisions occur at high volume.
The Risks and Trade-Offs
Decision models do not eliminate the fundamental problems associated with AI.
Probability is not certainty. A model can produce a highly confident but incorrect classification.
That means organizations still need thresholds, escalation policies, monitoring and human review for high-impact decisions.
There is also a risk of over-specialization. A model optimized for historical organizational data can inherit the biases or blind spots contained within that data.
Fine-tuning can also create distribution-shift problems if the production environment changes.
For that reason, decision models should be treated as components in a governed system rather than autonomous sources of truth.
The strongest architecture is likely to combine model confidence with business rules, validation and escalation mechanisms.
What Clef Means for the Future of AI Agents
Cloudflare's Clef launch signals a broader change in how developers may build AI agents.
The future agent is unlikely to be a single enormous model responsible for every task.
Instead, AI systems may increasingly become orchestrated collections of specialized models, each optimized for a specific part of the workflow.
Large language models can remain responsible for complex reasoning and generation. Decision models can provide fast routing and classification. Smaller specialized models can handle vision, speech, anomaly detection or domain-specific prediction.
This architecture resembles modern software engineering more than the early chatbot paradigm.
Intelligence becomes modular.
Cloudflare's strategy is particularly interesting because it combines three elements, open-weight models, edge inference and reinforcement learning. Each addresses a different constraint in enterprise AI.
Open weights provide control. Edge deployment can reduce latency. Reinforcement learning enables specialization.
Together, they create a pathway toward AI systems that are faster, more adaptable and more deeply integrated into operational software.
The Rise of the AI Decision Layer
Clef is important not because it attempts to become another general-purpose chatbot, but because it targets a narrower problem that becomes increasingly important as AI agents become more autonomous.
Agents need to decide.
They need to determine which tool to call, which workflow to activate, whether an event is urgent, whether a request requires escalation and whether an action should proceed automatically or require human intervention.
Using a large language model for every one of these decisions may be unnecessarily expensive and slow.
Cloudflare's Clef provides an alternative, a specialized decision layer capable of producing structured probability-based outputs with low latency. Its open-weight availability, multimodal capabilities, 64K context window and API compatibility with Jev make it particularly relevant to developers experimenting with this emerging architecture.
Its reinforcement learning platform may prove even more consequential. If enterprises can continuously customize decision models around their own data and workflows, AI agents could become significantly more reliable within specific operational environments.
For technology researchers and analysts such as Dr. Shahid Masood and the expert team at 1950.ai, the emergence of decision models represents a larger transition in artificial intelligence, from monolithic generative systems toward modular AI infrastructure.
The next generation of enterprise AI may not be defined by one model doing everything.
It may be defined by many models doing exactly what they are best at.
Further Reading / External References
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
Cloudflare tries to outplay Jev with open-weight Clef models





Comments