Skip to content
INTERNET INFRASTRUCTURE & NETWORKS

Cloudflare Expands Its AI Capabilities with the Release of Clef and Clef-Flash Decision Models

Cloudflare has introduced its first natively trained artificial intelligence models, marking a significant entry into the emerging category of decision models. Dubbed Clef and Clef-flash, the new models are hosted on Workers AI and designed to produce bounded, structured outputs quickly, cheaply, and consistently. The release arrives amid growing industry interest in specialized AI decision architectures—a concept recently popularized by systems like Typesafe AI’s Jev—that bridge the gap between traditional rigid classifiers and open-ended, non-deterministic large language models.

Unlike general-purpose large language models, which excel at broad reasoning, text generation, and tool calls for complex agentic workflows, decision models are built to make specific classifications based on input probabilities. When integrated into a technical pipeline, they allow automated systems to process incoming data, evaluate probabilities across defined categories, and execute programmatic actions instantly. This capability allows software agents to operate with a high degree of autonomy, routing tickets, triggering escalations, or making real-time determinations without requiring a human in the loop for every intermediate step.

Cloudflare’s internal testing demonstrates the practical advantages of this architecture. The Threat Intelligence team has deployed Clef alongside browser rendering tools to classify website domains automatically. By passing a domain URL into the model, the system quickly assigns probability categories—such as identifying a page with a high likelihood of being an e-commerce platform while accurately ruling out malicious threats like phishing. In internal benchmarks, Clef completed this fetch, render, and classification cycle in roughly 2.2 seconds, outperforming the company’s fastest general large language model, which required 4.7 seconds while returning fewer categorical options.

The name Clef pays homage to both music theory and the company’s corporate identity. In musical notation, a clef sets the pitch for the notes that follow, mirroring how a decision model establishes context and guides subsequent programmatic actions. The "CF" within the name also points directly to Cloudflare.

Clef enters a specialized market with several distinct technical advantages over existing decision models. Most notably, Clef features an integrated vision encoder, allowing it to process and classify visual content alongside textual inputs, whereas preceding models like Jev have been limited to text-only classification. Furthermore, Clef provides a 64-kilotoken context window, doubling the 32-kilotoken limit found in comparable models and enabling systems to ingest larger state inputs for evaluation.

Performance evaluations released alongside the launch indicate that Clef and Clef-flash perform competitively across multiple standardized evaluation suites. Evaluated against the Jev Decision Index and other industry benchmarks, the models excelled in core decision-making categories such as tool retrieval, API accuracy, domain-specific classification, and security incident management. In benchmark comparisons covering various complex agentic workflows—including invoice processing, customer service triage, and security incident response—Clef models routinely scored near the top while maintaining significantly lower latency profiles than traditional large language models.

The performance edge is further amplified by Cloudflare’s infrastructure. Because Clef models are hosted directly on Workers AI, they leverage the company’s distributed network of edge GPUs. This proximity minimizes network latency, allowing developers to place decision models directly into the critical path of autonomous agents, combining fast classification with downstream actions handled by complementary large language models.

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

To ensure broad accessibility, Cloudflare is fully open-sourcing the Clef models on Hugging Face under an Apache 2.0 license, allowing developers to run and experiment with the weights locally. For enterprise users, the models are immediately available through the Workers AI API, backed by privacy guarantees ensuring that customer requests and responses are neither stored nor used for training by default.

Under the hood, Clef represents a departure from traditional autoregressive generation. While initial industry experiments adapted models like DiffusionGemma to output probabilities by exposing logprobs, Cloudflare chose Qwen as the base backbone for its new family. During inference, Clef uses Qwen for a prefill-only pass and then scores valid schema choices in parallel. Because the decision step is non-autoregressive, the model bypasses the need to generate intermediate text token by token, resulting in significantly faster execution times.

The architecture relies on a specialized two-stage attention routing process. Every valid choice extracts context relevant to the prompt, enabling individual field parameters to cross-attend with other fields and back to the original payload prior to scoring. By leveraging a lexical prior, the model preserves semantic intent across options, uniting option-specific evidence routing, joint cross-field attention, and schema-bound scoring into a single unified process.

Cloudflare post-trained the models by freezing specific Qwen checkpoints and jointly optimizing a routing head alongside rank-256 low-rank adapters. Training utilized label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration, drawing on synthetic datasets with permuted field orders, prompts, and schema structures. Additionally, the team developed Reinforcement Learning for Calibrated Decisions to serve as a secondary optimization target, providing partial credit for adjacent ordinal choices, rewarding fully precise record outputs, and applying reference penalties to prevent distribution shift.

Alongside the model release, Cloudflare is introducing a reinforcement learning and fine-tuning service designed to help organizations adapt Clef to proprietary workloads. While general-purpose models offer broad utility, many internal Cloudflare use cases—such as evaluating Trust and Safety submissions, triaging support requests, or distinguishing legitimate web crawlers from malicious bots—benefit from domain-specific customization.

The company plans to offer fine-tuning support initially through a hands-on approach via its forward-deployed engineer team, before transitioning the capability into a self-serve platform. This upcoming service will integrate several existing components of Cloudflare’s broader AI ecosystem, including AI Gateway for capturing request and response traffic, secure containers for reinforcement learning environments, and infrastructure adapted from the company’s acquisition of Replicate to facilitate custom model deployment.

The introduction of Clef marks Cloudflare’s first natively trained machine learning model from its Workers AI team, aligning with the company’s broader corporate mission to serve as infrastructure for the agent cloud. Organizations interested in exploring the technology can access the developer documentation, test the hosted models via Workers AI, download the open-source weights from Hugging Face, or reach out to the engineering team regarding custom fine-tuning partnerships.

Leave a Reply

Your email address will not be published. Required fields are marked *