What looks wrong?

We say this article was researched and checked. If it is wrong, we want the counter-example.

Skip to content
Ahmad Hassan

Oct 2, 202613 min read

Cloudflare Clef is a new family of decision models. It’s engineered to replace general-purpose chat interfaces with strict tool-calling and routing logic for autonomous agents, much like how developers utilize Activepieces to orchestrate complex workflows through automated triggers and actions.

While legacy deployments use chat models to guess at intent, Clef is a 27B multimodal decision model. It turns a specific state and a schema of typed questions into actionable decisions.

Launch specialized decision models for agents

The shift from chat models to decision models

To reduce latency and token waste in automated workflows, decision models eliminate the "chatter" of LLMs.

PwC found that 264 out of 300 organizations surveyed are increasing their AI budgets. That means the pressure to move from experimental chatbots to production-ready agents is now a primary financial directive.

The path from AI budget to agent ROI

Only 138 of the 209 organizations that have adopted agents are reporting realized value, according to PwC.

This gap exists because general models often fail to execute precise API calls reliably. Activepieces bridges this by providing integrations that Clef can call as structured tools, ensuring the model outputs a specific Tool Call rather than a conversational response.

The model acts as a logic gate rather than a creative writer under this architecture.

[Flow diagram of the 'Decision Model' architecture: a 'State' object (JSON) and a 'Schema' of typed questions enter the Clef model, which outputs a specific 'Decision' (Tool Call) rather than a chat]

A workflow with three steps: Chat UI for human input, Extract Structured Data using Utility AI, and a third step below.

Developers prevent the hallucinated parameters that frequently break agentic chains by constraining the output to a validated schema.

Clef-1.5b and Clef-Flash availability

Cloudflare offers two distinct tiers to balance reasoning depth against execution cost, allowing developers to optimize their infrastructure spend based on the complexity of specific tasks, which means teams can allocate resources more efficiently without overpaying for simple operations.

Activepieces pricing page displaying four subscription tiers with features and costs.

Clef is a 27B parameter multimodal model hosted on Hugging Face for complex routing where high-dimensional state analysis is required. Clef-Flash is a 9B multimodal model available via Cloudflare Workers AI. It’s designed for high-throughput, low-latency decision tasks.

Post-trained to handle rapid-fire tool selection, Clef-Flash allows a standard worker to process state transitions in milliseconds so that real-time systems remain responsive under load.

Everything below works on Activepieces' free plan. Start without code or a credit card.

Performance benchmarks for tool calling and reasoning accuracy

On function-calling benchmarks, Clef outperforms general-purpose models. Cloudflare optimized its weights for structural precision rather than conversational fluidity. While frontier models like GPT-6 Astra handle open-ended reasoning, Clef is designed to turn a state and a schema of typed questions into decisions.

A row of different sized athletes at a starting line; while the others are in various relaxed poses or stretching, one…

Superior accuracy in function calling benchmarks

By narrowing their operational scope to specific schema matching, specialized models achieve higher reliability in regulated environments.

On the Hugging Face banking-specific reasoning benchmark, which tests a model's ability to map user intent to complex financial API calls, the performance gap between specialized and general models is stark.

Model Accuracy Score
Cloudflare Clef 94.2
Clef-Flash 90.9
Llama 3.1 8B 84.8
Llama 3.1 70B 79.7
GPT-4o-mini 74.3
Jev 14.3

Prices and plan limits checked against developers.cloudflare.com and huggingface.co and blog.cloudflare.com and huggingface.co on October 2, 2026.

Accuracy on banking-specific reasoning

1 8B across Python and REST API test cases.

Reliable outputs for an ERP system are easier to provide with a 27B specialized model than a 70B general model. This allows teams to downsize their compute requirements without sacrificing audit compliance.

Reduced latency for agentic loops

Multiple model calls are required to complete a single task in agentic workflows. That makes inference speed the primary bottleneck for user experience.

Cloudflare architected Clef 27B and Clef-Flash 9B for the Workers AI global network, hosting both models there.

Clef achieves these accuracy scores at a fraction of the parameter count of Llama 3.1 70B. A customer-facing bot can verify a transaction and trigger a refund API before the user session times out.

Specialized reasoning for complex domain tasks

By prioritizing structural integrity over conversational filler, Clef handles domain-specific logic. Multi-step financial or technical operations follow strict execution paths without the creative drift typical of larger models.

To prevent the silent data corruption that breaks automated reconciliation pipelines, Cloudflare architected Clef to fail safely or request missing parameters. This prevents the "fill in the blanks" behavior seen in general-purpose models when a schema is ambiguous.

Outperforming GPT-6 Luna in logic

Because it treats every prompt as a computational graph rather than a linguistic suggestion, Clef maintains higher fidelity to nested logic constraints.

When a workflow requires a conditional check, such as verifying a user’s regional compliance before triggering a cross-border wire, general models like GPT-6 Luna may prioritize the "helpfulness" of completing the request over the strict "if-then" requirements of the API documentation.

The stylistic nuances of human speech are ignored by Clef’s focus on tool-calling. The resulting JSON output matches the target system's requirements exactly. It reduces the need for expensive retry logic in the application layer.

The following table distinguishes between the two primary Clef variants to help architects select the appropriate latency-to-reasoning ratio for their specific stack.

Feature Clef Clef-Flash
Base Model Qwen 3.8-27B Qwen 3.5-9B
Context Window 64k —
Primary Use Case Complex Reasoning High-Speed Routing

Evaluating interdependent variables in a single pass requires the depth of the standard Clef model. Clef-Flash is a high-throughput gatekeeper for simpler redirection tasks.

The advantage of specialized training sets

Precision in these models is ensured by training data that favors structured documentation and code over general web crawl data. This eliminates the "hallucination noise" common in models trained to mimic human chat.

A small, rectangular card representing a JSON object containing structured text, positioned next to a list titled 'Schema'…

Developers gain a tool that understands the difference between a string and a boolean in a production environment by isolating the model’s learning to high-quality API specifications and logical proofs.

This specialization ensures that a banking bot won't confuse a transaction ID with a currency amount. That’s a mistake that would otherwise require manual intervention from a human auditor to correct.

Easier to see it running than to read about it: set it up free, no card.

Connect Cloudflare Clef to your automation workflows today

By functioning as a high-speed router that maps natural language intents to specific JSON-defined functions, Cloudflare Clef integrates into existing agentic stacks.

By offloading the logic of tool selection to a specialized decision model, you avoid the latency penalties of using a general-purpose model like GPT-6 Astra for simple routing tasks.

Activepieces makes these decision models actionable by exposing every connected integration as a tool schema via its per-project MCP server.

Because many integrations are community-contributed, any action defined in the open-source repository at packages/pieces is automatically available for Clef to call as a validated tool. This eliminates the need to manually re-integrate your catalog for every new agent.

Complex reasoning models only consume compute cycles when the decision model identifies a task requiring deep analysis under this architecture.

Step 1: Get your Cloudflare API credentials

A Cloudflare Workers AI API Token is required to access these decision models. It acts as the authentication layer for all inference requests.

  1. You must generate this token within the Cloudflare dashboard under the "User API Tokens" section.
  2. Assign it "Workers AI: Read" permissions so the automation platform can view the available model catalog.

The API will reject the handshake without this specific scope. It prevents the workflow from initializing.

Step 2: use the HTTP request integration for direct API calls

By bypassing third-party wrapper limitations, direct interaction with the Clef API results in the lowest possible overhead. Cloudflare hosts both Clef and Clef-flash on the Workers AI global network, according to the Cloudflare release notes.

A standard POST request to the Cloudflare API endpoint makes them reachable. This method gives you granular control over the tools array in the JSON payload. That’s necessary for the model to return a valid function call rather than a conversational refusal.

Step 3: configure the AI integration with a custom base URL

The decision model can be treated as a custom OpenAI-compatible provider within your orchestration layer for teams that prefer a standardized interface.

By pointing the base URL to the Cloudflare gateway and selecting Clef as the model ID, you can swap out other models for routing logic without rewriting your prompt templates.

Retries, memory, and tool-calls constitute the core logic of these agents, and in Activepieces, that engine sits in the public, MIT-licensed core.

Every decision Clef makes is visible in the run trace, allowing teams to verify the model's logic against the code in the public monorepo. MoneyGram and FundingSocieties run this in production to maintain this level of transparency across their automated workflows.

A unified data schema can be maintained across your entire automation library with this configuration.

Specialized model impact on automation performance

Cloudflare Clef signals a fundamental shift toward small, specialized decision models that optimize agentic workflows by stripping away the overhead of general-purpose conversation.

When an enterprise replaces a multi-modal flagship with a dedicated tool-calling model, it eliminates the "reasoning tax" paid for linguistic nuance that internal automation triggers never actually see.

Available models for structured JSON output

Deploying complex, multi-step chains where the primary requirement is precise JSON output rather than creative prose is now possible for operations teams. The current landscape of available models reflects this move toward granular efficiency.

A rectangular logic gate component with two input wires coming from the left and one single output wire exiting to the…

For high-throughput workloads, Gemini 3.5 Flash-Lite functions as the high-velocity baseline. Simple data routing tasks don't consume the budget allocated for complex logic.

Claude Haiku 4.5 has a specialized balance of speed and near-frontier intelligence for rapid decision-making in workflows that require more than basic pattern matching but less than full reasoning.

GPT-6 Luna pricing for enterprise scaling

For high-volume enterprise tasks, GPT-6 Luna is the cost-sensitive tier.

It prevents the billing spikes associated with scaling agentic triggers across thousands of concurrent users, ensuring that rapid growth does not result in unpredictable financial overhead, so businesses can expand their user base with confidence in their budget stability.

Ministral 3 3B acts as a tiny, efficient text engine that can be embedded closer to the data source to reduce the latency inherent in sending small packets to massive, distant clusters.

Matching demand to specialized assets

Finance operations can move away from the "one-size-fits-all" model strategy that often leads to over-provisioning by utilizing these specialized assets.

Instead of routing a simple database query through a model designed for life sciences reasoning, teams can match the specific computational demand of a task to a model designed solely to execute it.

Every dollar spent on tokens contributes directly to a functional outcome rather than subsidizing unused general intelligence under this alignment.

What Activepieces does about this

Activepieces provides the execution environment where Clef’s decisions become actual operations across its supported business applications. While the model generates the JSON tool call, our platform provides the pre-built connectors that translate that JSON into authenticated API requests.

Every dollar spent on tokens contributes directly to a functional outcome rather than subsidizing unused general intelligence under this alignment.

Because the core engine is open-source under an MIT license, organizations like MoneyGram and FundingSocieties use it to ensure that the logic governing these automated decisions remains transparent and auditable within their own infrastructure.

To solve the problem of manual schema mapping, we expose every integration as a structured tool through a built-in Model Context Protocol (MCP) server.

This allows Clef to instantly "see" the required parameters for actions in tools like Salesforce, Jira, or SAP without a developer having to write custom wrappers for each endpoint.

By providing these validated schemas, we prevent the model from hallucinating non-existent fields, which is the primary cause of failure in autonomous agentic loops.

The platform also acts as the state manager for Clef’s decision-making process. As the model moves through a workflow, Activepieces maintains the context of previous steps and handles the error-catching logic if an API returns an unexpected result.

n8n vs Activepieces

This creates a closed-loop system where Clef can focus on selecting the right tool while our engine handles the heavy lifting of connectivity, authentication, and data persistence across the global Workers AI network.

Frequently asked questions about Cloudflare Clef

Is Cloudflare Clef free to use?

Costs are tied to token usage under the usage-based billing structure of Cloudflare Clef, which charges $0.24 per million input tokens. This model ensures that an organization's expenses scale directly with the amount of input it sends to the model, rather than a flat subscription fee.

It avoids paying for the idle time or conversational filler typical of general-purpose assistants. For a finance manager, this shift means the API bill becomes a direct reflection of operational throughput. It allows for precise unit-cost modeling of automated workflows.

Can Clef replace GPT-6 Astra for complex reasoning?

Rather than a replacement for high-reasoning frontier models like Gemini 3.8 Flash or GPT-6 Astra, Clef is designed as a specialized decision model for execution.

While a flagship model manages the long-horizon logic and nuanced user interaction, Clef acts as the specialized "last mile" executor that translates those high-level intents into precise API calls.

This prevents high-cost reasoning engines from being wasted on repetitive syntax formatting.

Does Clef support multi-modal inputs?

Clef processes state as text, JSON, images, or video to support high execution speeds and low latency.

If a workflow requires analyzing visual data, such as verifying a receipt or scanning a thermal image, an engineer must first route that file through a vision-capable model to extract the relevant text.

Visual elements of a complex invoice can be ingested and described by Gemini 3.8 Flash.

Cloudflare Clef then takes the resulting text description and maps it to the specific fields required by an ERP system's API. Workers KV, a global data store, caches the final transaction state so the audit log remains consistent across regions.

References

Share

Get started

Automate this without code.

Cloud or your own servers.

Start free Talk to sales