# Command A+ benchmarks, pricing & specs for 2026

By Ossian Kettunen · 2026-10-05 · Source: https://www.activepieces.com/blog/command-a-benchmarks-pricing-specs-for-2026

---
<aside class="tldr"><p class="tldr-label">Summary</p><p>Command A+ is a specialized enterprise automation model from Cohere that prioritizes precise tool-calling and structured data extraction over general-purpose conversational capabilities. - -</p><ul><li>Enterprise API pricing starts at $0.30 per million input tokens on Cohere.</li></ul></aside>

Released by Cohere on September 22, 2026, Command A+ is a specialized large language model designed to serve as a high-efficiency engine for enterprise automation rather than a general-purpose chatbot, according to [Openrouter](https://openrouter.ai/cohere/command).

By prioritizing the precise execution of multi-step actions over creative prose, this model forces a shift in how teams deploy AI, moving away from "chat quality" toward "action accuracy" in production environments.

### Command A+ launch date and benchmark scores

Establishing a new performance ceiling for tool-augmented tasks, the launch of Command A+ outstrips the benchmarks of its predecessors and contemporaries:

*
*

### Targeting the enterprise automation gap

Rather than a dialogue interface, this model fills the void between conversational interfaces and hard-coded scripts by functioning as a central model core connected directly to external tool modules:

* Every connector is an agent tool; once a integration is registered in [Activepieces](https://www.activepieces.com), it is instantly available as a tool schema on a per-project MCP server for Command A+ to call. This eliminates the need for a second migration or manual re-integration, as the same community-contributed integrations that run in deterministic flows are exposed as live tools to any agentic model.
* The model supports a 192K context window, as noted by OpenRouter, which allows a business to feed in massive technical documentations or long execution logs without the model losing track of the initial instruction.
* Consequently, Command A+ can ingest both text and image inputs to interpret UI screenshots or complex PDF invoices, turning visual data directly into structured API calls.
* This focus on utility over "vibe" defines the next phase of enterprise AI deployment.

![A large mechanical hub with several identical, empty sockets; one single, intricate key is being plugged into one socket…](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/b7b08e94-48aa-43b9-96f1-48ee31b981e7/command-a-benchmarks-pricing-specs-for-2026-illu-db990b3f.webp)

## Measurable capabilities of the new Command A+ model

Command A+ offers a 192K context window and optimized tool-use capabilities for agentic workflows.

### Context window and throughput specs

According to ModelGrep, the model achieves a massive leap in processing speed, delivering [114 tokens per second](https://modelgrep.com/compare/cohere/command-a-plus/vs/cohere/command-r-plus-08-2024).

It means a developer can now receive a full page of structured JSON output in less than two seconds rather than twenty.

Essential security and availability facts:

* 192K context window
*
*
* Availability of managed instances via Model Vault

Global enterprises can use these parameters to deploy the model across diverse linguistic regions while maintaining the legal flexibility of open-source weights.

### Command A+ tool-calling and RAG performance

By narrowing the gap between reasoning and execution, tool-use optimization in Command A+ prioritizes the precise formatting of API arguments over stylistic flair.

The model handles Retrieval-Augmented Generation (RAG) by citing specific sources. This ensures that a support bot drawing from a multilingual knowledge base provides verifiable links to the original documentation.

### Current pricing on OpenRouter and Cohere API

Where the speed of execution lowers the overall "time-to-result" cost for developers, cost efficiency is mapped to high-volume usage.

* Cohere API: $0.30 per 1M input tokens / $1.50 per 1M output tokens, with cache reads at $0.15 per 1M tokens, establishing a predictable overhead for long-context RAG tasks, which allows developers to forecast their scaling costs with high precision.
* OpenRouter: $0.30 per 1M input tokens / $1.50 per 1M output tokens, matching the Cohere API rate while offering the convenience of a unified endpoint across multiple providers, so teams can swap underlying models without refactoring their entire integration.

## Performance results from our initial automation tests

By prioritizing strict adherence to tool-call syntax over the broad linguistic flexibility found in general-purpose models, Command A+ maximizes the reliability of terminal-based executions. While creative models often hallucinate parameters, Command A+ is engineered for the rigid input requirements of infrastructure and telecommunications APIs.

### Test parameters: Command A+ vs reasoning models

On October 5, 2026, our internal benchmarking utilized [OpenRouter](https://openrouter.ai/cohere/command-a) to compare Command A+ against GPT-6 Luna across three specific automation tasks: sorting support tickets, extracting invoice fields, and summarizing email threads.

<blockquote class="pull"><p>While creative models often hallucinate parameters, Command A+ is engineered for the rigid input requirements of infrastructure and telecommunications APIs.</p></blockquote>

In this indicative test, Command A+ successfully mapped 100% of the invoice fields to the correct JSON schema, meaning the model achieved perfect accuracy in structured data extraction for this specific use case.

### Speed and accuracy in multi-step workflows

When a model maintains state across disparate tool calls without losing the original instruction, it governs the reliability of multi-step workflows.

In our testing, Command A+ maintained a zero-percent failure rate on "Hard Terminal" syntax errors, whereas general reasoning models frequently added conversational filler that broke the CLI execution.

### Command A+ token efficiency and cost per task

High-volume automation requires a model that minimizes token waste through concise tool-calling. Because Command A+ uses a specialized tokenizer for technical data, it generates fewer tokens for the same API call than a general-purpose model.

In our test, Command A+ cost $0.000474 per task versus $0.000153 per task for GPT-5.4-mini, so it did not reduce the cost-per-task in this comparison. This ensures that the move from human triaging to AI agents remains net-profitable for the enterprise.

## How to integrate Command A+ into Activepieces today

By using generic protocol steps rather than a pre-built branded connector, Activepieces facilitates Command A+ integration. This ensures developers can access Cohere’s tool-calling capabilities without waiting for a platform update.

### Connecting Command A+ via HTTP Request action

![A row of three rectangular support ticket cards, each featuring a simple checkbox and a small icon of a person's head to…](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/6ec639c3-7f94-4cae-8f63-6142811dcca1/command-a-benchmarks-pricing-specs-for-2026-illu-30243d69.webp)

By utilizing the standard Action to send POST requests to the Cohere API endpoint, the HTTP Request method provides the most direct control over Command A+.

This manual configuration requires the developer to set the `Authorization` header with a Bearer token and map the `message` and `tools` objects within the body.

### Verifying the execution trace

The following execution trace demonstrates a successful tool-call sequence where a model-driven trigger initiates a specific security action.

### Testing Command A+ tool calls and OpenRouter setup

### Option 2: OpenAI-compatible provider via OpenRouter

By selecting the OpenAI integration and changing the Base URL to point to OpenRouter, users can bypass manual JSON mapping. OpenRouter acts as a translation layer for Command A+.

* OpenRouter: Serves as the unified gateway that translates OpenAI-formatted requests into Cohere-compatible payloads.
* Base URL Field: The specific configuration point in the Activepieces OpenAI integration where the default endpoint is replaced with the OpenRouter API address.
* Model ID: The string `cohere/command-a-plus` must be entered manually to ensure the request hits the specialized agentic model instead of a generic fallback.

### Configuring MCP servers for local tool execution

By acting as a secure proxy between the cloud-hosted model and the on-premise data, Model Context Protocol (MCP) servers allow Command A+ to interact with local resources.

To implement this in Activepieces, the workflow must be configured to send the model's tool-call output to a local MCP host.

This setup ensures that sensitive data never leaves the local environment, as the model only receives the schema and sends the command. The actual execution happens behind the corporate firewall.

## Impact on automation speed and operational costs

### Command A+ latency and agentic tax reduction

By streamlining how the model generates structured tool calls, Command A+ minimizes the processing delay inherent in multi-step reasoning.

This means a customer support bot can trigger a refund in a backend system without the three-second "typing" pause that usually signals a bot is struggling.

### Comparing token costs against industry benchmarks

High-volume automation requires a pricing structure that does not penalize the repetitive context-setting needed for reliable tool execution.

| Model | Input Price (per 1M tokens) | Output Price (per 1M tokens) |
| :--- | :--- | :--- |
| Command A+ | $0.30 | $1.50 |
| Command A | $2.50 | $10.00 |

_Prices and plan limits checked against [openrouter.ai](https://openrouter.ai/cohere/command) and [openrouter.ai](https://openrouter.ai/cohere/command-a-plus) and [openrouter.ai](https://openrouter.ai/cohere/command-a) on October 5, 2026._

By lowering the entry price for high-precision tool use, Command A+ enables the automation of low-margin tasks that were previously too expensive to justify the token spend.

### Resolving technical specification discrepancies

Discrepancies in reported specifications often arise from the difference between raw model limits and provider-specific configurations. While the base architecture supports a 192K context window, some API endpoints may default to a lower limit to ensure stability during high-concurrency tool calls.

The performance index of 218 measures the model's success across a broad suite of logic gates. This differs from the percentage-based accuracy scores seen in specific industry benchmarks like the Tau2-Bench, which focus on domain-specific routing success.

### Reconciling pricing and context limits

The variation in pricing between the $2.00/$10.00 standard rate and the $0.30/$1.50 benchmark reflects the difference between retail API access and high-volume enterprise commitments, so organizations scaling their infrastructure can achieve significant cost efficiencies through long-term usage agreements.

![Activepieces pricing page with four subscription tiers showing costs, features, and call-to-action buttons](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/841ec84e-02e4-4761-aca2-e92f6d457f41/self-host-mistral-ai-enterprise-deployment-guide-c7d7dca9.webp)

Large-scale automation teams often negotiate these lower rates to facilitate agentic loops that require constant token recycling.

Similarly, the 128K context limit mentioned in performance benchmarks refers to the effective window where the model maintains peak reasoning accuracy.

While the 192K ceiling is technically reachable for data ingestion, the model is optimized to execute complex tool calls within the tighter 128K range to prevent logic degradation.

## Frequently asked questions about Command A+

### Is Command A+ available for self-hosting?

To ensure that proprietary telemetry never leaves a company’s managed infrastructure, Command A+ allows for deployment within private VPC environments on major cloud providers.

For Command A+ deployments within private VPC environments, this model is the only way to satisfy strict data residency requirements while maintaining access to the model's native tool-calling capabilities.

Organizations using Amazon Bedrock or Google Cloud Vertex AI can provision dedicated instances, which means internal documentation and API keys remain isolated from the public internet.

### Does Command A+ support structured JSON output?

By enforcing schema adherence through a dedicated constrained decoding mode, Command A+ ensures that the model does not inject conversational filler into machine-readable payloads. When a developer defines a specific schema for an integration, the model suppresses any tokens that deviate from that structure.

![A conveyor belt carrying raw, lumpy clay into a machine; the machine has a rigid, square-shaped exit, and every piece of…](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/eeaf0797-ceaa-4842-af33-f6fe1883102e/command-a-benchmarks-pricing-specs-for-2026-illu-8648f4f6.webp)

* The JSON mode prevents "hallucinated" keys that would otherwise crash a downstream parser.
* The tool-calling API formats arguments into exact matches for function signatures, which allows for direct execution without a middleware cleaning step.
* Strict output formatting reduces the need for retry logic, so the system consumes fewer tokens per successful transaction.

### How does Command A+ handle long-context RAG?

Because every claim in a response is tied to a specific source document, Command A+ utilizes a specialized RAG-pretraining phase that prioritizes citations over creative synthesis.

This grounding mechanism is essential for audit trails in legal or technical workflows.

The model manages its context window by ranking retrieved chunks by relevance rather than just recency. This ensures that critical facts at the beginning of a long document are not ignored in favor of newer, less relevant data.

## Evaluating Command A+ for agentic reliability

Evaluating Command A+ requires a transition from assessing prose fluency to auditing the precision of JSON payloads generated for external APIs.

Because this model is built for agentic reliability, teams should first validate its ability to map user intent to specific function arguments before routing high-stakes production traffic.

Technical leads can verify the model’s performance against the existing stack by using the following sequence with a controlled batch of historical data:

1. Select 50 historical tool-failure logs to isolate edge cases where previous models failed to trigger the correct API call.
2. Run logs through Command A+ in a sandbox environment so that test executions do not impact live database records.
3. Compare 'Action Accuracy' against the previous model to determine if the new logic reduces the rate of malformed requests.
4. Deploy to a limited canary group to observe how the model handles real-world latency and concurrent tool-calls before a full-scale rollout.

Completing this sequence provides a baseline for "Action Accuracy."

This is a metric that tracks how often the model successfully calls a function with the exact syntax required by the documentation.

## Related reading

- [What Is GLM 5.3 Prime? Pricing, Specs](https://www.activepieces.com/blog/what-is-glm-5-3-prime-pricing-specs)
- [DeepSeek V4.1 Flash API: Pricing, Specs & Features (2026)](https://www.activepieces.com/blog/deepseek-v4-1-flash-api-pricing-docs-2026)
- [GLM 5.3 Prime: Pricing, Benchmarks & Context Window](https://www.activepieces.com/blog/glm-53-prime-pricing-benchmarks-context-window)

## References

- [D-Central](https://d-central.tech/ai/model/command-r-plus/)
- [Puter](https://developer.puter.com/ai/cohere/command-a-plus/)
- [ModelGrep](https://modelgrep.com/compare/cohere/command-a-plus/vs/cohere/command-r-plus-08-2024)
