# Step 5 Preview: How the New StepFun Model Impacts AI Automation

By ColetteDuprez · 2026-10-08 · Source: https://www.activepieces.com/blog/step-5-preview-how-the-new-stepfun-model-impacts-ai-automation

---
<aside class="tldr"><p class="tldr-label">Summary</p><p>Step 5 Preview enhances AI automation by consolidating multimodal reasoning and text processing into a single API call, enabling complex agentic workflows with a one-million-token context window.</p><ul><li>The model supports 1,000,000 tokens of context for processing exhaustive technical documentation. -</li><li>Invoice extraction tasks are completed in 1.2 to 3.9 seconds per document.</li></ul></aside>

Step 5 Preview is [StepFun’s](https://platform.stepfun.ai/) flagship model for agentic work, designed to handle complex multimodal reasoning within production automation pipelines.

This model functions as a high-reasoning engine capable of processing diverse data types, which allows teams using [Activepieces](https://www.activepieces.com) to consolidate vision and text tasks into a single API call rather than chaining multiple specialized services.

![A data definition form showing fields for extracting invoice issuer information with name, description, and data type…](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/2d0203e2-1833-4861-b048-548984d0a5a1/air-gapped-ai-deployment-how-to-run-mistral-2026-d52ce763.webp)

A tiered access structure governs the model’s utility in a live workflow, dictating how aggressively an organization can scale its automated operations. According to [APIs.io](https://apis.io/rate-limits/stepfun/stepfun-rate-limits/), the rate limits are defined by the account's financial commitment:

## Understand the Step 5 Preview model

*
*

This capacity ensures that a developer is not bottlenecked by artificial latency during high-volume batch processing.

While competitors like Claude Fable 5.1 or GPT-6 Astra offer similar multimodal capabilities, Step 5 Preview positions itself as a specialized alternative for teams requiring specific agentic performance. These performance characteristics suggest a model ready for integration into existing enterprise architectures today.

![Activepieces homepage showing flexible AI workflow automation for technical teams with hero graphic and use case cards.](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/ba8592d6-bc9a-4ace-83e4-839c44f8d423/comparing-clickup-alternatives-for-2026-screensh-deedd788.webp)

## Technical capabilities of the Step 5 Preview release

By processing vision and text inputs within a **every massive context window**, Step 5 Preview achieves industrial-grade utility that accommodates exhaustive technical documentation and high-resolution imagery.

This capacity ensures that complex automation scripts do not lose track of early instructions during long-running execution cycles, preventing the logic drift that typically occurs in smaller models.

By supporting native multimodal inputs, the model eliminates the need for external optical character recognition tools to interpret UI screenshots or PDF schematics, reducing the total number of failure points in a production pipeline.

![A workflow with three steps: Chat UI for human input, Extract Structured Data using Utility AI, and a third step below.](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/5984c238-cffe-4889-9d89-183af7c95a38/model-security-vs-data-security-in-ai-workflows-97b21eb2.webp)

The following table demonstrates how this capacity aligns with current frontier models, allowing teams to determine if their specific dataset fits within the model's operational memory.

| Model | Context Window Capacity |
| :--- | :--- |
| Gemini 3.1 Pro | 2,000,000 tokens |
| GPT-6.1 Sol | 1,050,000 tokens |
| Step 5 Preview | 1,000,000 tokens |
| Claude Sonnet 5.5 | 200,000 tokens |

_Prices and plan limits checked against [platform.stepfun.ai](https://platform.stepfun.ai/docs/en/guides/models/step-5-preview) and [stepfun.com](https://stepfun.com/step-5-preview) and [openrouter.ai](https://openrouter.ai/stepfun/step-5-preview) on October 8, 2026._

For a controller, these figures represent a threshold for auditability; a million-token window allows the model to reference the entire history of a transaction alongside the governing regulatory framework in a single pass.

![A single, extraordinarily long scroll of paper that a person is holding at chest height, with the bottom of the scroll…](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/26fbe5e1-ea3c-4a17-95c5-1f0d79edfa5f/step-5-preview-how-the-new-stepfun-model-impacts-dff76a35.webp)

This architectural choice directly mitigates the risk of "hallucination" caused by truncated data.

Beyond simple text ingestion, the model’s vision capabilities allow it to verify visual state changes in a browser or legacy application, providing a verifiable trail of successful execution that purely text-based models cannot offer.

## Performance benchmarks from our internal automation tests

Step 5 Preview delivers reliable accuracy for high-volume back-office tasks while maintaining a latency profile that supports synchronous process execution.

By utilizing a sparse Mixture-of-Experts architecture (a design that activates only specific neural pathways for each request) the model from [StepFun](https://stepfun.com/step-5-preview) reduces the computational overhead typically associated with flagship-class reasoning.

### Task 1: Structured data extraction from invoices

Preventing downstream reconciliation errors in the general ledger requires high field accuracy in our audit of invoice processing.

During internal testing on [OpenRouter](https://openrouter.ai/stepfun/step-5-preview), Step 5 Preview completed invoice extraction in 3.9 seconds, ensuring that automated accounts payable workflows can validate vendor data without triggering system timeouts.

This allows a controller to refresh a batch processing queue and see updated records without waiting for an overnight sync.

### Task 2: Multimodal reasoning for inventory checks

Verifying physical stock levels against digital records requires a model that can interpret visual discrepancies in warehouse imagery.

Step 5 Preview identifies mismatched SKU counts between a photograph and a manifest, providing the visual evidence necessary to close an audit trail.

Because the model is built for agentic work, it can navigate a browser to log these discrepancies directly into an ERP system, removing the need for manual data entry between disparate software platforms.

![A person standing between two high walls, reaching out to hold two handles—one on each wall—effectively acting as a human…](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/820d21e1-5de7-4fd7-9ede-b59409ba2231/step-5-preview-how-the-new-stepfun-model-impacts-16078523.webp)

### Task 3: Latency and cost per thousand tokens

GPT-5.4-mini was the faster option for simple string manipulations in earlier tests, while Step 5 Preview offers the stability required for complex, multi-step summaries. The following data highlights the performance trade-offs between these two models across three common automation categories.

| Task | Step 5 Preview Latency | GPT-5.4-mini Latency |
| :--- | :--- | :--- |
| Ticket sorting | 0.9 – 1.7s | 0.9s (Median) |
| Invoice extraction | 1.2 – 3.9s | 0.9s (Median) |
| Email summary | 0.8 – 17.7s | 0.9s (Median) |

A standard workflow can process approximately 15 documents per minute based on the median latency of 3.9 seconds for invoice extraction.

While the email summary task peaked at 17.7 seconds, this represents the model performing deep synthesis of long threads; for a manager, this means receiving a comprehensive briefing rather than a shallow, bulleted list that misses key context.

![A chef meticulously stirring a single, small pot on a large stove, while in the background, a conveyor belt rapidly moves…](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/c19f6b81-4d7e-4937-aad7-6e05e2912b11/step-5-preview-how-the-new-stepfun-model-impacts-7b71dc07.webp)

These results confirm that Step 5 Preview is a viable high-performance alternative for teams that prioritize reasoning depth over raw millisecond speed.

## Connecting Step 5 Preview to your current workflows

Integration with Step 5 Preview requires no specialized middleware because the model utilizes an OpenAI-compatible API structure. This compatibility ensures that automation architects can swap existing model calls for Step 5 Preview without refactoring the core logic of their sequences.

Activepieces provides the bridge for this transition by making every connected integration available as an MCP tool.

### Option 1: The HTTP Request method

![A workflow automation flow with four steps: MCP Tool, Get all Events from Google Calendar, Find Database Item in Notion…](https://ap-marketing-media.fra1.cdn.digitaloceanspaces.com/uploads/318b16a0-2f90-469e-8680-14a0f0808df4/what-is-an-agent-harness-architecture-and-terms-6191a2be.webp)

Direct API communication serves as the most stable integration path because it bypasses the need for third-party connector updates.

By using a standard web request utility, teams maintain full control over version headers and timeout settings, which prevents workflow failures when model providers update their documentation.

### Establish the API connection

To route requests correctly, the automation must point to the official StepFun gateway.

1. Generate API Key in StepFun [Platform](https://platform.stepfun.ai/docs/en/guides/models/step-5-preview).
2. Open the request configuration in your automation builder.
3. Set Method to POST and URL to the StepFun endpoint.
4. Add Authorization header with Bearer token.
5. Map the JSON payload to include the model name and prompt variables.

This manual configuration creates a permanent audit trail of the exact parameters sent to the model.

### Option 2: Using the AI Step with custom providers

Utilizing a dedicated AI module allows for cleaner visual mapping by leveraging the tool's native handling of chat history and structured outputs.

Most modern automation platforms include a "Custom" or "OpenAI Compatible" provider setting, which enables the use of Step 5 Preview within the standard interface rather than raw code blocks.

This Bring-Your-Own-Key availability, documented across all tiers including the free plan, prevents a platform from reselling the model and deciding your AI strategy for you.

### Connecting Step 5 Preview via MCP servers

For complex agentic workflows, connecting through a Model Context Protocol (MCP) server provides a standardized bridge between the model and local data sources.

An MCP server acts as a translator that allows Step 5 Preview to securely query internal databases or file systems without exposing those credentials to the public internet.

### Setting up your MCP server for Step 5

In practical terms, an MCP server is a lightweight piece of software, often a Node.js or Python script, that runs on your local machine or within a private cloud environment.

You can download pre-built servers for common tools like PostgreSQL or Google Drive from the official MCP repository or run them as Docker containers.

To connect this to your workflow, you simply provide the server's local address or transport command to your automation host.

Once the host recognizes the server, Step 5 Preview can automatically discover and execute the specific functions the server exposes, such as reading a local CSV or updating a private SQL table.

This architecture ensures that sensitive operational data remains behind the corporate firewall while still benefiting from the model's reasoning capabilities.

## Pricing and availability for the preview phase

Step 5 Preview is available through the StepFun platform by creating an API key and sending requests with cURL or Python.

This vetting process ensures that high-volume multimodal requests do not destabilize the infrastructure during the initial rollout, meaning teams must clear a compliance review before moving beyond sandbox testing.

Access is managed through a consumption-based credit system, where users purchase units upfront to draw down against their API calls.

The pricing structure for this preview period distinguishes between input and output tokens, with separate billing categories for text and image processing.

Because the model supports high-resolution visual inputs, the platform calculates costs based on the complexity and pixel count of the uploaded media rather than a flat fee per file.

This granular accounting allows controllers to audit the exact resource consumption of specific automation pipelines, preventing the "hidden cost" spikes common in unmonitored vision workflows.

Availability is currently restricted to specific geographic regions to comply with data residency requirements.

* StepFun Developer Console: The primary gateway for managing API credentials and monitoring real-time usage metrics.
* Token Usage Dashboard: A reporting tool that provides a line-item breakdown of costs per deployment environment, so managers can identify inefficient prompts that are inflating the monthly spend.
* Technical [Support](https://support.claude.com/en/articles/17154008-monthly-api-credits-for-max-and-team-plans) Portal: A dedicated channel for reporting latency issues or API errors during the preview phase, which limits downtime for teams integrating the model into production-adjacent systems.

## What to watch next for AI automation teams

Stable release certification and the removal of the preview designation are the primary indicators that these tools have met the rigorous reliability standards required for production environments.

This transition signals that the underlying infrastructure can support high-volume API calls without the risk of sudden schema changes, which means teams can move from experimental sandboxes to customer-facing deployments.

Beyond stability, the expansion of regional availability to additional cloud data centers is the necessary step for global compliance.

Until a vendor adds support for a specific jurisdiction, organizations subject to strict data sovereignty laws are effectively barred from routing sensitive logs through these newer endpoints.

Automation leads should monitor the integration of specific frontier models into these multimodal workflows to ensure the chosen engine matches the complexity of the task.

* Claude Fable 5.1 from [Anthropic](https://www.anthropic.com/news/claude-3-5-sonnet), for projects requiring demanding reasoning and long-horizon agentic work.
* GPT-6 Luna from OpenAI, for teams scaling cost-sensitive, high-volume workloads that require consistent logic at a lower price point.
* Gemini 3.8 Flash from Google, for long-horizon software engineering tasks where deep context is required to maintain code integrity.
* Mistral Large 4, for organizations prioritizing flagship, state-of-the-art open-weight multimodal models to maintain vendor flexibility.

The final milestone to watch is the introduction of specialized cybersecurity safeguards, such as those found in GPT-5.6 Cyber. The availability of these models within standard automation protocols will allow security teams to automate authorized vulnerability research without manual oversight.

## What Activepieces does about this

Activepieces provides the bridge for this transition by making every connected integration available as an MCP tool.

This unified catalog means an agent can call any of the available integrations immediately without a separate export step or manual re-wiring.

The platform runs the model you choose on your own provider key, ensuring that the high-volume multimodal spend of Step 5 Preview lands on your own account at your negotiated rates.

This Bring-Your-Own-Key availability, documented across all tiers including the free plan, prevents a platform from reselling the model and deciding your AI strategy for you.

By maintaining this direct relationship with the model provider, teams can scale their agentic workflows without the hidden markups often found in managed AI services.

For teams managing high-volume production pipelines, Activepieces offers a self-hosted option that keeps the entire automation engine within a private cloud.

This deployment model is critical for organizations using Step 5 Preview to process sensitive multimodal data, as it ensures that the orchestration logic and API keys never leave the corporate perimeter.

The ability to run the software on-premise allows for strict adherence to data residency requirements while still leveraging the high-reasoning capabilities of the StepFun flagship model.

This architectural approach allows a controller to maintain a verifiable trail of successful execution. Because the automation logic is decoupled from the model's reasoning, the system can log every visual state change and API call to an internal database for auditability.

This ensures that even as models like Step 5 Preview evolve, the underlying business processes remain stable, secure, and fully under the organization's control.

## Frequently asked questions

### Is Step 5 Preview available via an OpenAI-compatible API?
Engineering teams can swap existing model endpoints for this new tier without rewriting the underlying transport logic because Step 5 Preview utilizes a standard RESTful architecture.

This protocol compatibility aligns with the header and payload requirements of the OpenAI chat completions schema.

By maintaining this interface, the model integrates directly into established monitoring stacks like the Datadog observability platform, ensuring that every request generates a traceable audit log for compliance review.

### Does Step 5 Preview support image and video inputs?
Working as a native multimodal engine, the preview tier is capable of processing both static frames and temporal data streams alongside text instructions.

This capability allows the model to interpret visual state changes in a software interface, meaning it can validate that a UI element has rendered correctly before proceeding to the next step of an automated test.

Unlike text-only models, this version can ingest direct exports from the Veo 3.1 cinematic video generation tool to verify visual consistency across synthetic media workflows.

### What are the current rate limits for the preview tier?
A sliding window of requests per minute and tokens per day governs throughput for the preview tier to ensure equitable resource distribution during the initial release phase.

These constraints require developers to implement robust exponential backoff logic in their code to prevent service interruptions during high-concurrency events.

Because these limits are enforced at the organization level, teams must partition their API keys between development and production environments to prevent a single testing suite from exhausting the entire daily quota.

## Related reading

- [How to choose a Mistral model size for automation in 2026](https://www.activepieces.com/blog/choosing-a-mistral-model-size-for-self-host-automation)
- [Choosing the Best Model for Business Workflows](https://www.activepieces.com/blog/deepseek-r1-vs-v3-for-business-ai-2026-guide)
- [Xing4.0-29B-A4B: Agentic MoE Model Guide (2026)](https://www.activepieces.com/blog/xing40-29b-a4b-agentic-moe-model-guide-2026)

## References

- [StepFun](https://openrouter.ai/stepfun/step-5-preview)
- [APIs.io](https://apis.io/rate-limits/stepfun/stepfun-rate-limits/)
