AI DevelopmentNew Release10 min readPublished October 8, 2026

A hosted trial and a future deployment choice

Step 5 Preview: Access, Pricing and Open-Weights Plans

Evaluate Step 5 Preview for agent work: hosted access, dated OpenRouter prices, context limits and the distinction between API use and planned weights.

DA
Digital Applied Team
Research and practical guidance
PublishedOctober 8, 2026
Read time10 min
SourcesPrimary documentation

Step 5 Preview is a model you can evaluate through hosted access now, while its downloadable weights remain a separate promised milestone. That distinction matters more than a leaderboard position if you are deciding what your team can deploy this week. Build the trial around the endpoint that exists, and keep future self-hosting out of the current acceptance decision.

Product documentation and current access details were checked on October 11, 2026. The workflow examples are proposed evaluation designs, not reports of completed customer deployments.

Key takeaways
  1. 01
    Use an exact routeRecord the provider and model identifier instead of relying on a display name.
  2. 02
    Date the price observationOctober 11 catalogue values are not evidence of earlier or permanent terms.
  3. 03
    Keep weights in the future tenseStepFun says October 15; confirm the files and license when that date arrives.
  4. 04
    Evaluate the assembled agentCheck source selection, tool arguments and accepted outcomes, not just fluent answers.

01 — Release boundarySeparate the model from the access milestones

StepFun’s model page describes Step 5 Preview as a mixture-of-experts model with 600 billion total parameters, 27 billion active per token, a one-million-token context window and vision input. The page says hosted products and API access are available, with open weights planned for October 15. Those are vendor statements; they do not establish a successful deployment in your environment.

The OpenRouter route appeared in the catalogue with an October 8 creation timestamp. A catalogue timestamp is not proof of the model’s first public launch, and this guide does not treat it as one. The distinction prevents a listing event from becoming a misleading release story. For a buyer, the useful facts are where access works, what the endpoint accepts and what obligations come with it.

Keep three milestones in your evaluation record: announced capability, usable hosted access and downloadable artifacts. A team can complete a useful hosted trial without the third milestone. A team that requires its own infrastructure cannot complete its deployment review merely because the first two are satisfied. This separates a promising option from a dependency that might delay an already committed project.

Available to assess
Hosted endpoint
Current experiment

Test behavior and integration through the actual API route available to your team.

Present choice
Promised milestone
Downloadable weights
Future review

Check the published files, license, hardware guidance and serving behavior after release.

Separate decision

02 — Route identityRecord the endpoint you are really buying

The exact aggregator identifier is stepfun/step-5-preview. In the October 11 OpenRouter snapshot, it advertises text, image and video input with text output, a 1,000,000-token context length and a 64,000-token maximum completion value. These are route metadata, not a promise that every direct provider interface exposes identical controls or limits.

Before a pilot, send a small representative request through your intended account and inspect the actual result. Confirm the accepted input type, response structure and usage reporting. A multimodal product page can describe the model family while a particular endpoint supports a narrower subset. If your workflow depends on video, that request needs its own integration check rather than an assumption based on vision support.

Save the route configuration alongside the evaluation input. If routing or provider selection changes later, you need to know whether the behavior comparison used the same serving path. Our frontier-model selection guide covers the larger choice; the procurement detail here is that two requests with the same friendly model name can still differ in access, limits and billing. Record enough information to repeat the request meaningfully.

One-million-token context is capacity

A large request limit does not prove that the model will find the right exception in a long document set. Test evidence selection and answer quality separately from whether the API accepts the payload.

03 — Cost observationUse dated prices and explicit arithmetic

The OpenRouter catalogue checked on October 11 lists $1 per million input tokens, $2.70 per million output tokens and $0.05 per million cached-input tokens for this route. These are a dated aggregator observation. They should not be presented as a direct StepFun quote, an October 8 historical price or a guarantee that every request qualifies for cached billing.

For a hypothetical request using 30,000 uncached input tokens and 3,000 output tokens, those rates produce a token charge of $0.0381: $0.03 for input and $0.0081 for output. This illustration excludes tools, retries and review work. It is not an observed production bill. If the request needs a second attempt, the cost of producing an accepted result increases even when the advertised token price stays unchanged.

Long context makes this arithmetic more consequential. It can be tempting to attach every document because the request fits, but irrelevant material still contributes to input usage and can make evidence selection harder to judge. Our human-review cost guide explains the other half of the bill: the correction effort required after the model returns. Measure both rather than declaring the cheapest token a business saving.

Suppose the team expects a repeated policy prefix to receive cached pricing. Record the usage returned by the endpoint for the actual requests instead of subtracting an assumed discount from the forecast. The prefix may change when the policy changes, and the request shape may not qualify as expected. A procurement sheet should therefore keep an uncached baseline and a separately evidenced cached case. This makes the business estimate resilient to an ordinary document update, rather than depending on an optimization that nobody has confirmed on the chosen route.

OpenRouter route metadata checked October 11, 2026. The worked example is hypothetical and excludes all non-model costs; these are not asserted launch-day rates.
Observed ratePer million tokensImportant condition
Input$1.00Uncached input in the catalogue observation
Output$2.70Generated output tokens
Cached input$0.05Only when the endpoint reports eligible cached use

04 — Task designMake a small agent trial answer a real question

Choose a job whose outcome your team can inspect. A useful example is preparing a response from an order record and an approved policy. The agent reads the request, looks up the order, identifies the applicable rule and drafts an answer. The trial can remain read-only while still exposing important differences in reasoning, retrieval and tool handling.

Include cases where the information is incomplete. Give the agent two similar orders, a missing attachment, an outdated policy and a tool response that explicitly says the record was not found. The valuable behavior is preserving the uncertainty and taking the correct next step. A polished answer that quietly invents a missing detail should fail, even if it reads better than the current model’s response.

Keep the input set stable across candidates and record any prompt changes separately. Otherwise, the comparison mixes a model switch with a workflow redesign. Decide the acceptance criteria before reviewing results: correct record, correct policy, supported answer and no unauthorized action. These proposed tests make the release relevant to a business process without claiming that Digital Applied has already measured Step 5 against another model.

An illustrative difficult case could contain an order with two line items and a policy exception that applies to only one. Ask the agent to draft a response without changing the order. A clean pass preserves the item identities, applies the exception only where it belongs and tells the reviewer what remains unresolved. A partial pass might find the correct exception but apply it to the whole order. Recording that distinction is more useful than a single preference score, because it tells the team whether the model needs a clearer task boundary or whether the information supplied was insufficient.

  • Use ordinary tasks and known difficult cases.
  • Keep evidence, tool responses and final drafts together.
  • Score the errors that would create extra work for your team.

05 — Action evidenceInspect tool use before trusting the final message

An agent’s final answer can conceal a failed lookup or a malformed request. Review the actual tool arguments and responses, especially where the model has to select between similar records. If the tool rejects an identifier, the assistant should not turn that rejection into a confident statement about the customer. The test should preserve the failed interaction rather than scoring only the eventual prose.

Design the tool description so required information is explicit. If a change needs an order identifier and a selected date, the agent should obtain both before proposing the action. Avoid giving the model a vague tool that accepts arbitrary text and then blaming it for every ambiguous interpretation downstream. The model and the integration each need a clear contract that a reviewer can inspect.

Use the tool-error decision reference to separate a correct retry, a request for missing information and a stop. Repeated attempts at the same rejected action can consume budget without improving the outcome. A useful preview trial records those attempts as part of the result. It does not erase them because the final answer eventually sounds plausible.

Add a case where the lookup returns a valid but unexpected record. The model should compare the result with the user’s identifiers rather than accept any successful response as the intended order. This exposes an integration weakness that ordinary not-found tests miss. Keep the expected comparison fields explicit: a display name alone may not distinguish two customers, while an exact order identifier can. The test is about following the tool contract and preserving uncertainty, not asking the model to infer a person’s identity from incidental details.

Completion must be observable

For a write-enabled trial, verify the destination record. A tool call being generated, a request being accepted and the intended business change being completed are separate events.

06 — Hosting pathKeep the open-weights plan out of today’s promises

The announced October 15 milestone can be strategically interesting without being usable today. Downloadable weights may offer additional deployment choices, but the actual license, artifacts and serving instructions determine what an organization can do. Until those are available, describe self-hosting as a future evaluation rather than a current delivery option.

A future deployment review needs different evidence from the hosted trial. It must cover operational ownership, hardware capacity, updates, monitoring and the behavior of the chosen serving stack. A strong hosted result does not settle those questions. Even when weights share the same model name, repeat the business evaluation on the deployment you intend to run before assuming equivalent behavior.

Preserve the useful parts of the hosted experiment outside a vendor dashboard: test inputs, expected outcomes, scoring categories and accepted examples. That work remains valuable if the deployment route changes. It also makes a decision to stay hosted easier to defend. The goal is a dependable workflow, not satisfying an architectural preference after the model trial has already shown what the business needs.

  • Confirm the released files and license rather than relying on an announced date.
  • Assign operational ownership before choosing self-hosting.
  • Repeat task evaluation on the actual serving configuration.

07 — Evidence standardRead vendor demonstrations as leads for testing

StepFun’s page presents examples and benchmark results for software and professional work. Those can suggest useful tasks to investigate, but they remain vendor evidence. This guide does not reproduce their measurements or infer a universal advantage from them. A result obtained under a vendor’s task setup may depend on tools, prompts and evaluation rules your application does not share.

Ask what a promising example would look like in your own process. If the attraction is long-running work, test how the agent handles a dependency that remains unresolved. If it is document analysis, test conflicting versions and a relevant exception buried among similar passages. If it is coding, review whether the changes satisfy your project’s actual requirements and tests. Translate capability claims into observable acceptance conditions.

Also retain a reason not to migrate. If your current route already meets the requirement, a new model needs to justify the integration and review effort. Our AI transformation service focuses these evaluations on a defined operational outcome. The useful recommendation may be a narrow additional route, a later revisit or no change at all, depending on the evidence.

No invented winner

This article provides an evaluation design and dated access information. It does not claim a completed head-to-head quality test, a latency advantage or a guaranteed production saving.

08 — Adoption decisionTurn the result into a maintainable routing rule

Write the outcome of the pilot as an operating rule. For example, the model may prepare internal policy drafts from approved sources while external actions remain on the existing workflow. Name the exception categories that require review and the conditions that send the task back to the prior route. That is more useful than a general instruction to adopt the latest model everywhere.

A preview also needs a change boundary. Record the endpoint and evaluation date, and repeat the relevant cases when serving behavior or limits change. You do not need to rerun every experiment after an unrelated change, but a new input format, tool schema or routing path can invalidate the part of the test that depended on it. Keep the evidence connected to the configuration it actually assessed.

Step 5 Preview adds another candidate for agent work and a possible future weights-based deployment. Its value for a team will come from completed tasks and controlled failures on that team’s workflow. Keep the access facts precise, the cost arithmetic dated and the migration reversible, and the trial can produce a useful decision even if the final answer is to wait.

Set a practical trigger for the next evaluation. A changed tool schema should rerun tool-handling cases; a new document source should rerun evidence-selection cases; a different serving provider should rerun access, limits and representative behavior. Naming those triggers keeps maintenance proportional. It avoids both extremes: treating a preview trial as permanent proof or rerunning every historical experiment after an unrelated configuration edit.

  • Adopt only for the task class the trial supports.
  • Keep a fallback and the reason for using it visible.
  • Revisit the weights milestone with new evidence when the artifacts exist.
Your next step

Test the route that exists today

Record the exact endpoint and run one repeatable task comparison. Include missing information, tool failures and a check of the final destination.

Treat the October 15 weights plan as a separate follow-up decision. Hosted access can be useful now without turning a future release promise into a current commitment.

Put the method to work

Build a workflow your team can verify

Digital Applied helps teams turn a promising AI capability into a clear operating process, with useful evaluations, review points and a practical path to production.

Workflow designPractical evaluationsClear ownership
Work with us

From trial to useful work

  • →Define the task and its acceptance criteria
  • →Connect the right information and tools
  • →Review failures before expanding access
FAQ · Practical decisions

Questions before you start

The StepFun page checked October 11 says open weights are planned for October 15. This guide does not describe that future milestone as completed.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue exploring

AI Development

Mistral Large 4 Preview: Should Your Agent Team Try It?

Assess Mistral Large 4 for agent workflows: preview access, hosting choices, context limits, evaluation design and the promised open-weights release.

October 6, 2026 · 9 minRead
AI Development

Microsoft Decision 1 vs OpenAI Decisions API: A Guide

Compare Microsoft Decision 1 and OpenAI Decisions API by access, answer types, probability handling, evaluation design and the cost of a useful decision.

October 9, 2026 · 10 minRead
AI Development

Open Decision Models Compared: Clef, Decider 2B and Jev

Cloudflare's Clef and Amazon's Decider 2B put decision models like closed Jev into open weights. Licence, size, context, latency and benchmarks in one table.

October 1, 2026 · 6 minRead
AI Development

An Open Model That Makes Editable, Layered Designs

Ming-Image-0.1-Design makes UI, poster and infographic images with transparent backgrounds; its Layer variant splits a flat design into editable layers. MIT.

September 25, 2026 · 7 minRead
AI Development

AI Browser Landscape 2026: Atlas vs Comet vs Arc vs Dia

AI browser landscape 2026 — Atlas, Comet, Arc, Dia, Brave Leo, and Opera Neon. Feature matrix, market share estimates, and how agencies should prepare.

April 16, 2026 · 16 minRead
AI Development

AI Agent Marketplaces 2026: Discovery and Distribution

AI agent marketplace landscape — Claude Skills, GPT Store, MCP Hubs, Hugging Face Spaces, Replit Agent Market. Distribution strategy for agency builds.

April 16, 2026 · 16 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source