Step 5 Preview is a model you can evaluate through hosted access now, while its downloadable weights remain a separate promised milestone. That distinction matters more than a leaderboard position if you are deciding what your team can deploy this week. Build the trial around the endpoint that exists, and keep future self-hosting out of the current acceptance decision.
Product documentation and current access details were checked on October 11, 2026. The workflow examples are proposed evaluation designs, not reports of completed customer deployments.
- 01Use an exact routeRecord the provider and model identifier instead of relying on a display name.
- 02Date the price observationOctober 11 catalogue values are not evidence of earlier or permanent terms.
- 03Keep weights in the future tenseStepFun says October 15; confirm the files and license when that date arrives.
- 04Evaluate the assembled agentCheck source selection, tool arguments and accepted outcomes, not just fluent answers.
01 — Release boundarySeparate the model from the access milestones
StepFun’s model page describes Step 5 Preview as a mixture-of-experts model with 600 billion total parameters, 27 billion active per token, a one-million-token context window and vision input. The page says hosted products and API access are available, with open weights planned for October 15. Those are vendor statements; they do not establish a successful deployment in your environment.
The OpenRouter route appeared in the catalogue with an October 8 creation timestamp. A catalogue timestamp is not proof of the model’s first public launch, and this guide does not treat it as one. The distinction prevents a listing event from becoming a misleading release story. For a buyer, the useful facts are where access works, what the endpoint accepts and what obligations come with it.
Keep three milestones in your evaluation record: announced capability, usable hosted access and downloadable artifacts. A team can complete a useful hosted trial without the third milestone. A team that requires its own infrastructure cannot complete its deployment review merely because the first two are satisfied. This separates a promising option from a dependency that might delay an already committed project.
Hosted endpoint
Test behavior and integration through the actual API route available to your team.
Downloadable weights
Check the published files, license, hardware guidance and serving behavior after release.
02 — Route identityRecord the endpoint you are really buying
The exact aggregator identifier is stepfun/step-5-preview. In the October 11 OpenRouter snapshot, it advertises text, image and video input with text output, a 1,000,000-token context length and a 64,000-token maximum completion value. These are route metadata, not a promise that every direct provider interface exposes identical controls or limits.
Before a pilot, send a small representative request through your intended account and inspect the actual result. Confirm the accepted input type, response structure and usage reporting. A multimodal product page can describe the model family while a particular endpoint supports a narrower subset. If your workflow depends on video, that request needs its own integration check rather than an assumption based on vision support.
Save the route configuration alongside the evaluation input. If routing or provider selection changes later, you need to know whether the behavior comparison used the same serving path. Our frontier-model selection guide covers the larger choice; the procurement detail here is that two requests with the same friendly model name can still differ in access, limits and billing. Record enough information to repeat the request meaningfully.
A large request limit does not prove that the model will find the right exception in a long document set. Test evidence selection and answer quality separately from whether the API accepts the payload.
03 — Cost observationUse dated prices and explicit arithmetic
The OpenRouter catalogue checked on October 11 lists $1 per million input tokens, $2.70 per million output tokens and $0.05 per million cached-input tokens for this route. These are a dated aggregator observation. They should not be presented as a direct StepFun quote, an October 8 historical price or a guarantee that every request qualifies for cached billing.
For a hypothetical request using 30,000 uncached input tokens and 3,000 output tokens, those rates produce a token charge of $0.0381: $0.03 for input and $0.0081 for output. This illustration excludes tools, retries and review work. It is not an observed production bill. If the request needs a second attempt, the cost of producing an accepted result increases even when the advertised token price stays unchanged.
Long context makes this arithmetic more consequential. It can be tempting to attach every document because the request fits, but irrelevant material still contributes to input usage and can make evidence selection harder to judge. Our human-review cost guide explains the other half of the bill: the correction effort required after the model returns. Measure both rather than declaring the cheapest token a business saving.
Suppose the team expects a repeated policy prefix to receive cached pricing. Record the usage returned by the endpoint for the actual requests instead of subtracting an assumed discount from the forecast. The prefix may change when the policy changes, and the request shape may not qualify as expected. A procurement sheet should therefore keep an uncached baseline and a separately evidenced cached case. This makes the business estimate resilient to an ordinary document update, rather than depending on an optimization that nobody has confirmed on the chosen route.
| Observed rate | Per million tokens | Important condition |
|---|---|---|
| Input | $1.00 | Uncached input in the catalogue observation |
| Output | $2.70 | Generated output tokens |
| Cached input | $0.05 | Only when the endpoint reports eligible cached use |
04 — Task designMake a small agent trial answer a real question
Choose a job whose outcome your team can inspect. A useful example is preparing a response from an order record and an approved policy. The agent reads the request, looks up the order, identifies the applicable rule and drafts an answer. The trial can remain read-only while still exposing important differences in reasoning, retrieval and tool handling.
Include cases where the information is incomplete. Give the agent two similar orders, a missing attachment, an outdated policy and a tool response that explicitly says the record was not found. The valuable behavior is preserving the uncertainty and taking the correct next step. A polished answer that quietly invents a missing detail should fail, even if it reads better than the current model’s response.
Keep the input set stable across candidates and record any prompt changes separately. Otherwise, the comparison mixes a model switch with a workflow redesign. Decide the acceptance criteria before reviewing results: correct record, correct policy, supported answer and no unauthorized action. These proposed tests make the release relevant to a business process without claiming that Digital Applied has already measured Step 5 against another model.
An illustrative difficult case could contain an order with two line items and a policy exception that applies to only one. Ask the agent to draft a response without changing the order. A clean pass preserves the item identities, applies the exception only where it belongs and tells the reviewer what remains unresolved. A partial pass might find the correct exception but apply it to the whole order. Recording that distinction is more useful than a single preference score, because it tells the team whether the model needs a clearer task boundary or whether the information supplied was insufficient.
- Use ordinary tasks and known difficult cases.
- Keep evidence, tool responses and final drafts together.
- Score the errors that would create extra work for your team.
05 — Action evidenceInspect tool use before trusting the final message
An agent’s final answer can conceal a failed lookup or a malformed request. Review the actual tool arguments and responses, especially where the model has to select between similar records. If the tool rejects an identifier, the assistant should not turn that rejection into a confident statement about the customer. The test should preserve the failed interaction rather than scoring only the eventual prose.
Design the tool description so required information is explicit. If a change needs an order identifier and a selected date, the agent should obtain both before proposing the action. Avoid giving the model a vague tool that accepts arbitrary text and then blaming it for every ambiguous interpretation downstream. The model and the integration each need a clear contract that a reviewer can inspect.
Use the tool-error decision reference to separate a correct retry, a request for missing information and a stop. Repeated attempts at the same rejected action can consume budget without improving the outcome. A useful preview trial records those attempts as part of the result. It does not erase them because the final answer eventually sounds plausible.
Add a case where the lookup returns a valid but unexpected record. The model should compare the result with the user’s identifiers rather than accept any successful response as the intended order. This exposes an integration weakness that ordinary not-found tests miss. Keep the expected comparison fields explicit: a display name alone may not distinguish two customers, while an exact order identifier can. The test is about following the tool contract and preserving uncertainty, not asking the model to infer a person’s identity from incidental details.
For a write-enabled trial, verify the destination record. A tool call being generated, a request being accepted and the intended business change being completed are separate events.
06 — Hosting pathKeep the open-weights plan out of today’s promises
The announced October 15 milestone can be strategically interesting without being usable today. Downloadable weights may offer additional deployment choices, but the actual license, artifacts and serving instructions determine what an organization can do. Until those are available, describe self-hosting as a future evaluation rather than a current delivery option.
A future deployment review needs different evidence from the hosted trial. It must cover operational ownership, hardware capacity, updates, monitoring and the behavior of the chosen serving stack. A strong hosted result does not settle those questions. Even when weights share the same model name, repeat the business evaluation on the deployment you intend to run before assuming equivalent behavior.
Preserve the useful parts of the hosted experiment outside a vendor dashboard: test inputs, expected outcomes, scoring categories and accepted examples. That work remains valuable if the deployment route changes. It also makes a decision to stay hosted easier to defend. The goal is a dependable workflow, not satisfying an architectural preference after the model trial has already shown what the business needs.
- Confirm the released files and license rather than relying on an announced date.
- Assign operational ownership before choosing self-hosting.
- Repeat task evaluation on the actual serving configuration.
07 — Evidence standardRead vendor demonstrations as leads for testing
StepFun’s page presents examples and benchmark results for software and professional work. Those can suggest useful tasks to investigate, but they remain vendor evidence. This guide does not reproduce their measurements or infer a universal advantage from them. A result obtained under a vendor’s task setup may depend on tools, prompts and evaluation rules your application does not share.
Ask what a promising example would look like in your own process. If the attraction is long-running work, test how the agent handles a dependency that remains unresolved. If it is document analysis, test conflicting versions and a relevant exception buried among similar passages. If it is coding, review whether the changes satisfy your project’s actual requirements and tests. Translate capability claims into observable acceptance conditions.
Also retain a reason not to migrate. If your current route already meets the requirement, a new model needs to justify the integration and review effort. Our AI transformation service focuses these evaluations on a defined operational outcome. The useful recommendation may be a narrow additional route, a later revisit or no change at all, depending on the evidence.
This article provides an evaluation design and dated access information. It does not claim a completed head-to-head quality test, a latency advantage or a guaranteed production saving.
08 — Adoption decisionTurn the result into a maintainable routing rule
Write the outcome of the pilot as an operating rule. For example, the model may prepare internal policy drafts from approved sources while external actions remain on the existing workflow. Name the exception categories that require review and the conditions that send the task back to the prior route. That is more useful than a general instruction to adopt the latest model everywhere.
A preview also needs a change boundary. Record the endpoint and evaluation date, and repeat the relevant cases when serving behavior or limits change. You do not need to rerun every experiment after an unrelated change, but a new input format, tool schema or routing path can invalidate the part of the test that depended on it. Keep the evidence connected to the configuration it actually assessed.
Step 5 Preview adds another candidate for agent work and a possible future weights-based deployment. Its value for a team will come from completed tasks and controlled failures on that team’s workflow. Keep the access facts precise, the cost arithmetic dated and the migration reversible, and the trial can produce a useful decision even if the final answer is to wait.
Set a practical trigger for the next evaluation. A changed tool schema should rerun tool-handling cases; a new document source should rerun evidence-selection cases; a different serving provider should rerun access, limits and representative behavior. Naming those triggers keeps maintenance proportional. It avoids both extremes: treating a preview trial as permanent proof or rerunning every historical experiment after an unrelated configuration edit.
- Adopt only for the task class the trial supports.
- Keep a fallback and the reason for using it visible.
- Revisit the weights milestone with new evidence when the artifacts exist.
Test the route that exists today
Record the exact endpoint and run one repeatable task comparison. Include missing information, tool failures and a check of the final destination.
Treat the October 15 weights plan as a separate follow-up decision. Hosted access can be useful now without turning a future release promise into a current commitment.