Mistral Large 4 is worth a bounded trial if your team wants another capable model for tool-using workflows or a future self-hosting option. It is not a reason to replace a working agent overnight. Mistral announced the public preview on October 6, 2026, while downloadable weights were still a promised follow-up. Keep those two milestones separate when deciding what you can actually test.
This is an evaluation guide, not a report of our own head-to-head benchmark. The examples below are proposed business tests. They are designed to show whether a model change improves the whole job, including the checks and corrections a person still has to make.
- 01Preview is a useful testing boundaryStart with copied tasks and restricted tools before putting the model in a customer-facing workflow.
- 02Hosted access is not self-hostingAn API account and downloadable weights solve different operational problems.
- 03Judge completed workTrack correct actions, evidence and review effort alongside token charges.
- 04Keep the old route availableA narrow rollout needs a tested fallback and a clear reason to expand.
01 — Release scopeWhat the announcement actually changes
The Mistral announcement establishes the important launch facts: a public preview can be tried through the hosted API, while the weights are intended to follow by the end of October. That gives a buyer an immediate experiment and a possible future deployment choice. It does not make the second choice available today.
For an agency building a document assistant, the immediate experiment could be straightforward: give Large 4 the same approved documents, questions and tool descriptions as the current model. Compare the resulting work without changing access permissions or business rules. If the new model needs a different prompt, keep that as a separate trial so you know whether the improvement came from the model or the redesigned workflow.
A common mistake is to bundle every attractive part of the launch into one purchasing decision. A team may like the possibility of controlling its own deployment, yet need a managed endpoint this month. Those requirements should produce two evaluations with different owners. The hosted trial assesses behavior and integration; a later infrastructure review assesses whether running the model is economically and operationally sensible.
Treat the announced weights date as a plan until the files, license and deployment instructions are actually available. Do not make a customer delivery promise depend on an unreleased artifact.
02 — Business fitChoose the job before the benchmark
Start with a job that already has an answer key or an observable outcome. A useful first candidate is an assistant that reads an order, looks up a policy and prepares a draft response. You can check whether it selected the right policy, preserved the order details and avoided committing to something the business cannot deliver. That is a more useful test than asking each model a few impressive questions.
Keep a small collection of routine cases and difficult cases. Include a missing attachment, contradictory documents, an unavailable lookup and a customer whose request falls outside policy. A model that produces polished prose for the routine cases but invents a resolution for the exception is not an operational upgrade. The exception is precisely where a person needs the system to be predictable.
Decide the acceptance criteria before inspecting outputs. Otherwise, a compelling example can quietly change what the team means by success. Our guide to choosing a frontier model provides the broader selection context; this trial should narrow that question to one task, one information boundary and a defined reviewer decision. Keep quality categories separate rather than reducing every observation to a single score.
Draft with evidence
The agent prepares a response and cites the policy it used. A person verifies both before sending.
Act through a narrow tool
The agent may perform a specific reversible action after the draft workflow has demonstrated reliable checks.
03 — Action qualityTest the tool contract as well as the answer
An agent can sound correct while calling the wrong tool, omitting a required argument or interpreting a failed response as success. Preserve the exact tool schema in your comparison and inspect the structured call alongside the final message. If a customer has two open orders, the important result is not whether the explanation is friendly; it is whether the agent selected the intended order before preparing a change.
Use a test destination whose records you can inspect. A lookup can return a known result, a simulated timeout or an explicit not-found response. The reviewer should be able to trace the model's next step to that response. When the tool reports uncertainty, the agent should preserve it instead of smoothing it into an answer that sounds complete. This is a test of the assembled system, not a general judgment about the model's intelligence.
Also test what happens after a rejected action. Does the agent correct the argument, ask for missing information, try a different legitimate route or repeat the same failing call? Repetition can consume money and still leave the task unfinished. The tool-error decision reference is useful here because an integration needs deliberate behavior for each failure class, not one instruction to try harder.
- Record the requested action and the actual tool arguments.
- Keep the tool response and the agent's interpretation together.
- Check the destination state before accepting a success message.
04 — Document loadSeparate a large context from useful evidence
The current model documentation lists a large context window, but capacity is only the first question. A request can fit and still contain the wrong document, several conflicting versions or so much irrelevant material that the important exception is hard to identify. Evaluate evidence selection separately from whether the API accepts the payload.
Consider an assistant answering a question about a returns policy. The folder contains the published policy, an older draft and a regional exception. Loading all three gives the model more text without settling which source governs the customer. Your test should require the assistant to identify the applicable version and explain any unresolved conflict. A correct answer drawn from the wrong version should not count as a clean pass.
Use realistic document lengths, including the instructions and tool definitions that accompany them. Leave room for the output and any intermediate information your system adds. The document-reading evidence guide explains why a citation is only useful when it supports the claim actually made. A million-token headline does not replace that discipline, nor does it tell you how long the request will take on your selected endpoint.
Ask whether the model can find and apply the governing exception in your document set. Merely placing the exception somewhere inside a large request is not an evaluation.
05 — Trial economicsPrice an accepted result, not a headline token
Keep a dated price sheet for the endpoint you actually use. A direct API, an aggregator route and a future self-hosted deployment need not share the same charges or limits. Mistral's documentation checked on October 11 displayed sale rates of $0.68 per million input tokens, $0.07 for cached input and $2.09 for output. Treat those as a dated observation, not a permanent quote or proof of the exact launch-day terms.
For a hypothetical request with 20,000 uncached input tokens and 2,000 output tokens at those rates, the token charge is $0.01778: $0.0136 for input plus $0.00418 for output. That calculation excludes tools, retries, infrastructure and review. It is an illustration of the arithmetic, not a measured bill from a Large 4 trial. The invoice can differ when the selected provider or usage pattern differs.
Now include work the first answer creates. If the agent needs a second attempt, sends an unnecessarily large history or produces a draft that takes longer to correct, the cheaper token rate may not reduce the cost of the job. Record both the machine bill and the review burden. The existing human-review cost guide is the right companion to a price comparison because accepted output is the unit the business ultimately buys.
| Cost layer | Record | Why it matters |
|---|---|---|
| Model request | Input, output and cache usage | Makes the token calculation reproducible |
| Extra work | Retries and tool charges | Captures the cost of reaching an answer |
| Review | Time and correction category | Shows whether the result saves useful effort |
| Completion | Accepted or rejected outcome | Keeps failed jobs in the denominator |
06 — Hosting choiceTreat deployment control as a separate project
The possibility of downloadable weights can matter for buyers with infrastructure, customization or continuity requirements. It also transfers responsibilities that a managed service normally absorbs. A self-hosting decision should include capacity planning, model updates, monitoring, security controls and the people who will investigate a degraded service. A license permitting a deployment is not a guarantee that your organization can operate it well.
Keep the trial portable where that is practical. Store evaluation inputs, expected outcomes and scoring notes outside a provider's dashboard. Record the model and endpoint version used for each run. Those records let you repeat the business test if the hosted preview changes or a downloadable version becomes available. They do not prove the two versions will behave identically; that is exactly why you repeat the test.
For a small business, the strongest reason to try the hosted model may simply be supplier choice. There is no requirement to turn every successful trial into an infrastructure program. Conversely, a team that genuinely needs its own deployment should not assume a managed preview settles license, hardware or operational questions. Make those separate acceptance criteria rather than burying them in the model-quality score.
- Behavior: can the system complete the agreed work?
- Operations: can the chosen deployment sustain the required service?
- Governance: do the actual contract and controls meet the organization's needs?
07 — Change controlMake a preview rollout reversible
After the offline comparison, choose a narrow production boundary if the results justify it. Route a defined task class to the new model and keep a way to return it to the previous route. Avoid changing the retrieval system, tool permissions and model at the same time. When several moving parts change together, a failure becomes harder to attribute and the rollback becomes harder to execute.
A practical example is using Large 4 for internal response drafts while the current model continues to handle external actions. Review the drafts under the same policy as before and collect disagreements with the existing route. You are not trying to maximize the number of jobs on the new model. You are trying to learn where it is dependable, where it needs different instructions and where it should not be used.
Specify the rollback trigger in ordinary language. An unauthorized action proposal, repeated failure to select the governing document or a substantial increase in correction effort should cause investigation, not an argument about the launch benchmark. Our AI transformation service starts with this workflow boundary because a model substitution is valuable only when the team can explain and maintain the resulting process.
Preserve unsuccessful runs and difficult examples. Removing them after seeing the outputs makes a trial look cleaner while making the eventual routing decision less reliable.
08 — Decision pointThe useful outcome is a routing decision
A successful evaluation does not have to produce a universal winner. It may show that Large 4 is useful for document-heavy drafts, while another model remains preferable for a different tool workflow. It may show that the preview is promising but needs a specific integration repair. It can also show that switching would create work without delivering a meaningful benefit. Each is a legitimate business result.
Write the decision as a small operating rule: which tasks qualify, which information they receive, which actions remain restricted and who reviews an exception. Attach the examples that justify the rule. That is much easier to maintain than a broad instruction to use the newest model wherever possible. It also creates a clear next experiment when a new version or deployment option arrives.
Mistral's preview changes the set of options available to test. Your own task evidence should determine whether it changes the system you run. Keeping that distinction visible lets the team be interested in a new release without becoming dependent on promises, demonstrations or a result measured under someone else's conditions.
- Adopt for a defined task when the evidence supports it.
- Repair and retest when an integration issue obscures the result.
- Keep the existing route when the change does not improve useful work.
Run one comparison you can repeat
Choose a bounded workflow, preserve the inputs and review both the actions and the answer. Treat hosted access as available to evaluate and the promised weights as a separate future milestone.
The strongest reason to adopt Large 4 is a repeatable improvement in your own work, with an endpoint, cost record and rollback path that your team understands.