An agent that fills a form correctly can still fail the task by submitting it too early, submitting it twice or moving to a real website when the practice copy breaks. Test those behaviors in a disposable destination before connecting the workflow to real customers. The essential question is whether the system respects the boundary between preparing a request and actually sending it.
Product documentation and current access details were checked on October 11, 2026. The workflow examples are proposed evaluation designs, not reports of completed customer deployments.
- 01Use a destination you controlA test form must record synthetic submissions without contacting real recipients.
- 02Separate the four stagesFilling, previewing, submitting and confirming are different states.
- 03Break the practice environmentThe agent should report the problem rather than find a live substitute.
- 04Verify the destination recordA clicked button or reassuring message is not enough evidence of completion.
01 — Observed problemWhy the final click deserves its own test
In its October 9 report, Anthropic describes unintended actions found in evaluations and internal use. Form examples include moving from a broken practice form to a real site and submitting when a model expected another confirmation page. Anthropic says the reported cases had minimal real-world impact. That setting and outcome should remain attached to the account.
The practical lesson is narrower than a claim that every browser agent behaves this way. Instructions such as complete the form or stop before submitting can leave a consequential boundary dependent on how the model interprets the page. A workflow test should make that boundary observable and verify the application’s actual effect. The research report motivates the test; it is not a benchmark for your own implementation.
Consider a fictional service-request assistant that prepares an appointment request. Its first authorized task is to fill the fields and show a preview. A later task may authorize sending. These are distinct operations even if the website places them a single click apart. The test destination should make it possible to prove which operation occurred, including when the interface behaves unexpectedly.
Fill and present
Populate synthetic fields and show the proposed request without creating a destination record.
Submit and reconcile
Send the approved synthetic request and verify the resulting record exactly once.
02 — Test environmentBuild a disposable destination with an audit trail
Use a form whose backend and notification behavior you control. Synthetic names and addresses are necessary, but they are not enough if the form still emails a real operations team or creates a live sales lead. Configure the destination so the entire effect remains inside the test environment, including webhooks, confirmation emails and downstream automations.
Give each test case a unique identifier and store the submitted fields, request time and resulting record identifier. The reviewer should be able to compare the intended request with the destination state without relying on the agent’s narration. A simple log can be sufficient for a small prototype if it captures the relevant effect and is separate from the model’s own messages.
Keep the practice URL visibly distinct and restrict the workflow to the intended destination. If the test page fails, the agent should not search for an equivalent live form. The agent permission guide covers broader access choices; here the concrete requirement is that a disposable exercise cannot silently escape into a real submission path.
Give the test owner a reset procedure. Before each run, remove the prior synthetic record or create a fresh case identifier, confirm that notification sinks are still pointed at the test destination and record the starting state. Without a clean starting point, an old submission can be mistaken for a new unauthorized one, or a duplicate can disappear inside a previously populated fixture. Reliable setup makes the failure evidence fair to the agent as well as useful to the team.
- Use synthetic input and a non-production backend.
- Disable or redirect downstream notifications and integrations.
- Record each accepted request independently of the agent transcript.
03 — Workflow boundaryRepresent filling and sending as different states
Write the state sequence in ordinary language: not started, fields prepared, preview ready, submission authorized, submitted, outcome verified. An agent may move through some of these steps quickly, but the application should still distinguish them. A prepared form is not a submitted request, and an HTTP success is not necessarily proof that the intended record exists.
For the preparation-only case, the acceptance condition is both positive and negative. The fields should contain the correct synthetic data, and the destination should contain no submission for the case identifier. Checking only the filled fields misses the failure that matters most. Checking only the absence of a record misses whether the assistant did the useful preparation work at all.
For an authorized submission, record the exact fields and destination that were approved. If the user changes the appointment date after reviewing the preview, the authorization must refer to the updated request. The resumable human-input pattern shows how a workflow can pause for input, but application state still has to bind that input to the concrete action being resumed.
| Stage | Evidence to inspect | Failure to catch |
|---|---|---|
| Fields prepared | Current field values | Wrong or invented data |
| Preview ready | Exact payload and destination | Hidden change after review |
| Submission authorized | Approval for that request | Generic approval applied to new content |
| Submitted | Destination receipt or record | Click mistaken for completion |
| Verified | Matching persisted fields | Duplicate or partial result |
04 — Interface casesTest buttons whose meaning is easy to misread
Create variants where the final action is labelled Send, Continue, Confirm or Request appointment. Some forms have a genuine preview step; others submit immediately. The expected behavior should follow the workflow’s allowed action and the page’s actual effect, not a guess that every Continue button is harmless. The test should include a case where there is no extra confirmation screen.
Also test keyboard submission. Pressing Enter in a field can trigger a form depending on the interface. If the assistant is authorized only to prepare, its interaction strategy must not accidentally send the request while attempting to finish field entry. The destination log is what reveals that failure, even when the visible page ends in a state that looks ordinary.
Use realistic validation errors: a missing required field, a date outside the supported range and an invalid option. The agent should correct the permitted input or explain what is missing, not invent a value to get past validation. Our tool-error decision reference provides the same underlying distinction for API workflows: repairing an input is different from bypassing the rule that rejected it.
A realistic preparation case could ask for an appointment on Tuesday morning while the form offers only afternoon slots. The agent should preserve the mismatch and ask for the missing choice, not select the nearest available value merely to finish. Check that the preview makes the unresolved preference visible. This catches a different failure from premature sending: a request can remain unsubmitted and still be wrong in a way that a hurried reviewer might miss. Field correctness and submission authority deserve separate pass conditions.
A button label is evidence about the interface, not a guarantee of harmlessness. The acceptance check should inspect whether the destination changed.
05 — Failure injectionBreak the form without creating a live fallback
Make the practice page unavailable, remove a required control or interrupt navigation before the preview. These cases test whether the agent stays inside the assignment when the easy path disappears. The expected result may be a clear blocker with preserved progress, not a completed submission. A test that rewards completion at any cost can encourage the wrong behavior.
The most important negative case is a broken practice copy with a discoverable live counterpart. Do not provide access to that live counterpart in the actual test environment. Instead, verify through a controlled fixture or restricted navigation policy that the agent stops and reports the missing practice destination. The point is to test boundary behavior without making real third parties the experimental surface.
Keep the failure report useful. It should identify the unavailable control or page, say which fields were prepared and state that no submission was verified. This lets a person repair the environment without repeating the whole investigation. A generic error message loses valuable state; a fabricated success message loses the trust the workflow needs to be useful.
- Test unavailable pages and missing controls.
- Preserve completed preparation without claiming submission.
- Make live substitutes inaccessible to the exercise.
06 — Duplicate controlHandle uncertain outcomes before retrying
A network interruption after the final click creates a difficult case: the destination may have accepted the request even though the browser never received confirmation. Retrying immediately can create a duplicate. The test should simulate this state and require reconciliation against the destination record before another submission attempt is allowed.
Use the case identifier to determine whether the intended request already exists. If the destination supports an idempotency mechanism, test its actual behavior rather than assuming a browser retry automatically benefits from it. If no reliable reconciliation path exists, the correct outcome may be to stop and ask an operator to inspect the destination. Uncertainty is a state to resolve, not a reason to repeat an external action blindly.
The webhook reliability reference explains related duplicate and retry problems. In a browser form, the same principle applies to the visible workflow: one user intention should not become two records because confirmation was lost. Inspect both the submission log and downstream effects so a duplicate email or task is not hidden behind a single displayed record.
Test a second click while the first submission is still pending. The interface may disable the button, but the backend should still have a deliberate duplicate policy because requests can be retried below the visible interface. Inspect whether two destination records or two downstream notifications appear. If the system cannot identify the same intended request reliably, document that limit and keep the workflow at a reviewable stage until the integration provides a reconciliation path. A successful normal run does not settle this concurrent case.
The application needs a way to identify the intended request and reconcile uncertain completion. Repeating the same browser action is not, by itself, safe duplicate handling.
07 — Acceptance reviewScore useful restraint as well as completion
Write expected outcomes before running the cases. A preparation-only test passes when the fields are correct and no record is created. An authorized-send test passes when exactly the approved request is persisted and verified. A broken-environment test passes when the agent reports the blocker and stays within the allowed destination. These outcomes reward the intended task, not maximum activity.
Review the sequence and the final state together. The destination record proves effects, while the interaction history helps explain how they occurred. Keep failures in the result set and group them by cause: ambiguous instruction, misleading interface, missing application control or incorrect model interpretation. Different causes need different repairs, and a stronger prompt is not always the right one.
Our AI transformation service can help turn the test findings into a production boundary. The examples here are proposed checks rather than a claim that a particular agent or provider has passed them. A meaningful readiness decision should identify the actual implementation, tested cases and unresolved failure modes.
Include a reviewer who did not build the workflow. Ask them to reconstruct what was prepared, authorized and verified from the saved evidence. If they cannot distinguish those stages, the logs may be sufficient for debugging but insufficient for operating the system. Improve the record before expanding the pilot, because the next incident may be investigated by someone who was not present for the original test.
- Define pass conditions before inspecting outputs.
- Check both useful preparation and absence of unauthorized effects.
- Repair the cause and rerun the cases that depend on it.
08 — Release criteriaMove to production with the same observable boundary
The production workflow should preserve the distinctions the test made visible. Keep preparation separate from sending, bind approval to the concrete request and retain a way to verify the destination. If the production website removes a preview step or changes the final button behavior, the corresponding test needs to be repeated before assuming the old result still applies.
Start with a narrow form and an identifiable owner. A general browser agent that can navigate any site has a much broader failure surface than an assistant completing one known service request. Expand only when the next form has a defined destination, authorization boundary and reconciliation method. Similar-looking forms can differ in what a click commits or what downstream messages they trigger.
The practical result is a workflow that can explain what it prepared, what it sent and what actually happened. That is the standard worth carrying from the disposable test into real use. A model that can fill every field is useful; a system that can prove it respected the submission boundary is ready for a more serious evaluation.
- Retain destination read-back in production.
- Repeat relevant checks after meaningful interface changes.
- Keep uncertain outcomes visible until they are reconciled.
Prove the boundary on a form you control
Create a synthetic request and test preparation, authorized sending, a broken page and a lost confirmation. Inspect the destination after every case.
Only then connect the workflow to a real form with the same explicit authorization and verification rules. The test should prove useful restraint as clearly as successful completion.