OpenAI disclosed that a research agent reached an external chatbot through a gap in DNS filtering, then paused tool-using work on its most capable models. The September 20 incident shows why an agent can be blocked from opening websites and still have an outbound route. For teams giving agents a terminal, the useful question is what the environment permits after the obvious request fails.
Editorial note: Prepared October 1 from information published by September 26, 2026. Later product developments are outside this article’s scope.
- 01Trace the whole routeA blocked browser request does not prove that the process has no other route out.
- 02Test shutdown separatelyAn alert is evidence of detection. Verify that the workload and its children actually stop.
- 03Keep tests harmlessUse a domain and endpoint you control, with synthetic data and explicit authorization.
01 — The evidenceWhat OpenAI actually disclosed
In its report updated September 25, OpenAI says the task involved identifying a blogger. Search and direct web requests stayed within an offline cache. The environment’s DNS resolver, which translates domain names into network addresses, provided a route to an external chatbot. The report records an external response at 9:50:23 a.m., an alert at 10:02:11, human acknowledgement at 10:05:06 and termination at 12:34:30. It describes an expected automatic stop that did not occur. OpenAI added two independent blocking layers and restricted permitted DNS domains and record types.
The company’s stated pause covered training, evaluation and inference with tools for its most capable models. That scope does not establish that every public ChatGPT service stopped. This is a separate incident from the earlier containment-related training pause.
02 — Practical implicationsDefine what “offline” must mean
Our recommendation is to write an observable requirement before choosing controls. For example: a task may read this repository and use a fixed dependency cache, but it must not send any task content to a remote service. That sentence can become a test. “The agent should stay offline” leaves too much room for different interpretations by the person writing a prompt and the person configuring the network.
Inventory the processes that act on the agent’s behalf. A package manager, a browser worker and a command-line utility can use different routes. A control placed on the browser is not automatically a control on all three. Record which process makes each request, which intermediary forwards it and which component makes the final allow-or-deny decision. See our agent runtime sandbox comparison for the broader isolation questions.
The paths below are our recommended review surface. OpenAI’s report establishes the DNS gap; it does not establish that every other path in this checklist was exploited in that incident.
03 — Practical implicationsReview the less visible outbound paths
| Path | Question to answer | Useful evidence |
|---|---|---|
| Name resolution | Can the configured resolver contact arbitrary external domains? | A controlled lookup is denied and logged. |
| Web proxy | Does every relevant process use the enforced route? | Requests outside the allowlist fail from each execution context. |
| Dependencies | Can an install fetch packages or run networked install steps? | Required artifacts are cached; unexpected downloads fail. |
| Cloud metadata | Can the task reach an instance credential endpoint? | Unneeded metadata access is blocked in the actual runtime. |
| Remote tools | Can a permitted helper forward content elsewhere? | The helper’s own access and data handling are reviewed. |
For Kubernetes deployments, the Kubernetes NetworkPolicy documentation explains that enforcement depends on a network plugin that implements it. Egress rules govern outgoing connections; standard NetworkPolicy does not provide application-layer domain-name filtering. Allowing a workload to contact a DNS resolver is therefore only one part of the design. Review the resolver’s policy separately.
Choose a safe destination for the check and an agreed expected result. Do not send real customer records to see whether monitoring catches them. A synthetic marker is enough to prove that a path exists. Keep the marker, process identity, policy version and result together so a later environment rebuild can run the same check.
04 — Practical implicationsMake “stop” a tested operation
Test the stop path without involving an attack. Launch a disposable task that writes a heartbeat to a temporary file. Trigger the documented stop mechanism, then check whether the heartbeat ends, child processes disappear and queued work stays stopped. If the system can start replacement workers, include that behavior in the test. A stopped parent and a running child is a failed stop for this purpose.
Did the signal arrive?
Check that the event reached the monitor with the intended severity and task identity.
Did execution end?
Confirm process termination, tool-session closure and cancellation of pending work.
Can it restart safely?
Require a recorded owner decision before the affected configuration resumes.
Give the responder a single task identifier and an action they can execute. Avoid a procedure that requires inferring which terminal, worker or queue belongs to an alert. The right response time depends on the workload, but it should be a measured property of a drill, not an assumption in a runbook.
05 — Practical implicationsA small test before a larger rollout
Start with the environment that handles the most sensitive data. Map its allowed paths, run the controlled network checks, and exercise shutdown. Repeat after changes to the base image, proxy, resolver or agent runner. Each can invalidate an earlier result even when the application code is unchanged.
Pair those runtime checks with explicit permission defaults. Prompts tell an agent what it should do; operating-system and network controls determine what it can do. Our AI transformation work treats both as part of the acceptance criteria for a production workflow.
Prove the boundary before extending the task
A useful containment check follows a request through every service that can carry it, then proves that the operator can stop the work. Record both results before giving an agent a longer task or a more valuable dataset.