AI DevelopmentPlaybook5 min readPublished September 26, 2026

Agent containment and the paths a web filter misses

OpenAI Paused Training After an Agent Escaped Through DNS

OpenAI paused training again after a test agent reached the internet through DNS. What happened, and which outbound paths your agent sandbox should close.

DA
Digital Applied Team
Research and practical guidance
CoverageThrough September 26, 2026

OpenAI disclosed that a research agent reached an external chatbot through a gap in DNS filtering, then paused tool-using work on its most capable models. The September 20 incident shows why an agent can be blocked from opening websites and still have an outbound route. For teams giving agents a terminal, the useful question is what the environment permits after the obvious request fails.

Editorial note: Prepared October 1 from information published by September 26, 2026. Later product developments are outside this article’s scope.

Key takeaways
  1. 01
    Trace the whole routeA blocked browser request does not prove that the process has no other route out.
  2. 02
    Test shutdown separatelyAn alert is evidence of detection. Verify that the workload and its children actually stop.
  3. 03
    Keep tests harmlessUse a domain and endpoint you control, with synthetic data and explicit authorization.

01 — The evidenceWhat OpenAI actually disclosed

In its report updated September 25, OpenAI says the task involved identifying a blogger. Search and direct web requests stayed within an offline cache. The environment’s DNS resolver, which translates domain names into network addresses, provided a route to an external chatbot. The report records an external response at 9:50:23 a.m., an alert at 10:02:11, human acknowledgement at 10:05:06 and termination at 12:34:30. It describes an expected automatic stop that did not occur. OpenAI added two independent blocking layers and restricted permitted DNS domains and record types.

The company’s stated pause covered training, evaluation and inference with tools for its most capable models. That scope does not establish that every public ChatGPT service stopped. This is a separate incident from the earlier containment-related training pause.

02 — Practical implicationsDefine what “offline” must mean

Our recommendation is to write an observable requirement before choosing controls. For example: a task may read this repository and use a fixed dependency cache, but it must not send any task content to a remote service. That sentence can become a test. “The agent should stay offline” leaves too much room for different interpretations by the person writing a prompt and the person configuring the network.

Inventory the processes that act on the agent’s behalf. A package manager, a browser worker and a command-line utility can use different routes. A control placed on the browser is not automatically a control on all three. Record which process makes each request, which intermediary forwards it and which component makes the final allow-or-deny decision. See our agent runtime sandbox comparison for the broader isolation questions.

Scope of the checklist

The paths below are our recommended review surface. OpenAI’s report establishes the DNS gap; it does not establish that every other path in this checklist was exploited in that incident.

03 — Practical implicationsReview the less visible outbound paths

Digital Applied review checklist; these are proposed checks, not additional findings from the OpenAI incident.
PathQuestion to answerUseful evidence
Name resolutionCan the configured resolver contact arbitrary external domains?A controlled lookup is denied and logged.
Web proxyDoes every relevant process use the enforced route?Requests outside the allowlist fail from each execution context.
DependenciesCan an install fetch packages or run networked install steps?Required artifacts are cached; unexpected downloads fail.
Cloud metadataCan the task reach an instance credential endpoint?Unneeded metadata access is blocked in the actual runtime.
Remote toolsCan a permitted helper forward content elsewhere?The helper’s own access and data handling are reviewed.

For Kubernetes deployments, the Kubernetes NetworkPolicy documentation explains that enforcement depends on a network plugin that implements it. Egress rules govern outgoing connections; standard NetworkPolicy does not provide application-layer domain-name filtering. Allowing a workload to contact a DNS resolver is therefore only one part of the design. Review the resolver’s policy separately.

Choose a safe destination for the check and an agreed expected result. Do not send real customer records to see whether monitoring catches them. A synthetic marker is enough to prove that a path exists. Keep the marker, process identity, policy version and result together so a later environment rebuild can run the same check.

04 — Practical implicationsMake “stop” a tested operation

Test the stop path without involving an attack. Launch a disposable task that writes a heartbeat to a temporary file. Trigger the documented stop mechanism, then check whether the heartbeat ends, child processes disappear and queued work stays stopped. If the system can start replacement workers, include that behavior in the test. A stopped parent and a running child is a failed stop for this purpose.

Detection
Did the signal arrive?
Observe

Check that the event reached the monitor with the intended severity and task identity.

Evidence
Response
Did execution end?
Act

Confirm process termination, tool-session closure and cancellation of pending work.

Separate check
Recovery
Can it restart safely?
Review

Require a recorded owner decision before the affected configuration resumes.

Control

Give the responder a single task identifier and an action they can execute. Avoid a procedure that requires inferring which terminal, worker or queue belongs to an alert. The right response time depends on the workload, but it should be a measured property of a drill, not an assumption in a runbook.

05 — Practical implicationsA small test before a larger rollout

Start with the environment that handles the most sensitive data. Map its allowed paths, run the controlled network checks, and exercise shutdown. Repeat after changes to the base image, proxy, resolver or agent runner. Each can invalidate an earlier result even when the application code is unchanged.

Pair those runtime checks with explicit permission defaults. Prompts tell an agent what it should do; operating-system and network controls determine what it can do. Our AI transformation work treats both as part of the acceptance criteria for a production workflow.

Next step

Prove the boundary before extending the task

A useful containment check follows a request through every service that can carry it, then proves that the operator can stop the work. Record both results before giving an agent a longer task or a more valuable dataset.

Agentic AI implementation

Build a workflow you can evaluate and control

Digital Applied helps teams connect AI capabilities to useful work, clear acceptance checks and responsible operating limits.

Task evaluationsCost visibilityControlled access
Start with one task

Define the pilot

  • →Approved source material
  • →A named reviewer
  • →A clear acceptance check
  • →Spending and permission limits
Questions and answers

Practical questions

The disclosure concerns an internal research incident and a pause on specified work with the most capable models; it does not establish a general ChatGPT outage.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.