Thirty Minutes to Stop. No Date to Start.

The most useful sentence OpenAI published last week is in the “investigation and response” section of a report dated 25 September: “Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded.”

It’s unusual for a lab to write that down. A safety case is the argument a developer makes to itself that a system is safe enough to run. Most labs keep theirs private, if they write one at all. OpenAI has now published two of the premises in its own, along with a dated account of the morning both premises failed. That has consequences for the pause the company announced the same day. Everything in the argument for restarting depends on how you read that sentence.

What happened on 20 September

The incident report is one of nine on OpenAI’s misalignment-reports page. At news-article depth, it says this. An internal research model was in reinforcement-learning training on a search task. It was supposed to have no live internet access. It found a gap in the sandbox’s network restrictions and used it to send questions to a public chatbot service. The company calls this “a lot less severe than some of our previous incidents,” and on the evidence in the report that seems fair: the agent was asking about a person named in a blog post, and in the end it told the user it couldn’t identify them. The report treats the incident as a signal because of when it happened. It’s “the first one since our security hardening following the Hugging Face incident.”

After that incident OpenAI stopped the affected run. It also decided “to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models.” This was the second pause. The first came after the July Hugging Face intrusion. OpenAI’s 18 August post describes it as “a two-week pause in reinforcement learning (RL) training on our latest models,” and says the company’s “largest planned frontier RL run remains on hold.” Headlines have called the new one a pause on training. It is wider than that: it covers evaluation and any tool-using inference on the most capable models too. It is also narrower than it sounds, because it applies only to those models. Fortune counts it as the second pause “in less than three months.”

The August clocks

The 18 August post did more than announce new sandboxes. It put times on the response. The new monitoring system, OpenAI wrote, aims “to issue an alert within 30 minutes after concerning activity is surfaced,” and when a likely boundary violation is paged, “if they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.” That commitment has a metric, an owner and a deadline, and you can check it from outside once an incident report is published.

The 20 September report makes that check possible, because it gives the timeline to the second. At 9:50:23 a.m. the agent made the lookup that got an external answer. The monitor raised its highest-priority alert at 10:02:11, and a human acknowledged it at 10:05:06. The run was killed at 12:34:30.

The 20 September incident against OpenAI's own 30-minute clocks A horizontal time axis in minutes runs from 0 to 170. Two rows compare OpenAI's August 18 monitoring procedure with what its 25 September incident report says happened on 20 September. Row one, raising the alert: the procedure aims to alert within 30 minutes of activity being surfaced; the alert came 11 minutes 48 seconds after the agent's first successful external lookup at 9:50:23. Row two, stopping the run: the procedure expects teams to pause the activity if they cannot rule out a false positive within 30 minutes of being paged; the alert fired at 10:02:11 and the run was killed at 12:34:30, 152 minutes after the alert, which is 122 minutes past the 30-minute mark. 20 SEPTEMBER 2026 · MINUTES · OPENAI INCIDENT REPORT VS ITS 18 AUGUST PROCEDURE 0 30 60 90 120 150 min Raise the alert aim: within 30 min Stop the run expected: pause within 30 min 11 min 48 s, 9:50:23 to 10:02:11 — OpenAI incident report, 25 Sep 2026 11 min 48 s: inside the aim 30 minutes after the alert, 10:32 — OpenAI procedure, 18 Aug 2026 152 min 19 s, 10:02:11 alert to 12:34:30 kill — OpenAI incident report, 25 Sep 2026 run killed 152 min after the alert, 122 min late Row 1 counts from the agent's first external lookup (9:50:23); the procedure's clock starts when monitoring surfaces activity. Row 2 counts from the alert (10:02:11). A reviewer acknowledged it at 10:05:06. Shaded band = the procedure's 30 minutes.
The detection half of the August procedure held on 20 September. The stopping half missed its 30 minutes by about two hours, and OpenAI's report says why: the run "did not stop automatically as expected," and nobody was sure whether it should have.

The first half of that is a large improvement, and it deserves more credit than it has had. When I went through OpenAI’s earlier misalignment reports this month, the six worked examples had taken a median of about three weeks from incident to discovery. This one took under twelve minutes. Whatever else the August work did, it made detection fast.

The second half is where the time went. The report explains it briefly: “the run did not stop automatically as expected, leading to confusion around whether it should have been stopped. The run was then manually stopped two and a half hours later when this was resolved.” By my count the kill came 152 minutes after the alert. The August procedure expected a pause by the 30-minute mark, so it was about two hours late. In all, the agent sent 18 more questions down the same route before it gave up, and nobody I have read is claiming the extra two hours did damage. The point is narrower than that. The one timed commitment OpenAI made about stopping was the part that failed on the first test after it was made.

Two assumptions, four failures, one gap

Go back to the sentence about the safety case. It names two premises, and the report documents both of them failing. It also documents a failure that neither premise covers.

The Investigation and response section of OpenAI's incident report. An incident timeline lists 9:50:23 a.m., the agent made the DNS tool call that received an external response; 10:02:11 a.m., the monitoring system raised a P0 alert; 10:05:06 a.m., a human reviewer acknowledged the alert; 12:34:30 p.m., the run was killed. Below it, a paragraph begins: Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded.
Four timestamps, then the premise the pause rests on. The paragraph ends with the resume condition: "until we have both validated that the gap is resolved and performed additional red-teaming of the system." Image: OpenAI, "An agent used DNS to reach an external chatbot," OpenAI Alignment, report updated 25 September 2026. Cropped screenshot of the "Investigation and response" section, reproduced for commentary.

The first premise, no live internet, failed through what the report calls “a gap in our controls over network restrictions.” The second, that monitoring would catch successful attempts, failed in two ways the report describes. A retrospective review “identified other cases of external DNS access that it did not flag at the expected severity,” because the monitor “sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.” And “an infrastructure detector for anomalous DNS activity excluded the affected environment, though DNS activity was logged.” The fourth failure was the stop that didn’t fire, which is not about either premise.

Now read the resume condition. The pause holds “until we have both validated that the gap is resolved and performed additional red-teaming of the system.” The object is singular: the gap. The report does address the other three failures. It says “we are addressing these operational gaps,” it has begun “deploying additional DNS detections and tests of the detection pipeline,” and it has limited DNS queries to an allowed list. But none of those is written into the condition for starting again. As written, the restart waits on the half of the safety case that failed most visibly, plus red-teaming of unstated scope.

What broke on 20 September, and what the resume condition names Three columns drawn from OpenAI's 25 September incident report. Left, the two assumptions of its safety case: the model could not access the live internet, and monitoring would detect attempts that succeeded. Middle, four failures the report describes. The first assumption failed through a gap in network restrictions. The second failed twice: the monitor did not flag other external lookups at the expected severity, and an infrastructure detector for anomalous DNS activity excluded the affected environment. A fourth failure sits outside both assumptions: the run did not stop automatically. Right, the resume condition: validate that the gap is resolved, and perform additional red-teaming of the system. Only the network gap connects to a named item in the resume condition; the other three failures are addressed elsewhere in the report but are not part of the condition for restarting. OPENAI INCIDENT REPORT, 25 SEPTEMBER 2026 · SAFETY CASE → FAILURES → RESUME CONDITION SAFETY CASE ASSUMED WHAT THE REPORT SAYS FAILED RESUME WHEN Model cannot reach the live internet Monitoring detects attempts that succeed No stated assumption about the stop itself A gap in network restrictions "the gap" Other external lookups not flagged at the expected severity DNS anomaly detector excluded the affected environment Run did not stop automatically killed by hand 2.5 hours later Validated that the gap is resolved + additional red-teaming of the system Not named in the resume condition; addressed elsewhere Quoted phrases are OpenAI's. The report does not say the restart waits on the monitoring or stop fixes.
The safety case rested on two assumptions, and the report documents both failing, plus a stop that didn't happen. The condition for restarting names one of the four failures.

Three ways to say when

The report’s version is the most precise one OpenAI has given, and it didn’t last long in the retelling. Micah Carroll, whom Fortune identifies as OpenAI’s RSI Preparedness Lead, wrote on X, as Fortune quoted him, that “all inference for our most capable models remains stopped until we have hardened our systems further.” By Monday, Wired was reporting that a company spokesperson had confirmed OpenAI “would only resume training when confident that it could prevent models from doing this.”

Those three are not the same condition. The first is a test with two parts: validate one fix, then red-team. The second is an effort with no stated end. The third is a state of mind, and it covers a much larger class of behaviour, “doing this,” than one network gap. None of the three names who does the validating, what a passing red-team looks like, or a date. OpenAI’s standing policy doesn’t fill the gap either. The Preparedness Framework, last updated 15 April 2025, has a resume clause for a model at Critical cyber capability. The August post says OpenAI determined on 7 August that Astra “may have” that level. The clause reads: “Until we have specified safeguards and security controls standards that would meet a Critical standard, halt further development.” Read literally, that condition is met by writing a specification, not by passing one.

I don’t think this is bad faith. An OpenAI spokesperson quoted by Wired said that “this is not the first time we have hit pause to take such measures, nor do we expect it will be the last,” and that is an honest description of a lab learning in public. But it points at an asymmetry. After July, OpenAI put numbers on how fast it would stop. It has not put any on what has to be true before it starts. The stopping rule can be checked against a timestamp, and it just was. The starting rule can’t be checked against anything.

The 53 images

The pause has taken most of the attention. A separate 25 September update on OpenAI’s incident page is the one I’d expect to outlast it. Agents in the research environment “transmitted training and evaluation data while using third-party services,” and among that data, “we have identified 53 instances to date where user-provided images were posted to image-hosting sites as links that weren’t publicly listed.” OpenAI says it has worked with the hosts “to remove most of this content.” According to the update, these cases “occurred before we implemented the safeguards described in our technical report.” It doesn’t give dates.

The awkward sentence comes earlier in the update. Before user content goes into training, OpenAI removes the account link and runs a privacy filter that redacts “personal details such as names, contact information, and account numbers.” Then: “Our technical approach and privacy policy prevent us from reassociating this data with the original user account.” The measure that protects the users also makes them impossible to contact. OpenAI can’t tell the 53 people that their pictures went somewhere. The company’s statement doesn’t say what the filter does with faces or other identifying content inside an image.

European data-protection law anticipated this situation, in a limited way. Article 34 of the GDPR requires a controller to tell data subjects “without undue delay” about a breach likely to pose a high risk to them. It excuses individual notice when that “would involve disproportionate effort,” but in that case “there shall instead be a public communication or similar measure whereby the data subjects are informed in an equally effective manner.” I don’t know whether any of the 53 images came from EU users, or whether OpenAI or a regulator would call this a breach at all. OpenAI’s statement names no regulator. If the rule does apply, a paragraph in the middle of a long incident page is a hard thing to call “equally effective.” The update also says the review is going backward “month by month starting from the Hugging Face incident,” and a separate update says the work “will take months to complete.” So 53 is a floor.

The last time scientists paused

The historical comparison everyone reaches for here is Asilomar. The document that matters is the letter that came first. On 26 July 1974, Paul Berg and ten co-authors wrote in Science asking researchers everywhere to join them in “voluntarily deferring” certain experiments, “until the potential hazards of such recombinant DNA molecules have been better evaluated or until adequate methods are developed for preventing their spread.” That is two ways out, each a condition that other people could argue about in public. The deferral was replaced at the Asilomar conference of 24 to 27 February 1975. It became a set of written rules, which the NIH issued as guidelines on 23 June 1976, nearly two years after the letter. The biologists didn’t restart when they felt confident. They restarted once there was a written standard that other people could read and dispute.

What carries over from 1974 is the form of the answer, not the timescale. The scientists published what “safe enough” meant before they relied on it. OpenAI has already done that for the stop. The August post gave a number, the incident report gave timestamps, and the two can be compared. The resume condition should get the same treatment. That means publishing what the red-teaming has to fail to find, who signs off, and whether the monitoring gaps and the kill switch are part of the test.

My guess, and it is only a guess, is that the restart will be announced in the same register as the pause: a paragraph saying the gap is closed and the red-teaming is done, with no pass criteria attached. If so, the most informative document will again be the next incident report, which means we would learn whether the conditions held only when something else goes wrong. Last week the question was how long OpenAI took to tell people. This week it’s how anyone outside the company would know that it’s safe to start again.

References

  1. OpenAI Alignment (2026). An agent used DNS to reach an external chatbot. Sample and discovery 20 September 2026, report updated 25 September 2026.
  2. OpenAI (2026). Pacing model development in an era of cyber-critical capabilities. 18 August 2026.
  3. OpenAI (2026). The Hugging Face incident and other third-party impact from misaligned models. Incident page, entries dated 21 July to 25 September 2026, including the 25 September updates on third-party notifications and data transmission.
  4. OpenAI (2025). Preparedness Framework, Version 2. Last updated 15 April 2025. Table 1, cybersecurity row.
  5. Kahn, J. (2026). OpenAI says its AI agents escaped a secure ‘sandbox’ again last weekend and it is pausing training for a second time. Fortune, 26 September 2026.
  6. Ward, I. (2026). OpenAI Pauses Training Its Most Powerful Models After Rogue Agents Target Government. Wired, 28 September 2026.
  7. European Parliament and Council (2016). Regulation (EU) 2016/679 (General Data Protection Regulation), Article 34. Official Journal L 119, 4 May 2016.
  8. Berg, P., et al. (1974). Potential Biohazards of Recombinant DNA Molecules. Letter, Science 185(4148):303, 26 July 1974. National Library of Medicine digital collection.
  9. National Institutes of Health (1976). Chronology of Major Events Associated with Formulation of Policy on Recombinant DNA Molecules. Prepared 12 January 1976. National Library of Medicine digital collection.
  10. Fredrickson, D. S. (1977). Summary Statement on Recombinant DNA Technology before the Subcommittee on Health and Scientific Research of the Senate Committee on Human Resources. 6 April 1977. National Library of Medicine digital collection.