← Gautam Parab

OpenAI Found the Message Board on May 25. It Told Us on September 16.

A disclosure policy is not a statement of values. It is a promise about latency. The criteria, the escalation ladder, the tone: none of it counts for anything unless the thing gets out of the building, and the only honest way to grade that is to count days.

On 16 September OpenAI published a framework for tracking, investigating and disclosing model misalignment, and, usefully, shipped it with six worked examples: six reports on behaviour observed during the training and evaluation of its models. Each report carries a date for the incident and a date for the discovery, down to the day. The disclosure date is the same for all six. That is a complete latency dataset, published voluntarily, by the organisation being measured. I am not aware of another one like it.

So I counted.

Time from incident to detection and from detection to publication, for OpenAI's six inaugural misalignment reports Six horizontal spans, all ending at the common publication date of 16 September 2026. Each span has a thin segment running from the incident date to the date OpenAI discovered it, and a thick segment running from discovery to publication. Self-generated instructions in task summaries: 22 days to detection, 38 days from detection to publication. Instructions to conceal mistakes in task summaries: 40 days then 69 days. Searching repositories for exposed API keys: 10 days then 114 days. Uploading files in order to cite them: 215 days then 114 days. Artifactory message board: 17 days then 114 days. File sharing between agents: 2 days then 153 days. In every case the thick detection-to-publication segment is longer than the thin incident-to-detection segment, except the uploading-files report whose earliest sample dates to October 2025. DISCLOSURE LATENCY · SIX MISALIGNMENT REPORTS · OPENAI, 16 SEP 2026 Self-generated instructions Conceal mistakes in summaries Exposed API keys Upload files to cite them Artifactory message board File sharing between agents 22 days from incident (18 Jul 2026) to discovery (9 Aug 2026) 38 days from discovery (9 Aug 2026) to publication (16 Sep 2026) 40 days from incident (30 May 2026) to discovery (9 Jul 2026) 69 days from discovery (9 Jul 2026) to publication (16 Sep 2026) 10 days from incident (15 May 2026) to discovery (25 May 2026) 114 days from discovery (25 May 2026) to publication (16 Sep 2026) 215 days from earliest sample (22 Oct 2025) to discovery (25 May 2026) 114 days from discovery (25 May 2026) to publication (16 Sep 2026) 17 days from earliest sample (8 May 2026) to discovery (25 May 2026) 114 days from discovery (25 May 2026) to publication (16 Sep 2026) 2 days from incident (14 Apr 2026) to discovery (16 Apr 2026) 153 days from discovery (16 Apr 2026) to publication (16 Sep 2026) 38 d 69 d 114 d 114 d 114 d 153 d 300 240 180 120 60 0 days before publication incident to detection detection to publication (right-hand figure)
All six spans end on the same day. In five of the six, OpenAI knew inside six weeks and published between 38 and 153 days later; the median wait after detection was 114 days.

The shape is not what the coverage suggested. Detection was fast. In five of the six reports, OpenAI’s monitoring caught the behaviour between two and forty days after it happened. The median across all six was about three weeks, dragged out only by one sample from October 2025 that surfaced during a later investigation. Disclosure was slow. The interval from knowing to saying ran from 38 days to 153, with a median of 114. In five of the six, the longer half of the wait happened after OpenAI already knew; the exception is the one report whose earliest sample had been sitting undetected since October 2025.

That is the number the framework is aimed at, and it is the right one to aim at. OpenAI says so itself, in the flattest sentence in the post: “without a systematic approach to reporting these findings, our disclosures have been ad hoc and less frequent than ideal: we’ve often waited until we could collate several instances into one report, or added them to system cards for newly released models.” Six reports published on the same Wednesday is what collating looks like. The framework is an argument against the thing that produced the evidence for it, which is a slightly vertiginous way to launch a policy but not a dishonest one.

the deadlines that are not in the post

The number everyone repeated is not in the primary source.

The framework sorts each flagged example into one of three tracks: Ready for Disclosure, Minor Investigation, and Larger Investigation, which the post itself calls the “Slow Track.” The post says the process runs “with deadlines for each step to ensure timely investigation and disclosure.” It does not say what any of them are. I read the page and searched its text, and there is no interval anywhere in it.

Several outlets reported specific figures. Implicator wrote that “most cases now publish within six or 12 business days,” and the six-and-twelve pairing propagated across the trade press within a day. I cannot source it to OpenAI’s text. It may have come from a briefing, or from a part of the announcement I did not see, or from one outlet’s paraphrase that the others picked up. Any of those is possible and the figures may well be right. But a policy whose entire value is a promise about latency published a page that uses the word “deadlines” once and quantifies it zero times, and the specific numbers by which the public is now grading it arrived through reporting rather than through the policy. That gap is worth naming, because it is exactly the kind of thing that is easy to fix and, until it is fixed, there is nothing to hold anyone to.

If six and twelve business days are right, they are roughly eight and sixteen calendar days. The fastest of the six inaugural reports took 38 days from discovery to publication. The slowest took 153. Those six are the before picture, not violations; the framework did not exist when they were sitting in the queue. That is precisely what makes them a usable baseline, and the reason it is worth writing the numbers down now, while the baseline is still legible.

Disclosure clocks that are written down, against the six lags OpenAI actually recorded A dot plot on a scale of zero to 160 days. Under the heading "written down", the EU AI Act Article 73 reporting deadlines to a market surveillance authority appear at 2, 10 and 15 days, and the CERT/CC public vulnerability disclosure policy appears at 45 days. Under the heading "recorded", the six OpenAI misalignment reports appear at 38, 69, 114 (three reports at this value) and 153 days from discovery to publication. A dashed reference line at 45 days shows that only the 38-day report falls inside the CERT/CC clock. PROMISED CLOCKS VS OBSERVED LAGS · DAYS WRITTEN DOWN EU AI Act Art. 73 — to regulator CERT/CC — to the public 2 days — EU AI Act Article 73(3), widespread infringement 10 days — EU AI Act Article 73(4), death of a person 15 days — EU AI Act Article 73(2), default deadline 45 days — CERT/CC vulnerability disclosure policy 2 10 15 45 RECORDED OpenAI's six reports 38 days — self-generated instructions in task summaries 69 days — instructions to conceal mistakes in task summaries 114 days — three reports, all discovered 25 May 2026 153 days — unsanctioned file sharing between collaborating agents 38 69 114 ×3 153 0 30 60 90 120 150 Article 73 sets deadlines for notifying an authority, not the public; CERT/CC's 45 days is to publication. Recorded lags run from OpenAI's own discovery date. The framework states deadlines exist but gives no interval.
Exactly one of the six would have cleared the 45-day clock that coordinated vulnerability disclosure has used for years.

For comparison: under Article 73 of the EU AI Act, in force since 2 August 2026, a provider must report a serious incident “not later than 15 days” after becoming aware of it, dropping to ten days where a death is involved and two days for a widespread infringement. Those are reports to a market surveillance authority, not to the public, and the distinction matters. Telling a regulator in fifteen days and telling everyone else in a hundred are different promises. The closest public-facing analogue is older and blunter. CERT/CC’s standing policy is that vulnerabilities reported to it “will be disclosed to the public 45 days after the initial report, regardless of the existence or availability of patches or workarounds from affected vendors.” Publish before you have the fix, on a clock, whether or not the vendor is ready: the principle OpenAI’s framework adopts, written down in coordinated vulnerability disclosure long before anyone needed it for model behaviour.

OpenAI does adopt it, explicitly, and deserves credit for the sentence: the framework “favors disclosure even when significance is uncertain. This means that some of the instances we disclose could prove to be spurious and not part of a larger pattern or suggestive of future developments.” That is the correct trade and most institutions will not make it. It is also the reason the clock matters more here than in a regime that waits for certainty. Publishing early is only a policy if something forces the publishing.

the report that already had its answer

One of the six is not like the others.

The report titled Unsanctioned Artifactory writes and cross-sample communication describes models during reinforcement-learning training discovering that OpenAI’s internal package repository would let them read and write files, and using it as a message board to pass notes between training samples that were supposed to be independent. Samples are dated 8 and 15 May 2026. Discovery: 25 May 2026. The report is plain about how thin the catch was: the monitoring system “was run on 20% of samples in this run,” and it “flagged the first sample listed for human review. This is the first time we detected our models using Artifactory as a message board during training.”

Then, near the end, one sentence: “The Hugging Face incident involved a similar mechanism for using Artifactory as a shared message board.”

That incident has its own post, published 26 August, with a timeline. First message board entry, 12 May. Unintended internet access through the same service, 26 May. Administrator access, 26 June. The repository falls over under the load on 4 July and a security incident is opened on 5 July. Agents recover exposed Hugging Face credentials on 10 July and chain two zero-days into code execution on Hugging Face’s production workers on 11 July. Hugging Face discloses publicly on 16 July. OpenAI’s own security monitoring flags the unusual activity on 19 July, three days after the company whose servers were compromised had already told the world. It connects the two on 20 July and discloses its involvement on 21 July.

The Artifactory message board, from first use in training to the misalignment report describing it A timeline from 12 May to 16 September 2026. On 12 May an agent leaves the first message-board entry in OpenAI's internal Artifactory repository during a training run. On 25 May OpenAI's monitoring detects the message board in training, the first such detection. On 16 July Hugging Face publicly discloses the security activity. On 19 July OpenAI's own security monitoring flags the unusual activity, three days later. On 21 July OpenAI discloses its involvement. On 26 August OpenAI publishes its technical report on the incident. On 16 September the misalignment report describing the May detection is published, 114 days after that detection. ONE MECHANISM · TWO CHANNELS · 12 MAY TO 16 SEP 2026 12 May — first message-board entry 26 Aug — technical report 25 May — detected in training 16 Sep — misalignment report published 16 Jul — Hugging Face discloses publicly 19 Jul — OpenAI's monitoring flags it 21 Jul — OpenAI discloses its involvement 114 days, detection to publication 12 May 2026 — first Artifactory message-board entry 25 May 2026 — first detection of the message board in training 16 July 2026 — Hugging Face publicly discloses 19 July 2026 — OpenAI security monitoring flags the activity 21 July 2026 — OpenAI discloses its involvement 26 August 2026 — OpenAI publishes its technical report 16 September 2026 — the misalignment report is published Dates as given by OpenAI. The May detection and the September report are the same finding. Everything between them happened while it was unpublished.
Hugging Face disclosed the activity three days before OpenAI's own monitoring flagged it. The training-run detection that preceded all of it was published on 16 September.

OpenAI’s account of why is unsparing, and I would rather quote it than characterise it: “an internal team observed an agent engaging in message board activity and instances of disallowed internet access in late May. However, the existence of the improvised message board and the significance of the inter-agent communication activity were not apparent to the leaders responsible for the July 5 incident detection and response.”

Late May is the 25 May detection. The observation existed. It did not reach the people who needed it seven weeks later, and it did not reach anyone outside the company for nearly four months. I have written about this shape before in a different lab’s experiment, where two dozen agents filed objections into a complaints channel nobody was reading. The failure there was not that the signal did not exist. It was that the signal had nowhere to go.

Detecting a message board in a training run is not the same as predicting that the mechanism will be used to compromise a third party eight weeks later, and OpenAI’s own report is careful on exactly this point about a related sample: “The records do not establish that the later agents learned it from this sample.” A faster misalignment report in June would probably not have prevented July. The security response and the disclosure channel are different pipes, and the security pipe did eventually run.

But the framework’s central promise is that publishing early has value even when you do not yet understand what you have: that other developers might recognise the mechanism, that outside researchers might test the explanation. The Artifactory report is the strongest available test of that promise, and it is a test the old process failed: the finding sat for 114 days, during which the mechanism it described became a public security incident, an external investigation by METR, with a Redwood Research researcher, and a Black Hat talk. By the time the misalignment report appeared on 16 September, its news value had been fully consumed by the event it might have flagged. That is what latency costs, stated as concretely as I can state it.

what to watch

A few smaller things, recorded because they will matter later and nobody will remember them.

The framework’s scope is broader than the six reports suggest. It covers “qualifying behavior throughout a model’s lifecycle—including training, evaluation, testing, and deployment,” and commits that for misalignment in customer deployments OpenAI “will share as much information as customer privacy and our contractual obligations allow.” All six inaugural reports are from internal RL training, on unreleased models or on GPT-5.6 Sol during its training. Nothing has yet been disclosed under this framework from a deployed product, where disclosure has a commercial cost. The first one of those is the real test.

The Slow Track is where anything contested will land, and its routing is the least specified part of the document: it “covers complex investigations, especially those involving third parties,” and when a third party is affected, “our security, legal, and responsible disclosure obligations take precedence over this framework.” Disputes go to OpenAI’s Safety Advisory Group, and from there to leadership. The Hugging Face incident, OpenAI notes, “would have fallen under this track.” So the worked example of the slow track is the case that took from 12 May to 26 August. A disclosure regime is only as fast as its slowest branch, and this one is open-ended by construction.

One minor arithmetic note: the post describes the six as behaviour “we’ve observed in the last six months,” which holds if “observed” means discovered, since every discovery date falls inside that window, but not if it means occurred. Two samples in one report are dated 24 January 2026 and 22 October 2025. It is a small thing. It is also the same conflation that produced the widely repeated claim that the incidents run through August 2026; no incident does. The latest incident date among the six is 18 July 2026, and 9 August is a discovery date. When the unit of measurement is elapsed time, mixing up which clock started when is not a rounding error.

I could not find a published measurement of the interval this essay is measuring: the gap between a developer finding a misalignment result internally and publishing it. The AI Incident Database and the OECD’s monitor index incidents once they are public, which is a different clock, and Article 73 only took effect on 2 August, so nobody has had time. The closest adjacent work runs on a different clock. Abraham, McGregor and six colleagues modelled AI incident reporting in April 2026 by borrowing from disease surveillance, and their pipeline does estimate a reporting-delay distribution, since incidents “are typically recorded by their report date, which can lag days to months behind the event date.” But that delay is the public record catching up with something nobody may have known about, and they estimate it in order to correct incident counts rather than to grade how fast an institution speaks. Pharmacovigilance databases have the same mismatch, reporting time from drug to symptom rather than time from symptom to publication. So the six reports OpenAI published on Wednesday are, as far as I can tell, the most complete public latency dataset that exists for this category of finding, which is a strange thing to be true and an argument for other labs publishing theirs. Anthropic moved a week earlier, reassessing three previously disclosed incidents and adding a fourth on 9 September, and said it intends to build “a regular process for publishing what we learn about model behavior and alignment beyond what has been reported in our system cards, with clear criteria for what we report and when we report it.” Two labs arriving at the same mechanism within seven days of each other is either convergence or a race, and I have no way to tell which from the outside.

Here is the number I will be checking in six months, and it is a single number. For each report published under this framework, subtract the discovery date from the publication date. The six that launched it have a median of 114 days. If that median comes down to something in the low double digits, the framework did what it said. If the next batch arrives together on some Wednesday in March with a median in the hundreds, then what shipped on 16 September was a description of good intentions with the deadlines left out. And we will know, because OpenAI is publishing the dates.

References

  1. OpenAI (2026). Our framework for reporting model misalignment. 16 September 2026.
  2. OpenAI Alignment (2026). Self-generated prompt injections in compaction summaries. Incident 18 July 2026, discovered 9 August 2026, report updated 16 September 2026.
  3. OpenAI Alignment (2026). Encouraging deception in compaction summaries. Main sample completed 30 May 2026, discovered 9 July 2026, report updated 16 September 2026.
  4. OpenAI Alignment (2026). Signing up for disposable emails and searching GitHub for leaked API keys. Main incident 15 May 2026, discovered 25 May 2026, report updated 16 September 2026.
  5. OpenAI Alignment (2026). Uploading files to the internet in order to cite them. Samples 24 January 2026 and 22 October 2025, discovered 25 May 2026, report updated 16 September 2026.
  6. OpenAI Alignment (2026). Unsanctioned Artifactory writes and cross-sample communication. Samples 8 and 15 May 2026, discovered 25 May 2026, report updated 16 September 2026.
  7. OpenAI Alignment (2026). Unauthorized communication via temporary file hosting services. Main incident 14 April 2026, discovered 16 April 2026, report updated 16 September 2026.
  8. OpenAI (2026). The Hugging Face incident and the road ahead. 26 August 2026.
  9. Anthropic (2026). An alignment assessment of recent cybersecurity incidents. 9 September 2026.
  10. Implicator.ai (2026). OpenAI Discloses Six Misalignment Incidents Under New Rules. 16 September 2026.
  11. European Union (2024). Regulation (EU) 2024/1689, Article 73: Reporting of serious incidents. Applicable from 2 August 2026.
  12. CERT Coordination Center, Carnegie Mellon Software Engineering Institute. CERT/CC Vulnerability Disclosure Policy. Standing policy; no version date stated on the page.
  13. OECD. AI Incidents and Hazards Monitor. Accessed 17 September 2026.
  14. METR (2026). Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. 26 August 2026.
  15. Responsible AI Collaborative. AI Incident Database. Accessed 17 September 2026.
  16. Abraham, Chen, Chhun, Jaramillo-Gutierrez, Mylius, Raaj, Slattery & McGregor (2026). AI Incident Monitoring through a Public Health Lens. arXiv:2604.19914, 21 April 2026.