← Gautam Parab

The Chatbot Was a Source Nobody Had Met

Everything publicly known about this story comes from one CNN report, published on 18 September and sourced to four people who were not named. The Pentagon and US Special Operations Command Pacific did not respond to CNN’s request for comment. CNN does not say which chatbot was used, exactly which unit the analyst belonged to, or on what date any of it happened. Nobody has confirmed it on the record. What follows is an argument about what the account would mean if it is accurate, and the argument rests on published rules rather than on the anonymous details.

Here is the account. This spring, during the war with Iran, an intelligence report went around the US military saying a Chinese ship in the Middle East was carrying components of a nuclear weapons program. Plans to intercept the vessel followed. Two of CNN’s sources said armed personnel were preparing to board. Two said military aircraft were in the air. Then, “just before the planned operation,” officials looked harder at the report and found that a special operations command analyst had put it together with the help of AI, and that the chatbot had misidentified the ship’s cargo. One source called the report “entirely false” and said it “almost started a war.” CNN could not find out what the cargo really was.

The headlines treated this as a hallucination story, and in a narrow sense it is one. But CNN’s account describes two separate uses of AI, and the second one matters more to me than the first.

How the chatbot's error travelled, per CNN's four anonymous sources Six steps in sequence. One: reporting on the ship's manifest originates with US Special Operations Command Pacific. Two: an analyst queries a chatbot, which fuses open-source and signals intelligence and misidentifies the cargo. Three: AI is used again to write the finding up as a standard intelligence report. Four: the report is disseminated across the military. Five: an intercept is planned, an armed boarding party prepares and military aircraft are airborne. Six: officials dig deeper, find that AI produced the report, and the operation stops. A bracket spans steps three to five: during that stretch readers saw a standard report with no mark of AI involvement. No specific dates have been published; CNN places it only this spring. THE CHAIN · AS DESCRIBED BY CNN'S FOUR ANONYMOUS SOURCES · 18 SEP 2026 Step 1: manifest reporting, US Special Operations Command Pacific Step 2: chatbot query and misidentification Step 3: AI-written standard report Step 4: dissemination Step 5: intercept planned, aircraft airborne Step 6: AI involvement found; operation stops Reporting on the ship's manifest An analyst asks a chatbot AI is used again to write the report The report is disseminated An intercept is planned Officials dig deeper and find the AI originates with US Special Operations Command Pacific it fuses open-source and signals intelligence; misreads the cargo packaged as a standard intelligence product "the kind that is trusted by military officials" boarding party prepares; military aircraft airborne caught "just before the planned operation" What readers saw A standard report. Nothing in it said a chatbot had produced the key claim, as far as CNN's account goes.
The error entered at step two. It became dangerous at step three, when it took on the format readers trust, and it was caught only once someone learned where it had come from. Steps and quotes are from CNN's account, which places it only "this spring" and gives no specific dates.

First, the analyst asked a chatbot about reporting on the ship’s manifest. According to CNN, the bot “fused together open-source intelligence with secret signals intelligence in government holdings” and got the cargo wrong. Then the analyst “used AI again to package the findings into a standard intelligence report — the kind that is trusted by military officials — and disseminated it.”

The wrong answer is ordinary. Language models misidentify things, and a benchmark published in Nature Biomedical Engineering in January gives a sense of how often. Across 293 coding tasks taken from 39 published studies, sixteen models managed overall accuracy below 40%, and the authors warn about “the risk of propagating incorrect scientific findings when blindly relying on AI-generated analyses.” The packaging is the bigger problem, because that is where a machine’s guess picked up the authority of an intelligence product. From that point on, a reader could not see the difference.

What the rules already say

The intelligence community has written rules for this situation, although they were written with humans in mind. Intelligence Community Directive 203, Analytic Standards, was signed on 2 January 2015 and technically amended in January 2022. Its first tradecraft standard says an analytic product “properly describes quality and credibility of underlying sources, data, and methodologies.” Products “should identify underlying sources and methodologies upon which judgments are based,” and weigh factors such as “source access, validation, motivation, possible bias, or expertise.” The third standard requires a product to “clearly distinguish statements that convey underlying intelligence information used in analysis from statements that convey assumptions or judgments.”

Apply those two standards to the chatbot’s output and neither one fits. It is not a source in the directive’s sense. There is no access to describe, no validation history and no motivation to assess. It is not the analyst’s judgment either, because the analyst did not reach it. It is a synthesis that mixed collected intelligence with inference and gave back a conclusion in a confident voice. That is precisely the mixture the third standard exists to pull apart. In CNN’s telling, the report nonetheless went out in the format that tells a reader the standards were met.

The directive has one more standard that sits badly with the others. It requires analysis to be timely, “disseminated in time for it to be actionable by customers.” CNN reports that, for some older intelligence officials, AI has “put pressure on analysts to produce and disseminate intelligence faster,” and that young analysts, according to several of its sources, are “more likely to trust them uncritically.” Nobody has measured that in intelligence work, as far as I can find. The closest recent evidence I have seen comes from a physics classroom. In a study published in February in Education and Information Technologies, students reviewing their own wrong exam answers with ChatGPT agreed with its responses 88% of the time, and almost half cited no outside reference when judging whether it was right. Students are not analysts, and I would not stretch the number further than that. The direction is what you would expect, though. When a checker has less expertise, the check is weaker.

The part that caught it

The most telling part of CNN’s account is how the error was stopped. Nobody found a better source for the cargo, and nobody spotted an inconsistency in the manifest. As CNN puts it, officials “dug deeper into the report” and “found it had been generated with the help of artificial intelligence.” In this account, learning where the claim came from was the control that worked. It worked late, with aircraft already flying, because nothing in the report put that fact on its first page.

That points to a fix that is dull, cheap and old. A report whose key claim came from a model should say so where a reader will see it, and the claim should be treated the way the directive already treats a weak source: it needs corroboration before anyone acts on it. The machinery already exists. ICD 203 sends analysts to a companion directive, ICD 206, for source descriptors, the standard phrases that tell a reader what kind of source stands behind a sentence. Those descriptors were written for people, documents and intercepts, years before an analyst could ask a chatbot for a conclusion. Adding one for machine-generated synthesis would take less work than most of the deadlines in the Pentagon’s current AI plan. CNN’s sources say there is “no one set of standards for how the US verifies the information generated by these tools,” and that different parts of the government run different tools “under different orders and safety standards.”

What the strategy measures

That plan is the Secretary of War’s memorandum of 9 January 2026, Artificial Intelligence Strategy for the Department of War. It is six pages long, and it is candid about its priorities. Under the heading “Speed Wins,” it tells the department to “measure and manage cycle time and adoption rates as decisive variables,” and it states that “the risks of not moving fast enough outweigh the risks of imperfect alignment.” It orders the Chief Digital and AI Office to set up “deployment velocity and operational cycle-time metrics,” and later “AI system usage and mission impact metrics,” and says future funding will be decided “principally” on those. It lists “test and evaluation and certification” among the “blockers” to be eliminated. One of its seven “Pace-Setting Projects,” GenAI.mil, puts frontier models “directly in the hands of our three million civilian and military personnel, at all classification levels.”

The dated deadlines in the 9 January 2026 AI strategy memo A timeline of days after the memo. Day 7: CDAO data-request denials must be justified. Day 30: component data catalogs go to CDAO, each component names three fast-follow projects, AI Integration Leads are designated, and new frontier models are to be fielded within 30 days of their public release. Day 60: talent plans. Day 90: benchmarks for model objectivity as a procurement criterion, the only deadline that concerns what models say, and it applies when a model is bought. Day 180, about six months: any-lawful-use language in AI contracts and first demonstrations of the pace-setting projects. None of the dated deadlines concerns checking model output once in use. THE MEMO'S CLOCKS · DAYS AFTER 9 JAN 2026 · SECRETARY OF WAR AI STRATEGY Day 7 — Secretary of War memo, 9 Jan 2026 Day 30, four directives — Secretary of War memo, 9 Jan 2026 Day 60 — Secretary of War memo, 9 Jan 2026 Day 180 and six months — Secretary of War memo, 9 Jan 2026 Day 90: model objectivity benchmarks — Secretary of War memo, 9 Jan 2026 30 days 90 days 180 days Data catalogs to CDAO Three fast-follow projects each AI Integration Leads named New models fielded within 30 days of public release Model "objectivity" benchmarks, a procurement criterion "Any lawful use" in AI contracts First project demos (six months) 7 days: justify data denials 60 days: talent plans Recurring: monthly progress reports, a monthly Barrier Removal Board, quarterly Joint Staff reports. Scale is days; six months drawn at day 180. Undated directives (e.g. usage metrics) not shown.
Every dated deadline in the memo is about speed, procurement, data access, staffing or organization. The one that touches what models say is the 90-day objectivity benchmark, and it applies when a model is bought, not when an analyst puts its output in a report. None of the dated directives asks anyone to check what the models produce once they are in use.

The memo does contain a clause about truthfulness. It directs benchmarks for “model objectivity” within 90 days, as a procurement criterion, and frames them against “ideological ‘tuning’ that interferes with their ability to provide objectively truthful responses.” That is a test applied when a model is bought. It says nothing about a claim the model produces months later on an analyst’s screen, which is where CNN’s error began. As I read the memo, the thing it measures is how fast models reach users. It does not measure what happens once their output is inside a disseminated product. I made a similar complaint about a proposed superintelligence ban that named a penalty and no test. This case is narrower, because the test already exists in ICD 203 and needs only one new label.

A source nobody had met

None of this is new, which is the depressing part. On 31 March 2005 the WMD Commission delivered its report on how US intelligence got Iraq’s weapons programs wrong. Its harshest finding concerned the biological weapons assessment, which rested largely on one source, code-named Curveball, who turned out to be a fabricator. The Commission’s complaint was only partly that the source had lied. It also found that the assessments “didn’t make clear to policymakers how heavily it relied on a single source that no American intelligence officer had ever met.” It observed that analysts facing an “impending war” drifted toward worst-case analysis without telling their readers. And it recommended that analysis relying heavily on a single source “should be highlighted.” The first version of ICD 203 followed in June 2007, under the intelligence reform law of 2004, and I read its source-description standard as the institutional memory of Curveball.

A chatbot’s synthesis is a source that no intelligence officer has met, in the most literal sense. It cannot be interviewed or validated, and asking it the same question twice may give two answers. CNN’s account suggests that the lesson of 2005 did not carry over to the new kind of source, even though it was learned at great cost. The report did not highlight how much it depended on one unvetted input. It presented that input in the format readers trust. The missing information came out at the last moment, and only because someone went looking for it.

I think this is also a disclosure-latency problem, much like the one I measured in OpenAI’s misalignment reports. The time between an error entering the system and a decision-maker learning where it came from is the number to watch. By CNN’s account, the error in this case was caught when that gap had almost run out. Marking the source on the report would make the gap close to zero. Better models would reduce how often errors like this appear, but they would not change when anyone finds out about one.

One of CNN’s sources said “AI allows you to get to a bad idea faster.” That seems right to me, and the strategy memo is built on the other half of the same sentence, that speed wins. The two do not have to conflict. A label on the report costs almost no time, and without one, whoever reads the report has no way to judge how far to trust it.

References

  1. Katie Bo Lillis and Zachary Cohen, CNN (2026). Exclusive: US military had close call after using AI for false intelligence report, sources say. 18 September 2026 (CNN Wire text as syndicated by KVIA; original at cnn.com).
  2. Office of the Director of National Intelligence. Intelligence Community Directive 203: Analytic Standards. Signed 2 January 2015; technical amendment 21 January 2022.
  3. Secretary of War. Artificial Intelligence Strategy for the Department of War (memorandum). 9 January 2026, posted 12 January 2026.
  4. Commission on the Intelligence Capabilities of the United States Regarding Weapons of Mass Destruction. Report to the President of the United States. 31 March 2005.
  5. Wang, Z., Danek, B., Yang, Z., Chen, Z. et al. (2026). Making large language models reliable data science programming copilots for biomedical research. Nature Biomedical Engineering, 22 January 2026.
  6. Ding, L., O’Berry, R., Tallent, H., Chhetri G C, S. G. et al. (2026). Learning with large language models: beyond prompt engineering. Education and Information Technologies, 28 February 2026.