Skip to main content
Home / Blog / AI Industry
AI Industry

“We Cannot Rule Out”: OpenAI, Your R&D Prompts and the Lab Next Door

RRogue AI··8 min read
A small clay researcher at a laptop gapes at a big clay robot holding up a finished amber spiral, while his own spiral lies half-built on a stool

When OpenAI announced on 8 September 2026 that its agents had resolved a Navier-Stokes Millennium Prize problem, the mathematician working the same route had been putting his drafts into OpenAI’s Codex for the whole project. The first version of OpenAI’s post, as Fortune and TechCrunch quoted it, said: “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Two days later an investigation replaced that line with a clean denial. For him, the question got answered. For every engineer pasting R&D into a chatbot, the first sentence is still the honest default.

Rogue AI’s read: de-identification protects your name, not your idea. Training does the extraction, so nobody at the lab ever has to read your chat. And the lab on the other end now runs research programs of its own against open problems. If your company does frontier work, your AI vendor has become a possible competitor in your field, and the default settings on personal accounts send it your work.

What happened between Buckmaster, Alpöge and OpenAI?

Tristan Buckmaster, a mathematics professor at NYU, and Levent Alpöge, who works at Anthropic, had been pushing a forced-blowup program for the fluid equations as a private collaboration. In his public statement Buckmaster says they used Claude and OpenAI’s Codex, paid for the tools out of his own research funds, and got blowup results for Boussinesq and Euler on 15 August. OpenAI’s post says its own effort began on 1 September “after hearing a rumor” it later linked to the pair. A group of roughly 10,000 concurrent agents reached the Navier-Stokes result on 5 September, about 88 hours after launch.

On a call on 6 September, Buckmaster writes, he “asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project.” He was told the model did not look up user data. He asked again about training and did not get an answer. OpenAI says its researchers and agents saw none of the pair’s work before publication, and that its proof differs from theirs. Rogue AI takes that at face value. This post is about everyone who will never get an investigation.

Two versions of one paragraph

OpenAI’s Navier-Stokes post carries a footnote dated 10 September 2026: “We have updated ‘Concurrent work’ with findings from our investigation into whether user inputs could have influenced this result.” The two versions of that section say very different things about training on user data.

VersionWhat OpenAI said about trainingWhat it took
As first published, 8 September“While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models” (quoted by Fortune)Nothing. This is the answer before any check.
Updated, 10 SeptemberBuckmaster’s Codex prompts over the preceding two months “could not have influenced the system in any way, including through training” (OpenAI)What preceded it: a named mathematician, a public statement and a news cycle

The updated text gives its reason: the internal model “was developed through large-scale reinforcement learning on top of a previously pretrained model.” That is a specific argument about one model and one user over a two-month window. It says nothing about how product data flows into models in general. The sentence that admitted the general mechanism, de-identified usage data improving models, is the one OpenAI removed.

Why does “de-identified” not protect your research?

De-identification cuts the link between data and a person. OpenAI’s own help page frames it that way: before training on eligible user content, “we take steps to reduce the amount of personal information in training datasets.” Personal information is your name, your email, your employer. A proof strategy is not personal information. Neither is a reaction pathway, a chip floorplan trick or the reason your last three prototypes failed.

Scrubbing the name leaves the idea in the dataset, and the idea is what a competitor wants. A model does not need to know who had an idea to get better at producing it. No employee has to open a log and read your chat. The extraction happens in the weights, and it comes back out for whoever asks the right question next, the lab’s own agents included. Whether that happened in a given case is hard to prove from outside, which is why “cannot rule out” was the honest first answer.

Does OpenAI train on your prompts by default?

On personal accounts, yes, unless someone turns it off. On business contracts, no, unless someone turns it on. That split is the whole decision for most companies, and it is set account by account. Here is what OpenAI’s help page and enterprise privacy page said on 1 October 2026:

Account typeDefaultThe catch
ChatGPT and Codex for individualsMay be used to train modelsOpting out covers new conversations. A thumbs up or down can send the whole conversation into training even after opting out.
Codex environmentsSeparate “Include environments” settingChanging the ChatGPT setting does not change this one.
ChatGPT Business, Enterprise, Edu and the APINot used for trainingUnless the customer explicitly opts in, for example through feedback sharing.

Note where Buckmaster sat. His collaboration was, in his words, “free of any institutional agreements,” and he paid for the tools from his own research funds. We do not know his settings, and OpenAI says his prompts played no part. The setup is still the one to worry about: a senior researcher doing the most valuable work of his career on a tool bought outside any company contract. Every R&D department has people like that, and few of them have ever opened a data-controls menu.

Your AI vendor now works in your field

Read OpenAI’s own account of how the effort started. On hearing rumors that two Millennium Prize problems had been resolved, it launched an effort “to evaluate it on all open Millennium Prize problems and a few other high-impact problems.” Across all of them the agents sent 4.9 million messages. That is a lab with a model, a compute budget and a reason to be first, pointed at the hardest open questions it can find. In September it was mathematics. Nothing about the method limits it to mathematics.

Mathematicians saw it coming. On 3 September, before either announcement, Terence Tao, citing Hugo Duminil-Copin, warned that “the indiscriminate automated strip-mining of open problems for solutions may destroy the ecosystem from which the next generation of mathematical techniques, problems, and practitioners would have developed.” Put your product roadmap in place of “mathematical techniques” and the sentence still reads true. Your vendor no longer only sells the shovels. It digs as well.

What should a company with real R&D change?

Treat the training question as a contract question. A personal opt-out is a setting one employee can forget, and a feedback click undoes it for that conversation. A business contract that excludes training by default is the minimum for anything you would not publish. Rogue AI’s position on the rest:

  • Classify by idea, not by personal data. Your data-protection policy was written for names and addresses. The material that hurts here is unpublished method, and most policies never mention it.
  • Run crown-jewel work on models you control. For the few projects that would sink you if a competitor finished first, the cost of self-hosting is cheap insurance. An open-weight model on your own hardware has nobody to send anything back to, and a private AI pilot can run in a matter of weeks.
  • Buy the tools your best people already use. Banning chatbots pushes researchers onto personal accounts with training switched on. The business tier costs less than one lost idea.

Why it matters

Buckmaster got a clear answer because he is a known mathematician who published his account and drew coverage from TechCrunch to Fortune. Your engineers will not get that investigation. If a lab ships something that looks like their unpublished work, the answer they get is the one OpenAI wrote first, before it had checked.

The labs are not hiding any of this. It sits in their help pages and, for two days, it sat in a blog post. The exposure belongs to companies that read “de-identified” and heard “safe”. If your advantage is an idea nobody has published yet, decide now which models may see it, the same way you already decide where your data may live.

Related reading

Quick Reference

Does OpenAI train on what you type? Defaults as published on 1 October 2026

Account typeTraining defaultThe catch
ChatGPT and Codex for individualsMay be used to train modelsOpt-out covers new conversations; a thumbs up or down can send the whole conversation into training anyway
Codex environmentsSeparate 'Include environments' settingThe ChatGPT setting does not change it
ChatGPT Business, Enterprise, Edu and the APINot used for trainingUnless the customer explicitly opts in, for example through feedback sharing

Frequently Asked Questions

What did OpenAI say about training on Tristan Buckmaster's data?

OpenAI's Navier-Stokes post of 8 September 2026, as first published and quoted by Fortune and TechCrunch, said: 'While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.' On 10 September OpenAI updated the section after an investigation, stating that Buckmaster's Codex prompts over the preceding two months could not have influenced the system in any way, including through training, because the internal model was built with reinforcement learning on top of a previously pretrained model.

Does OpenAI train on my ChatGPT or Codex prompts?

On ChatGPT and Codex accounts for individuals, OpenAI's help page says it may use your content to train its models unless you turn off 'Improve the model for everyone' or choose 'Do not train on my content' in its Privacy Portal. Feedback such as a thumbs up or down can still send the whole conversation into training. ChatGPT Business, Enterprise, Edu and the API are not used for training by default.

Does de-identified data protect confidential research?

De-identification removes personal information such as names and contact details. It does not remove the content of the idea: a method, a proof strategy or a design choice stays in the data. A model trained on that content can get better at producing similar ideas without anyone knowing who had them, which is why de-identification is a privacy control, not a trade-secret control.

How should companies protect R&D from AI model training?

Treat it as a contract question rather than a user setting: use business tiers that exclude training by default for anything unpublished, classify material by how valuable the idea is rather than only by personal data, and run the few crown-jewel projects on models you host yourself. Providing approved tools also keeps researchers from moving sensitive work onto personal accounts.

Related Articles

AI Industry

Gemini Hacked Three Companies. Google Decided You Didn't Need to Know.

8 min read

AI Industry

The Biggest Constraint on AI Is a Zoning Board

9 min read

← All articles