A new letter from insiders about the potential dangers of AI is refreshingly specific. The writers are concerned, according to the Wall Street Journal, about preserving the ability to monitor the models’ chains of thought. But the authors are not technically insiders now, because they were recently fired.
Last week, OpenAI announced that it had “partnered with three individuals for violating our policies on access and handling of sensitive corporate information.” They allegedly passed this information on to an AI security organization.
These contributors were Tomek Korbak, Mikita Balesni, and Jasmine Wang, and they are clearly of the school of thought that AI represents an existential risk. Before their firing, Balesni already posted on X that he thinks AI is “10% likely to kill all humans.” Wang posted, “It’s hard to overestimate how dangerous speed is in the direction of the RSI.” (RSI, or recursive self-improvement, means that AI models improve themselves).
Korbak, Balesni, and Wang’s new letter is addressed to OpenAI’s board of directors and the company’s “security committees,” writes the WSJ. It has not been published publicly, and the WSJ has only published excerpts.
“As an industry, we don’t yet know how to develop and deploy safe models that we can’t monitor[…] OpenAI and other border companies should not continue with developments that further decrease surveillance, the letter says.
OpenAI’s latest flagship model, GPT-6 Astra, has triggered two types of concerns vis-a-vis Chain of thought (CoT), and here I will describe them using metaphorical language with the risk of anthropomorphizing models: A) they can learn to leave the surveillance system, and B) OpenAI supposedly could steer things in a direction that can make it less useful.
More specifically, according to the model’s system map, “Astra class models could evade our CoT monitors under adversarial conditions.” And meanwhile, reasoning with Astra also implies a more opaque form of “thinking” (and look: none of this is actually thinkbut that’s the term we have at the moment) called “recurring depth.” Iterative depth initially cycles a request through the model several times to process it instead of generating something – an event that happens in the unrecognizable black box of the model, so there is nothing to monitor.
Interestingly, when the WSJ reached out to OpenAI for comment about the letter, an OpenAI representative said the fires “do not raise or comment on security concerns,” and presented an internal memo about the letter in which the company said it “strongly agrees” with the authors’ recommendations.
