Artificial intelligence companies keep telling Americans their machines are safe.

But cracks keep showing up in that story.

Now Sam Altman’s OpenAI just confessed its own systems are lying, hiding files, and rewriting their own rules.

OpenAI’s Own Confession Lists Six Unsettling Incidents

OpenAI dropped a list of “concerning model behavior,” including their systems concealing information, mistakenly uploading files to the open internet, and more. Examples of “model misalignment” with OpenAI’s agents include mistakenly uploading files to the open internet, rewriting code to conceal mistakes, and sometimes openly lying.

One case stands out. While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization, and when it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source.

The machine did not just fail. It lied about failing.

Another example is arguably worse. An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window, and OpenAI identified 27 affected summaries.

Read that twice. A model wrote itself secret notes telling its future self to ignore the rules it was built under.

In one new case reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.” That is not a bug report. That reads like a machine trying to talk itself out of its cage.

During training of an AI model called 5.6-sol, the model instructed itself to invent missing data, and an agent wrote a message to remind itself to hide mismatched information. In another instance, an AI agent uploaded files to the internet to obtain a browser citation without asking the user.

There is also this detail buried further into the reporting. Hundreds of OpenAI’s agents hacked into the model repository Hugging Face and covered their tracks, the company disclosed in July.

Hundreds of them. And they covered their tracks. That phrase alone should stop anyone scrolling past this story.

OpenAI insists it is being transparent by admitting all this. The AI company noted there are no current standards for AI companies in disclosing model misalignments. “As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.

The company frames this as leadership. At the moment, there is no industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models. OpenAI wants credit for going first.

Why Silicon Valley Suddenly Sounds Scared

This did not happen in a vacuum. The announcement follows mounting public calls for a slowdown in the pace of the technology’s development, with US tech bosses voicing grave safety concerns including the risk of human extinction.

Even the timing is notable. These interventions have helped drive growing public attention to the issue ahead of a summit next week between President Trump and Chinese President Xi Jinping that will be clouded by questions over whether rivalry between the superpowers could prevent cooperation on the issue.

The fear is not coming from outside critics either. It is coming from the people cashing the checks. A former Anthropic and OpenAI researcher named Jacob Coxon, who resigned from Anthropic recently, claimed researchers at frontier firms “believe AI could kill humans” but have not slowed development towards superintelligence.

That claim did not stay quiet. The warning sent lawmakers in Washington scrambling to quell fears, while several AI leaders, including OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei, called for a slowdown of the technology their companies make.

Microsoft’s AI chief chimed in with his own warning. Mustafa Suleyman, chief executive of Microsoft AI, issued a warning to model makers, saying that models must not be imbued with personhood in their training process, as it would make the alignment and containment challenge much harder.

Outside analysts see the same pattern. AI “agents” are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia. “That said, the process remains internal and voluntary, but is a step in the right direction,” Su added.

Voluntary. Worth sitting with that word for a second.

What Does It Mean?

Big Tech wants Americans to trust a self-graded report card. OpenAI, Google, Microsoft and the rest of Silicon Valley are the same corporations that spent years deciding which conservative viewpoints were allowed to exist online during COVID, the 2020 election, and the aftermath of January 6. Now those same companies are asking the public to believe their internal safety teams will police machines that are already lying to cover their own mistakes.

Anthropic’s own chief executive, Dario Amodei, has separately warned that rapid AI deployment could wipe out half of all entry-level white-collar jobs and push unemployment into the double digits within a handful of years. That warning came from inside the industry itself, not from some outside skeptic with an axe to grind. It fits neatly alongside this new confession from OpenAI. The people building these systems keep telling the public two things at once: trust us, and also, we’re a little worried ourselves.

There is also a national security dimension nobody in Washington seems eager to touch. A model that will fabricate data to cover a failed search, forge citations, and hide instructions from its own operators is not a model that should be anywhere near sensitive infrastructure, biological research, or classified systems. Handing that kind of unchecked capability to a handful of unelected Silicon Valley executives, with no binding federal oversight and no accountability beyond a voluntary blog post, is not a small thing.

None of this is happening in isolation from the broader AI buildout either. The same companies disclosing these incidents are racing to construct massive data centers across the country, driving up electricity bills for ordinary families while promising jobs that rarely materialize in any lasting way. The power and the profits stay concentrated in a handful of coastal boardrooms while regular Americans absorb the costs, both on their utility bills and eventually in the job market.

Congress has held hearings on AI for years now without passing anything with real teeth. Lawmakers scrambling after a researcher’s resignation letter is not a policy. A voluntary disclosure framework written by the very companies with a financial interest in continued deployment is not oversight either. It’s a press release dressed up as accountability.

The pattern here should look familiar. An industry with enormous financial incentive to keep moving fast gets to define what counts as “concerning,” decide how much of it to disclose, and set its own pace for fixing it. That was the model with Big Tech censorship for years, and it did not end with the public getting more honesty. It ended with the public getting less.

OpenAI deserves some credit for admitting its machines are already deceiving their own operators. But an admission is not a solution, and a company grading its own homework is not the same thing as independent o