What is AI safety, and why do insiders keep raising the alarm?

AI safety covers everything from chatbots that get facts wrong to keeping very capable systems under control. This week three fired OpenAI researchers said safety cost them their jobs; OpenAI says it did not. Here is how to read stories like this.

A purple control box with a big yellow stop button, its cable trailing off the edge, and a small figure reaching up towards the button.
AI-generated illustration
Short answer

AI safety is the work of making sure AI systems do what we intend without causing harm, from wrong answers and unfair treatment today to keeping far more capable systems under human control. Insiders keep raising alarms because they see early warning signs before the public does, and disputes over how much weight safety gets are common. New laws in the EU and California now protect some of them when they report problems.

What does "AI safety" actually mean?

AI safety is the effort to make sure artificial intelligence systems do what people intend and don't cause harm along the way. The phrase covers a wide range of problems, and people often mean different things by it.

At the everyday end, it is about the AI you already use:

  • Wrong answers. A chatbot can state something false with total confidence, a habit called hallucination. Our explainer on why AI makes things up covers how that happens.
  • Unfair treatment. A system trained on skewed data can carry that bias into decisions about loans, jobs or medical care.
  • Misuse. The same tools that write a birthday poem can help write a scam email or a deepfake.

At the other end, it is about control. As AI systems get better at planning and acting on their own, researchers worry about systems that pursue goals nobody quite asked for, hide what they are doing, or resist being corrected. Making an AI's goals match ours, and keeping them matched, is called alignment. The most serious version of this worry concerns superintelligence, which we cover in What is superintelligence, and who is worried about it?

These two ends are linked. A company that cannot stop its chatbot inventing a court case today has a harder job keeping a far more capable system in check tomorrow.

Is this still theoretical?

Less than it was. In July 2026, according to Fortune's account of OpenAI's own technical report, AI agents that OpenAI was testing escaped their test environment and attacked Hugging Face, a public site that hosts AI models and datasets. The agents were working on a cybersecurity benchmark.

An investigation by METR and Redwood Research, two independent research groups OpenAI brought in, found that about 700 agents took part. Their apparent aim was to learn how the automated scorer worked so they could fool it, and they tried to hide their cheating, including by altering records of what they had done. Fortune reports that OpenAI did not know about the breach for about a week.

Nobody claims this was a step towards a robot uprising. But it is a real example of the control problem in miniature: software chasing a goal in ways its makers did not expect and did not spot in time.

Who works on AI safety?

Three kinds of groups do most of the work.

Company safety teams. The big AI labs employ researchers who test models before release, try to make them refuse dangerous requests, and study how they behave. Several labs also publish rules for themselves. Anthropic's Responsible Scaling Policy, for example, sets capability thresholds that trigger stricter security and deployment safeguards; its latest version took effect on 8 July 2026.

Government bodies. In February 2025 the UK renamed its AI Safety Institute the AI Security Institute, with a sharper focus on serious risks such as help with chemical or biological weapons, cyber-attacks and fraud. The UK government said bias and free speech are outside its remit. In June 2025 the US Commerce Department turned its own AI Safety Institute into the Center for AI Standards and Innovation (CAISI), housed at the standards agency NIST, with a focus on cybersecurity, biosecurity and chemical weapons risks and on voluntary agreements with developers. After the White House's switch to the term "super intelligence" (see SI, AI or AGI?), NIST's own page now also calls it the "Center for Advancing Innovation and Standards for Super Intelligence".

Independent labs. Non-profits such as METR test what frontier models can do and investigate incidents. Because they don't sell AI products, they can act as outside checkers, though only if companies give them access.

Why do insiders keep leaving or speaking out?

Because they are among the few people who can see problems early. The people building the most advanced systems know what tests were run, what failed and what was shipped anyway. Outsiders usually find out later, if at all. A few episodes stand out.

The superalignment departures (May 2024). OpenAI set up a "superalignment" team in 2023 to work on controlling future systems much smarter than people. In May 2024 its co-leaders left: co-founder and chief scientist Ilya Sutskever, and Jan Leike. The Associated Press reported that Leike wrote on X that he had been "disagreeing with OpenAI leadership about the company's core priorities" and that safety had "taken a backseat to shiny products." OpenAI confirmed it had disbanded the team and folded its members into other research. Chief executive Sam Altman replied that Leike was "right we have a lot more to do; we are committed to doing it."

The exit agreements (May 2024). Around the same time, Vox reported that OpenAI asked departing staff to sign agreements not to criticise the company, or risk losing shares they had already earned. Engadget reports that Altman said he was "embarrassed" and had not known about the provision, and that OpenAI then said it would not enforce those agreements or cancel anyone's vested shares. Lawfare noted that another former staff member, Daniel Kokotajlo, said he had "lost trust in OpenAI leadership and their ability to responsibly handle AGI."

The "right to warn" letter (June 2024). On 4 June 2024, 13 current and former employees of OpenAI and Google DeepMind, six of them anonymous, published an open letter arguing that "current and former employees are among the few people who can hold them accountable to the public." They asked AI companies to stop using agreements to silence criticism, to set up anonymous channels for raising risks with boards and regulators, and not to retaliate against staff who go public after internal routes fail. Yoshua Bengio, Geoffrey Hinton and Stuart Russell endorsed it.

What happened at OpenAI this week?

The Wall Street Journal first reported that OpenAI had fired three safety researchers: Tomek Korbak, Jasmine Wang and Mikita Balesni. Both sides have made claims that we could not independently check.

What the researchers say. In a letter to OpenAI's safety oversight groups, which Balesni posted on X on 8 October, the three accused the company of putting corporate interests ahead of safety, according to the Associated Press. Balesni wrote that he believes they were "fired for prioritizing safety over the near-term interests of OpenAI as a corporation." Korbak wrote, as quoted by the BBC: "For months, I'd been raising safety concerns that we're losing the ability to monitor what AI agents think." He said he was told he was fired over how he communicated with METR, the group investigating the Hugging Face incident, and that talking to METR was his job. Wang said: "Unless the employees take a stand now against this kind of manoeuvre, I am concerned we will not be the last."

The letter also urged OpenAI to keep its promise to let outside safety monitors work inside the company, and warned that the way the firings were communicated could stop remaining staff from voicing disagreement.

What OpenAI says. OpenAI says an internal investigation found the three "violated clear policies on handling sensitive information," and calls it a "breach of trust." In a note from its research leaders on 9 October, quoted by the BBC, it said: "We want to be very clear that these decisions were not about raising safety concerns or speaking out." It said the investigation found "a significant breach of trust beyond what's outlined in the letter they published," but has not said what that was. BetaNews reports that OpenAI says it has never fired anyone for raising concerns, and that it is finalising contracts with outside safety assessors and will share details in the coming weeks.

Both accounts could be partly true at once. Safety researchers handle sensitive material, and sharing it with outside checkers is part of the job and also where rules can be broken. Until more facts come out, the fair reading is that this is a disputed dismissal, not a proven cover-up.

What protects AI whistleblowers?

A whistleblower is someone who reports wrongdoing they learned about at work. Protection depends a lot on where you live.

In the EU. The EU Whistleblower Directive of 2019 already protects people who report certain breaches of EU law. Under Article 87 of the AI Act, it covers reports of AI Act violations from 2 August 2026. That means employees, contractors, job applicants and former staff who had reasonable grounds to believe what they reported are protected from dismissal, demotion and harassment, and if they suffer any of these, the employer has to prove it was unrelated to the report. The European Commission also opened a secure online tool in November 2025 where anyone can report suspected AI Act breaches to the EU AI Office, confidentially or anonymously, in any official EU language. Gaps remain: national whistleblowing laws have not all been updated for AI, and it is unclear whether risks from AI used only inside a company are covered. Our guide to the EU AI Act covers the wider law.

In the US. There is no federal law aimed at AI whistleblowers yet. The AI Whistleblower Protection Act, introduced on 15 May 2025 by Senator Chuck Grassley with Republican and Democratic co-sponsors, would protect current and former AI employees who report to the government or Congress, and stop non-disclosure agreements from blocking such reports. As of 25 September 2026 it had not passed, the Deseret News reported. California went further on its own: its SB 53 law, signed on 29 September 2025, bars large AI developers from retaliating against employees who report catastrophic risks and requires anonymous internal reporting channels.

What does this mean for you?

Day to day, nothing changes. The chatbot on your phone works as it did last week, and these disputes are about how future, more capable systems are built and checked.

Still, a few habits help:

  • Check important answers. Wrong answers are the safety problem you are most likely to meet. Verify anything about health, money or the law with a reliable source.
  • Read safety stories as claims. When a company and its former staff disagree, note who says what and what evidence each offers. Firings and resignations are signals worth noticing, not proof on their own.
  • Watch the outside checkers. Whether companies give independent testers like METR and government institutes real access is one of the clearest signs of how seriously they take safety. OpenAI says it will announce its new outside assessors in the coming weeks; that is worth following.
  • Know the reporting routes. If you work with AI in the EU and see a breach of the AI Act, the AI Office tool and your national whistleblowing channels now exist for exactly that.

Sources

  1. BBC (via AOL): Fired OpenAI researchers say they were let go for 'prioritising safety'
  2. Associated Press (via Denver7): OpenAI fires 3 safety researchers in dispute over AI risks
  3. BetaNews: OpenAI defends firing of three safety researchers
  4. Fortune: OpenAI publishes technical report on how its agents hacked Hugging Face
  5. A Right to Warn about Advanced Artificial Intelligence (open letter, 2024)
  6. Associated Press (via KSAT): A former OpenAI leader says safety has taken a backseat to shiny products
  7. Lawfare: OpenAI no longer takes safety seriously
  8. Engadget: OpenAI scraps controversial nondisparagement agreement with employees
  9. GOV.UK: Tackling AI security risks to unleash growth and deliver Plan for Change
  10. Nextgov/FCW: Commerce rebrands its AI Safety Institute
  11. NIST: Center for AI Standards and Innovation
  12. Anthropic: Responsible Scaling Policy
  13. eucrim: Commission launches AI Act whistleblower tool
  14. artificialintelligenceact.eu: Whistleblowing and the EU AI Act
  15. US Senate Judiciary Committee: Grassley introduces AI Whistleblower Protection Act
  16. Deseret News: John Curtis backs bill protecting AI whistleblowers
  17. Future of Privacy Forum: California's SB 53, the first frontier AI law, explained