Last week was a bad one for AI systems. OpenAI disclosed six cases in which its artificial intelligence models behaved in ways the model maker didn’t expect. One searched public repositories for an exposed application programming interface (API) key and used it without permission. Another uploaded a file to the internet so it could cite it. Some instances even wrote instructions into their own summaries telling models that followed to conceal mistakes from users. Then The Wall Street Journal revealed that Google’s Gemini model had also misbehaved, hacking companies’ IT systems during routine tests.
Both incidents, plus a slew of others in which AI systems have exceeded their test boundaries, are warning lights flashing on a dashboard. But how we learn about them is also a concern. Invariably, the company that built the system is also the one that decides to reveal what happened, frames how serious it was, and picks when we learn about it.
It has led some to wonder whether we need an alternative approach. More than 100 AI experts signed an open letter last week calling for independent safety evaluators to police the leading AI labs.
One approach could be to borrow from the aviation industry, where aircraft are tested before they fly and, when something goes badly wrong, separate accident investigators step in afterward. “At least in some categories of incidents, there should be an independent incident investigation,” says Michael Chatzipanagiotis, an assistant professor of private law at the University of Cyprus who has studied how aviation-style incident reporting could apply to AI.
Right now, AI regulation often muddles together two different jobs, Chatzipanagiotis says: letting the world know that something has happened and investigating why. The latter is often put in the hands of a small number of organizations with industry links. And as AI systems gain the ability to browse the web, use tools, and act with little human oversight, companies still face potential liability and reputational damage from disclosing problems. Asking them to investigate themselves gives them good reasons to present what happened in the most favorable light.
Marius Hobbhahn, CEO and cofounder of the AI safety organization Apollo Research, believes an independent investigatory group would be useful. But he thinks we are some way from making it work. “In the status quo, a third party evaluator would not have enough access,” he says. “We need to get to a point of full transparency for the investigator; otherwise we cannot make a good assessment.”
Hobbhahn says investigators would need training records alongside logs of the incident itself: transcripts, timelines, the model weights being used, the computing cluster involved, and the safeguards that fired, failed, or did nothing—something equivalent to aviation’s black box.
AI testing also has one advantage over aviation investigations: Researchers can rerun the same circumstances to see whether the behavior replicates. That, Hobbhahn says, can help determine whether the action was a one-time thing or something baked into how an AI system works.
The push for independent evaluation is growing; a major announcement was made at the U.N. General Assembly this week. On Monday, Rumman Chowdhury, a former member of the U.S. Artificial Intelligence Safety and Security Board under the Biden administration, launched the Independent AI Evaluation Foundation (IAEF), backed by $10 million in philanthropic funding initially focused on education. The aim is to professionalize the community of independent AI evaluators, making independent evaluation something people can have a career doing and, in the process, improving its standards.
The goal is also to help fill the gap in AI checks after models are deployed. “Right now that loop doesn’t close,” Chowdhury says. An AI company can disclose an incident, then “it sort of floats off into the ether once that is announced, and there’s nobody to catch that,” she explains. Her hope is that independent evaluators could work alongside regulators and AI safety institutes on predeployment testing, ongoing testing as models change, and postincident investigation.
The foundation is trying to build some of the infrastructure that would make that possible. Part of its funding will create an open-access evaluation environment bringing together existing benchmarks, testing tools, and evaluators. Other money will go toward training, fellowships, and funding model evaluations themselves.
Chowdhury wants independent evaluation to become an actual profession rather than something performed ad hoc by academics or researchers with privileged access to AI companies. IAEF also wants to separate evaluation from validation. Rather than asking only what a test found, it will also ask whether the evaluator ran the right test and asked the right questions in the first place.
Gabriella Waters, who is working with IAEF on the validation layer of the foundation’s education effort through the Center for Responsible AI, argues that investigations also need to look beyond technical failure and toward real-world problems that emerge in use. “Why isn’t evaluation part of the layer of AI development?” Waters asks. She points out that as models keep changing, model evaluation needs to be an ongoing process, too.
No one is suggesting that AI needs a carbon copy, tomorrow, of the National Transportation Safety Board in the United States or the Air Accidents Investigation Branch in Britain. Chatzipanagiotis, the University of Cyprus professor, warns that aviation’s system resulted from decades of work within the industry and a level of international agreement that doesn’t exist for AI, particularly given the geopolitical conflict arising from the U.S.-China competition over the technology.
But the post-hoc, haphazard approach to major AI incidents so far isn’t sufficient. “Usually, we have near misses, we have warning signs, which if we ignore, we will pay the price later on,” says Chatzipanagiotis.