AI Incident Reporting

Mark BarclayMark Barclay·Founder & Curator, SynaBot·

Why It Matters and How It Works

This article is part of my series on AI safety and governance. If you haven't read the pillar article, it'll help orient you to where this fits.

Why I Called This the Unglamorous Cousin of AI Safety

In the pillar article, I described incident reporting as the part of this whole field that nobody really wants to talk about, and I think that's because it's fundamentally about failure. Risk assessment and red-teaming, the topics I covered earlier in this series, are about trying to find problems before they happen. Incident reporting is about what happens after something has already gone wrong, and I think that makes it feel like an admission, which I suspect is part of why it gets less attention than it deserves.

But I've come to think incident reporting might be one of the most consequential pieces of this whole puzzle, precisely because it's the mechanism through which the field is supposed to learn from its mistakes. Without it, I think every organization is left rediscovering the same failure modes independently, with no shared record of what's already gone wrong elsewhere.

The Comparison I Keep Coming Back To

I think the comparison to other industries is genuinely useful here, and it's one I've seen others draw too. Aviation has decades of mandatory incident reporting, where airlines and manufacturers are legally required to disclose accidents and near-misses to regulators, who investigate and maintain centralized records. The same is true for medical devices, where failures get reported to regulatory bodies that can track patterns across products and manufacturers.

I think the value of this kind of system is that it lets a whole industry learn from any single failure, not just the organization where that failure happened. If a particular failure mode keeps recurring across different products from different companies, a centralized system makes that pattern visible in a way that no individual company's internal records could.

I think AI is, in an important sense, missing this. There's currently no equivalent of the FAA's accident database or the FDA's adverse event reporting system specifically for AI. What exists instead are a set of voluntary, civil-society efforts, built largely on public reporting rather than the kind of mandatory first-hand disclosure that aviation and medical devices require.

What Exists Today

The project I think is most established in this space is the AI Incident Database, run by a nonprofit called the Responsible AI Collaborative. I think of it as something like a public catalog: it collects reports of AI-related harms from public sources, indexes them, and makes them searchable, with the explicit goal of helping the field learn from documented failures. By the time researchers reviewing its dataset had published an analysis of it, the database had cataloged several hundred incidents, and I've seen it referenced widely in academic and policy discussions of AI harms.

I think the OECD's AI Incidents Monitor is a useful complementary effort, operating at a larger scale, drawing on a substantial number of reported incidents and hazards, sourced from news coverage and classified along dimensions like the type of harm, who was affected, and what kind of AI system was involved. I find the distinction it draws between an "incident," where harm has actually occurred, and a "hazard," where there's a plausible risk of future harm but nothing has happened yet, to be a genuinely useful framing, because I think it captures something important: not everything worth tracking and learning from is itself a harm that's already materialized.

What I think is important to understand about both of these efforts is that they're built on secondhand, publicly available information. Their quality depends on what journalists, researchers, and other contributors happen to notice and report, not on companies being required to disclose what they know internally. I think this is a real limitation, not because the people running these databases aren't doing careful work, but because it means the record is necessarily incomplete in ways that are hard to fully characterize. Incidents that don't generate public attention, or that companies have strong incentives not to discuss, may simply not appear.

The Push for Mandatory Reporting

I think this gap, between what voluntary, publicly-sourced databases can capture and what a true incident reporting system would capture, is exactly why I've seen growing calls for mandatory reporting requirements, modeled more explicitly on the aviation and medical device examples.

I've seen proposals that focus specifically on high-consequence events: things like deaths, serious injuries, critical infrastructure disruptions, or incidents involving particularly dangerous capabilities, as the kind of narrower category that might be most realistic for a mandatory regime to start with. I think the logic behind narrowing the scope this way is twofold: it's more politically feasible than trying to mandate reporting of every AI-related issue, and it focuses attention on the events where learning from failure matters most.

I've also seen real tension in these discussions around the question of liability. Some have argued that a well-designed incident reporting system could actually reduce litigation risk for companies, by creating a structured, less adversarial channel for surfacing problems. Others have pushed back, suggesting that any mandatory disclosure regime creates exactly the kind of paper trail that makes litigation easier, which I think is part of why building political consensus around mandatory reporting has been slow going.

What the EU AI Act Does Here

I think the EU AI Act is the most concrete example I've found of incident reporting requirements actually being written into binding law, though I'd flag that the timeline for these specific obligations is notably long, even relative to the rest of the Act's already-extended rollout that I discussed in my article on global governance frameworks.

The Act defines a category of "serious incident" with criteria that go well beyond what I think most people would consider a typical IT incident: things like death or serious injury caused by an AI system, which could include scenarios like a medical misdiagnosis that delayed treatment, an autonomous vehicle accident, or a critical infrastructure failure, as well as cybersecurity breaches involving AI systems that compromise personal data or system integrity.

What I find notable is just how far out the actual reporting obligations under this part of the Act are scheduled to take effect, with full incident reporting obligations for standalone high-risk systems not applying until December 2027, and an even later date for high-risk AI embedded in regulated products. I think the practical takeaway here is that even in the jurisdiction with the most developed binding framework, AI-specific incident reporting obligations are still mostly in the future. In the meantime, AI-related harms remain subject to whatever existing sectoral law already applies, things like product liability, data protection law, or medical device regulation, depending on the context.

Why I Think the Editorial Challenges Are Underrated

One thing that struck me in looking at how existing incident databases actually operate is how much of the work isn't collection, it's interpretation. I've seen this described directly by people who've worked on these databases for years: there's an unavoidable amount of uncertainty in classifying incidents, around questions like what actually caused a given harm, how severe it really was, or even what kind of AI system was involved in the first place.

I think this connects back to the explainability themes I covered earlier in this series. If it's hard to explain why a specific AI system produced a specific output, I think it follows that it's also often hard to definitively establish that a given AI system caused a given harm, especially when the system is one part of a larger chain of events involving human decisions, other software, and real-world circumstances. I don't think this uncertainty is a flaw in how incident databases are run; I think it's an inherent feature of trying to attribute harms to AI systems at all, and I think any future mandatory reporting regime will have to grapple with the same problem.

What I Think This Means for Organizations

Even setting aside the question of what's legally required, and given how far out most binding requirements currently are, I think there's a strong case for organizations to build their own internal incident tracking now, rather than waiting for external mandates to force the issue.

I think this connects directly to the monitoring and feedback loop themes I discussed in my articles on risk assessment and responsible AI development. An organization that's already tracking its own incidents, classifying them, and feeding that information back into how it builds and deploys systems is, in effect, already doing the hard part of what a future mandatory reporting regime would require. The main thing a mandatory regime would add is the obligation to share that information externally, which I think is a meaningfully smaller lift for an organization that already has the internal infrastructure than for one that's starting from nothing.

Where I Land on This

I think incident reporting for AI is currently in a state that I'd describe as "necessary but not yet built," at least not in the comprehensive way that exists for aviation or medical devices. The voluntary databases that exist are valuable, and I think they've done real work in making AI-related harms visible and giving researchers something to study. But I don't think they're a substitute for the kind of mandatory, first-hand disclosure regime that other safety-critical industries rely on, and I think the timelines for binding requirements, even where they exist, suggest this gap will persist for a while yet.

I'd also say this is the piece of the broader safety and governance picture that I think depends most on all the other pieces working. Incident reporting only produces useful learning if there's a clear definition of what counts as an incident, which connects to the explainability challenges I've discussed. It only changes outcomes if organizations have review processes, like the ones I described in my piece on responsible AI development, that can actually act on what gets reported. And it only builds the kind of cross-industry trust that aviation has if it eventually moves from voluntary and reactive to something closer to mandatory and systematic. I don't think any of that happens quickly, but I think it's worth being clear-eyed about the gap, rather than assuming the existing voluntary efforts are doing more than they actually can.