Building an AI Ethics Board

Mark BarclayMark Barclay·Founder & Curator, SynaBot·

A Step-by-Step Guide

This article is part of my series on AI safety and governance. It's also the most operational piece I've written, so if you haven't read the pillar article or my piece on responsible AI development practices, I'd suggest those first, since this article is really about how to formalize a lot of what I cover there.

Why I Wanted This to Be the Last Piece in the Series

I've spent most of this series talking about ideas, alignment, governance frameworks, risk assessment, transparency, and about practices, documentation, review, staged rollouts. What I haven't really addressed is the structural question: who, inside an organization, is actually responsible for holding all of this together?

I think an AI ethics board, or whatever an organization chooses to call it, is one answer to that question. I don't think it's the only answer, and I don't think standing one up automatically makes an organization's AI practices good. But I've come to see it as a useful forcing function: a structure that, if built well, creates the conditions for a lot of the other practices I've discussed to actually happen.

Step One: Get Clear on What the Board Is For

I think the most common mistake I've seen organizations make is starting with structure, who sits on the board, how often it meets, before getting clear on purpose. I think this leads to boards that exist but don't do much, because nobody's quite sure what they're supposed to be evaluating or deciding.

I'd encourage starting with a small set of concrete questions. What kinds of decisions does this board actually need to weigh in on? Is it reviewing specific products or features before launch? Is it setting organization-wide policy? Is it investigating incidents after the fact? I think these are different functions, and while one board can sometimes do all of them, I think it's worth being explicit about which ones you're building for, because the answer shapes everything else.

I also think it's worth being honest about what the board is not. I don't think an ethics board should be the only place safety considerations get addressed, that would put too much weight on a single structure and create bottlenecks. I think of it more as a backstop and an escalation point: a place where decisions that are genuinely hard, or that cut across teams, or that carry significant risk, get a level of scrutiny that wouldn't happen otherwise.

Step Two: Think Carefully About Composition

I think who sits on this board matters more than almost anything else about how it's structured, and I think this is where I've seen the most variation in how seriously organizations take the exercise.

I'd think about a few dimensions. First, expertise: I think a board needs people who understand the technical side of what's being built, people who understand the ethical and societal dimensions, and often people with legal or regulatory expertise, since a lot of what gets discussed will have compliance implications. Second, independence: I touched on this in my article on responsible AI practices, but I think it's worth restating here specifically. If everyone on the board has the same incentives as the teams whose work they're reviewing, namely, shipping products, I don't think the board will function as a meaningful check, regardless of its formal authority.

I think this is why some organizations bring in external members: people from outside the organization who don't have a stake in any particular product's timeline. I've seen this done through external advisors, academic partnerships, or formal outside seats on the board itself. I think external perspective is valuable, but I'd also say it only works if those external members actually have access to the information they need and aren't just there for appearances.

Third, I think diversity of perspective matters in a broader sense than just professional background. The kinds of harms an AI system might cause often fall disproportionately on groups that aren't well represented in a typical engineering organization. I think a board that doesn't have some mechanism for bringing in perspectives from outside the organization's usual demographic makeup is more likely to miss things, not because of anyone's individual failings, but because of how blind spots tend to work.

Step Three: Define Authority Clearly

I think this is the step that determines whether a board is real or symbolic, and I'd encourage being very direct about it from the start.

Can the board actually delay or block a launch? Can it require changes to a product before it ships? Or is its role purely advisory, offering input that teams can take or leave? I don't think there's one right answer here, different organizations will land in different places depending on their size, structure, and culture. But I think whatever the answer is needs to be explicit and known, because I've seen real damage done by boards that everyone assumed had more authority than they actually did, right up until the moment that assumption was tested and found wanting.

I'd also think about escalation paths. If the board flags something serious and a team disagrees, what happens next? Is there someone more senior who makes the final call? I think having this worked out in advance, before there's a live disagreement with deadlines and pressure attached, makes a real difference in how these situations get handled.

Step Four: Build the Review Process

Once purpose, composition, and authority are clear, I think the next step is actually designing how reviews happen. I'd think about this in terms of inputs, process, and outputs.

For inputs, I think the board needs access to the kind of documentation I discussed in my responsible AI practices article: what's being built, what testing has been done, what the known risks and limitations are. I don't think a board can do meaningful review without this, and I think one of the most common failure points is a board that's asked to weigh in without being given enough information to do so.

For process, I think it's worth deciding upfront how reviews actually unfold. Does the board meet regularly to review a queue of items, or are reviews triggered by specific events, like a new product reaching a certain stage? I think regular cadence has the advantage of predictability, teams know when review will happen and can plan for it, but I think trigger-based review can be more responsive to things that don't fit a regular schedule, like an emerging issue that needs urgent attention.

For outputs, I think the board's conclusions need to be documented in a way that's useful later. Not just "approved" or "not approved," but the reasoning behind it, any conditions attached, and what would need to change for a different outcome. I think this connects back to the documentation theme that's run through this whole series: a decision that isn't documented is much harder to learn from, revisit, or hold anyone accountable to later.

Step Five: Plan for How the Board Learns and Adapts

I don't think an ethics board should look the same a year after it's created as it did on day one, and I'd be skeptical of one that does. I think the kinds of issues that come up will change as an organization's products and the broader landscape evolve, and I think the board's processes need to be able to change with them.

I'd build in some mechanism for the board to reflect on its own performance. Are decisions actually being acted on? Are issues that should be reaching the board actually reaching it, or are things slipping through? Is the board spending its time on the things that matter most, or getting bogged down in lower-stakes reviews while bigger issues don't get the scrutiny they need?

I think this connects to the feedback loop theme from my responsible AI practices article in a fairly direct way. A board that doesn't have a way of learning from its own track record is, in a sense, missing its own feedback loop, even while it's helping ensure other parts of the organization have theirs.

What I Think Success Actually Looks Like

I want to end this piece, and this series, with what I think is the most honest answer to "how do you know if this is working."

I don't think the signal is the absence of problems. I think organizations that build genuinely functioning review structures will find more issues over time, not fewer, at least initially, simply because they're looking harder and have a structure that surfaces things that previously went unexamined. I think the signal I'd actually look for is whether the board's findings change outcomes: whether products get delayed, redesigned, or in some cases not shipped at all, based on what the board surfaces. And I'd look for whether people across the organization, not just the board members themselves, see the board as a place where hard questions get a real hearing, rather than a formality to get through on the way to launch.

I think this is a fitting place to end the series, because it brings together everything I've discussed: the technical concepts like alignment and explainability, the practices like risk assessment and red-teaming, and the external context of governance frameworks and regulators. None of those things matter much if there isn't some structure inside an organization responsible for holding them together and making sure they actually inform what gets built and deployed. I don't think a board is a complete answer to that, but I think it's one of the more concrete steps an organization can take toward having one.

If you've made it through this whole series, I hope it's given you a grounded sense of where things stand, what's genuinely hard, and what's actually within reach. I plan to keep updating these pieces as the landscape shifts, because I expect it will keep shifting for a long time to come.