The report most companies write is not the report the board asked for
The typical AI risk report runs thirty slides and opens with a table of every model in production, with a red, amber or green dot beside each. It is accurate. It is expensive to produce. It is also close to useless for the decision a board is making, because a board is not managing models. A board is deciding whether the company's exposure to a category of failure is bounded, whether somebody with a name and a budget owns it, and whether the position got better or worse in the last ninety days. An inventory answers none of those.
The tell is what happens in the room. The report goes up, the audit committee chair asks one question that the thirty slides do not answer, and the meeting spends fifteen minutes reconstructing an answer from memory. The question is almost always some version of: if this thing produced a wrong output tomorrow, on the worst customer, what is the largest number we could end up writing a check for. That question has a shape. The report should have been built around it.
The rest of this piece is what we build when a company asks for that report. It works for a ratings firm scoring credit, a health-tech company routing clinical documentation, a manufacturer running vision inspection on a defense-adjacent line, or an integrator putting a model in front of a government customer. The industries differ. The four questions do not.
Question one: which of these exposures has no ceiling
Directors sort risk into two piles without being taught to. The first pile is bounded: the worst case is a known number, and the number is survivable. The second pile is unbounded, or unknown, which a board treats as the same thing. Almost all the attention should go to the second pile, and almost all of a typical report goes to the first.
An exposure is bounded when three things are true. The population of affected decisions is countable. The per-decision consequence has a cap, whether from a contract limitation of liability, an insurance layer, a regulatory penalty schedule, or the simple size of the transaction. And there is a detection path that finds the problem before the population grows further. Miss any one of the three and the exposure belongs in the second pile until somebody does the work to move it.
Write each exposure as a sentence with a number in it. "The pricing model touches 2.1 million quotes a year; a systematic error is capped by the contractual liability limit at the greater of fees paid or $2M per customer, and we detect drift within one billing cycle." That is a bounded exposure, and a director can read it once and move on. Now the other kind: "The clinical summarization feature writes into the record for roughly 40,000 encounters a month. We do not currently know how many summaries a clinician accepted without reading, and the liability is not capped by contract." That is four lines and it will occupy the committee for the rest of the session, which is the correct outcome.
What a director actually reads for, by attention paid
Our editorial weighting from practitioner reading and board-material reviews. Illustrative ordering, not a measured statistic.
Question two: who owns this, by name
A risk with no owner is a risk the board has to hold itself, which is the thing directors most want to avoid. So the second column of the report is a person, not a function. Not "Engineering." Not "the AI Governance Working Group." A name, a title, and the budget line the mitigation is funded from.
This is where most reports quietly break. Ownership of an AI failure tends to be split three ways by accident: the model is built by a data team, deployed by a platform team, and the consequence lands on a business unit that was not in either conversation. Each of the three believes one of the other two owns it. The report is the forcing function that finds this out, because writing a single name next to an exposure requires somebody to agree to be that name.
The federal frameworks are unusually blunt about this, and the language transfers cleanly to a commercial board. NIST AI RMF puts accountability structures in the Govern function rather than treating them as a technical matter, and the Govern 2.1 subcategory is specifically about roles, responsibilities and lines of authority being documented and clear. ISO/IEC 42001, the auditable AI management-system standard published in December 2023, requires top-management commitment and defined roles as a certification condition, which means an auditor will ask for the name. Banking supervisors got there twelve years earlier: SR 11-7, the Federal Reserve and OCC guidance on model risk management issued in April 2011, assigns explicit roles to model owners, model developers and independent validators, and treats the separation between the last two as the point of the exercise. None of these frameworks is telling you how to build a model. They are all telling you to write down who is responsible when it is wrong.
Question three: what changed since last quarter
A board reads the same report four times a year. The value is in the delta, and a report that regenerates its ratings from scratch each quarter destroys the delta. Directors notice. The third question is always some form of "is this getting better."
So the report carries a change log, and the change log has exactly four kinds of entry. New exposure appeared, with the reason it appeared. Exposure closed, with the evidence that closed it. Rating moved, with what moved it. Nothing changed, which is itself an answer and sometimes the most important one, because an exposure that has sat amber for four consecutive quarters is telling the board something about resourcing that no single-quarter snapshot conveys.
Keep the ratings stable across periods even when a new framework arrives. If the company adopts a new control catalog mid-year, map the old ratings forward and show both, rather than restating history. Restated history reads as a reset, and a board that suspects a reset stops trusting the trend, which was the only part of the document doing real work.
Question four: what is coming that we do not control
The last thing a board wants is a date it did not know about. Regulatory and customer-driven deadlines belong in the report even when nothing is wrong, because they set the schedule the mitigation work has to fit inside.
For a company whose product touches government, the calendar items that keep showing up are FedRAMP authorization timing for anything sold as a service to a federal agency, NIST SP 800-171 flowdown when the contract touches controlled unclassified information, and CMMC assessment requirements arriving through defense subcontracts. For a company selling into Europe, the EU AI Act high-risk obligations moved once already and now sit at December 2027 for the Annex III categories and August 2028 for AI embedded in regulated products. For anyone in financial services, the SR 11-7 validation cadence is a standing commitment rather than an event.
Two rules make this section useful instead of decorative. Each date carries the source and the day somebody last checked it, because these dates move and a stale citation in a board document has a long half-life. And each date carries the internal decision that has to happen before it, with its own earlier date. A compliance deadline in eighteen months usually means an architecture decision in four.
How to write the exposure lines so they survive the room
Every exposure gets one paragraph with five parts, in this order: the population, the failure mode, the detection path, the ceiling, and the owner. That order is deliberate. It moves from the concrete to the accountable, and it puts the number before the name so the person accepting ownership knows what they are accepting.
The failure mode has to be specific enough to be wrong. "Hallucination risk" is not a failure mode; it is a category. "The retrieval step returns a superseded policy document when two versions share a title, and the generated answer cites the old one with the current one's date" is a failure mode. It can be tested, it can be detected, and it suggests its own fix. Categories cannot be closed, which is why reports full of categories never show progress.
The named-framework vocabulary earns its place here and only here. The OWASP Top 10 for LLM Applications gives you shared names for prompt injection, insecure output handling, training data poisoning and excessive agency, which means a security lead and a business owner can discuss the same thing without translating. MITRE ATLAS does the same for adversarial machine-learning techniques against deployed systems, with a case-study base that lets you say "this has happened to someone" rather than "this is theoretically possible." Use the names when they save words. Do not include the framework diagram; nobody on a board has ever needed it.
| Instead of this line | Write this line | Why it lands |
|---|---|---|
| "Model accuracy: 94%" | "6% of 180,000 monthly decisions are wrong. 4,300 of those reach a customer before review catches them." | Converts a metric into a population and a consequence |
| "Hallucination risk: medium" | "Retrieval returns a superseded document when versions share a title; the answer cites the old text with the new date. Seen twice in Q2 testing." | A failure mode can be closed. A category cannot. |
| "Vendor concentration risk" | "One provider supplies the model behind 71% of AI revenue. We have no evaluation suite that would catch a quality regression on a switch, so we cannot move." | Names the reason the risk is structural, not just present |
| "Data governance: in progress" | "Two of nine customer contracts permit training on their data. The corpus loses 38% of its volume if the other seven enforce their terms." | Turns a status into an amount |
| "Controls implemented" | "Human review is designed for all high-severity outputs. Tested in March on 200 sampled items; 31 had no reviewer action recorded." | Separates designed from tested, which is the audit question |
| "Compliance roadmap on track" | "Annex III obligations apply December 2027. The architecture decision that determines whether we qualify is due February 2027." | Gives the board the date it can actually act on |
Designed, tested, and monitored are three different states
The single most common way an AI risk report misleads without lying is by reporting design as if it were operation. A control that exists in a document, a control that was tested once against real items, and a control that produces a signal somebody watches are three different levels of assurance, and a green dot flattens all three into one.
So the report uses three states and never one. Designed means it is written down and someone could follow it. Tested means somebody ran it against real items on a date, with a result, and the result is cited. Monitored means it emits a measurement on a cadence, and there is a threshold that triggers action. Most controls in most companies are in the first state, and saying so plainly is what buys the report credibility for the exposures where the answer is better.
This distinction is also where an outside reviewer pays for itself, because the people who designed the control are the wrong people to test it. That is the whole logic of independent validation in SR 11-7, and it is not a banking peculiarity. A developer testing their own control tests the cases they had in mind when they built it, which are by construction the cases it handles.
Where an AI control usually sits when we first look at it
Editorial weighting from practitioner reading, illustrating the usual gap between designed and operating. Not a measured statistic.
Six pages, and what goes on each one
Length is a governance decision, not a formatting one. A report the committee reads in full before the meeting produces a different conversation than a deck skimmed during it. Six pages is the size that gets read.
Page one is the four unbounded exposures, one paragraph each, with owners. Page two is the change log against last quarter. Page three is the control state table, designed against tested against monitored, with the tested column carrying dates. Page four is the external calendar with sources and last-checked dates. Page five is the incidents and near-misses, including the ones that were caught, because a caught near-miss is evidence a control worked and it is the only kind of good news in this document that a board believes. Page six is the ask: the specific decisions or funding the report is requesting, each tied to an exposure on page one.
Everything else goes in an appendix nobody is required to read. The full model inventory, the framework mappings, the benchmark tables, the architecture diagrams. Keep them current, hand them to the auditor, and keep them out of the six pages.
What it costs and how long it takes the first time
The first cycle is the expensive one because the answers do not exist yet. For a company with somewhere between three and fifteen models in production, building the first report generally runs four to eight weeks of work, most of it spent on two things: getting a defensible count of the decision population under each model, and finding the person who will put their name in the owner column. The technical work is the small part. In our experience that first pass lands somewhere in the range of $40,000 to $120,000 depending on how many business units are involved and whether the contract review has to be done alongside it.
The quarterly refresh, once the structure exists, is a fraction of that, usually one to two weeks. The cost drops because the format is stable and the work becomes updating deltas rather than establishing baselines. That asymmetry is the argument for building the structure properly the first time rather than assembling something ad hoc for the next meeting and rebuilding it every quarter forever.
The number that matters more than either is the one a bad report costs. A board that cannot answer a regulator's or an acquirer's question about AI oversight from its own materials ends up commissioning the work under deadline, at a premium, with somebody else setting the scope. Diligence teams and examiners both ask the same first question: show us the last four board reports. The value of having them is almost entirely in having them already.
What we do
We build the report and the machinery under it. That means reading the contracts to find where liability is actually capped, counting the decision populations from the systems rather than from a slide, running the independent tests that move a control from designed to tested, and sitting in the room where owners get assigned. We write the six pages, we write the appendix that backs them, and we set up the quarterly refresh so the next cycle is a delta rather than a rebuild.
We work in the vocabulary the auditors and the federal customers already use, because a report written to NIST AI RMF and ISO 42001 structure answers a commercial board and survives a government reviewer without being written twice. If your product touches a government customer at any point, that dual-audience property is worth building in from the first draft.
Bottom line
A board wants four things from an AI risk report: the exposures with no ceiling, a named owner for each, the delta since last quarter, and the dates it does not control. Everything else is appendix. The report that answers those four in six pages will get read, will get funded, and will hold up when a regulator or an acquirer asks to see the last year of them. The thirty-slide inventory will not, no matter how accurate it is.
Frequently asked questions
Quarterly for the standing report, with an out-of-cycle note when a new unbounded exposure appears or a material incident occurs. Quarterly is frequent enough to show a trend and infrequent enough that the delta is meaningful.
Usually the audit or risk committee owns the detail and the full board gets the page-one exposures and the ask. Write it so the first page stands alone, since that is the page most directors will read.
No. The standard is useful as a structure for roles, documentation and management review whether or not you certify. Certification becomes relevant when a customer or an insurer asks for it, and building the report to that structure means the audit is not a separate project.
If you already run SR 11-7 style validation, most of the machinery exists and the gap is usually coverage of systems nobody classified as models, plus the board-facing summarization. Generative and retrieval systems frequently sit outside the model inventory that the validation process reads from.
That is a finding, and it goes on page one. An exposure whose population cannot be counted is unbounded by definition, and counting it is usually a few days of instrumentation rather than a project.