Research
AI Is Reviewing Sustainability Data. Who Reviews AI?
AI hasn't earned its way into sustainability review; it was pulled in because manual review cannot scale to CSRD, LCA, and EPD volumes. The real question is not whether AI should check the data, but whether AI-assisted review can be governed so that a named, accountable person still answers for the result.
Life cycle assessment (LCA) studies and disclosures under the EU’s Corporate Sustainability Reporting Directive (CSRD) runs through thousands of individual claims: emission factors, product category rule (PCR) clauses, double materiality assessments covering both a company’s environmental impact and the financial risks it faces. Each claim requires a qualified human sign off. However, there simply aren’t enough qualified reviewers to do that by hand at the volume the frameworks now require. Sustainability reporting has a volume problem before it has an accuracy problem, and that is the opening AI walks into.
Necessity before trust
AI hasn’t earned its way into the review role. It was pulled in because the alternative, manual review at current volume, doesn’t scale. That distinction matters, as trust in AI as a reviewer isn’t established, it’s assumed by necessity, and the two are not the same thing. An institution can adopt a tool because it has no better option and still owe itself a hard look at whether the tool deserves the role it’s been handed.
The volume hasn’t eased either. The EU’s Omnibus I Directive narrowed how many companies fall under CSRD, but the companies still in scope face the same double materiality reporting requirements as before. Fewer filers, same depth per filer. The pressure that pulled AI into review in the first place hasn’t relaxed, it’s just concentrated on a smaller set of reports that are each just as dense as they were. A narrower scope was never going to be the thing that settled the question of whether AI belongs in the review chain.
What AI actually changes is what gets checked at all. A human reviewer working through a report has to sample: pull a handful of emission factors, spot check a few PCR clauses, and hope the pattern holds across the rest. Fatigue sets in a few hundred datapoints into a report that runs to thousands, and the inconsistencies that slip through tend to be the quiet ones, a boundary assumption that shifted halfway through, a clause addressed in spirit but not in the specific wording the standard requires. AI can cross-check every emission factor against its source, flag every PCR clause that’s only partially addressed, and catch exactly those quiet inconsistencies. That’s not a marginal improvement on the old process. It’s a different process altogether, checking a different fraction of the report, and a wider fraction checked isn’t automatically a more reliable check. Coverage and trust are two different questions. One asks how much of the report got checked and the other asks how much confidence we have on the checker itself. AI answers the first convincingly however, the second still remains unresolved.
Which raises the actual question: whether or not AI should review sustainability data, since the volume math already answered that, but whether AI review can avoid becoming the failure point the review step exists to catch.
A trust built for a person
The standards governing this work were not written with a non-human reviewer in mind. The International Organization for Standardization’s (ISO) 14044 critical review, a PCR verifier’s sign off, Greenhouse Gas (GHG) Protocol assurance: all three are built around a named, accountable person. The artifact of trust in each case isn’t just the number that got approved but the professional judgment of the person who approved it, and the fact that they can be asked, later and by someone else, to walk through why.
AI doesn’t fit cleanly into that role. It can flag an anomaly, but it can’t be cross examined about its reasoning the way a verifier can. A verifier who accepted an unusual boundary assumption can explain, on the record, why that assumption was reasonable for this product and this context. A model that flagged or cleared the same assumption has no equivalent account to give. Whether that gap matters comes down to governance: governance is what decides if AI stays a support function underneath the accountable person, or quietly erodes the premise that a human is the one actually answering for the result.
The chain that has to hold
Good governance follows a specific sequence. AI flags an issue, a qualified human confirms or overrides that flag, and provides reasoning. That decision gets logged against the exact model and prompt version that produced the flag, because a model updated between one reporting cycle and the next can flag differently on the same underlying data, and a decision log that doesn’t say which version was in use can’t tell anyone whether that’s what happened.
Break any one link and the chain stops doing its job. An AI flag nobody acts on is noise, indistinguishable from a false positive. There is no way to tell whether or not it was actually catching something real, because acting on it is what proves its value. The same logic applies to a human override, with no stated reasoning, nobody can reconstruct why the call was made, neither by an auditor checking the file nor by the reviewer’s own future self trying to remember the call. A model update that isn’t tracked against the decisions it produced makes last year’s results and this year’s results incomparable, even when nothing else about the underlying report has changed.
Governance is the difference between “AI touched this” and something a person can actually trace, question, and defend. Without it, the first phrase is all anyone can honestly say, and it isn’t an answer to the question an auditor is actually asking.
The template already forming
CSRD assurance itself is still being defined. Omnibus I dropped what had been a planned move toward reasonable assurance, a higher evidentiary bar than the limited assurance most disclosures still face. The European Commission is due to publish harmonized limited assurance standards by mid 2027. Whatever bar eventually applies, the assurance statement is still issued by a named, accredited provider who is accountable for it, regardless of whether or not AI touched the file along the way.
The clearest version of the pattern so far is showing up earlier in the pipeline, on the preparation side of environmental product declarations (EPD). AI tools are already interpreting PCR requirements, flagging data gaps, and suggesting system boundaries. What hasn’t moved is the step after: an independent, accredited verifier still has to review the submission and sign off before an EPD can be published. This pattern isn’t a settled template across sustainability reporting yet, but more a shape both frameworks keep landing on: whatever AI touches earlier in the process, a named person still has to be the one who signs at the end.
Judgment, not scale
This division changes what the human reviewer’s job actually is, in both directions at once. It narrows in scope: judging the anomalies AI surfaces rather than trying to catch everything cold across a report, rather than reading too closely. Such procedures deepens accountability: the reasoning behind a confirmation or an override has to be explicit enough to survive an audit, not just sit in the reviewer’s head as professional intuition nobody asked them to write down.
Trust in AI assisted review will end up resting on the strength and traceability of such governance trails, not on how fast or sophisticated the underlying tool is. A faster model that skips the logging step is not a more trustworthy reviewer. It’s an unaccountable one that happens to work quickly, and speed was never the thing the review step was built to protect. The review step exists to catch failure and a faster failure point is still a failure point nonetheless.
- AI
- CSRD
- Assurance
- Governance
- EPD
- PCR
Written by
Verdatir
Research and perspectives from the Verdatir team on verification, interoperability and the governance of environmental data.