A KYC score, a fraud flag, a CV rejection: each is a durable, timestamped judgement about a named person or company. The reasoning that produced it evaporates in months. The record lasts years, and one day someone with authority will ask who signed it.
A mid-sized bank’s transaction monitoring system generates a few thousand contested judgements before lunch. Every KYC adverse-media match, fraud-propensity flag, supplier risk score, and CV screening rejection is a durable, timestamped verdict about a named person or company. With 52% of DIFC firms now using AI, that production line is already running across the district.
Here is the asymmetry that should worry a compliance head more than model accuracy: the model’s reasoning is ephemeral, its output is permanent. Three years from now, in a dispute or a supervisory review, the score is the evidence. Everything that would make it defensible is gone.
The artefact outlives the argument that produced it
A classifier that scored a vendor high-risk in 2024 leaves behind the score. It does not leave the model version, the prompt, the threshold, the training distribution, or the analyst who set the policy. The judgement that produced the entry evaporates within months. The entry itself lasts years, and reads to a hostile audience like a verdict your firm handed down.
Most enterprise content and risk systems capture the outcome and discard the reasoning. That was a storage decision, usually made by whoever configured the pipeline. It is now the difference between a defensible record and an indefensible one.
GARM was ended by the cost of discovery, not by a finding
The cleanest case study available is commercial and recent. GARM, the advertising industry’s brand-safety body, produced shared taxonomies of content categories advertisers would avoid. Entirely voluntary, entirely commercial. X Corp sued in August 2024, and GARM discontinued operations within days. Its parent federation said plainly that it lacked the resources to defend the litigation.
No court ever ruled the taxonomy wrong. The organisation was ended by the cost of the discovery process.
Read that carefully, because it is the actual risk model. You do not need to lose. You need only be unable to afford the reconstruction. If your team cannot produce, on demand, the model version, threshold, input, output, and named policy owner for one decision made 30 months ago, a subpoena is an operational impossibility rather than a legal problem.
Nobody signs a probability
Most AI governance programmes invest heavily in explainability: feature attributions, attention maps, model cards. They invest almost nothing in attribution of authority. That is backwards. No inquiry has ever been satisfied by an explanation of how a model arrived at 0.83. Inquiries want a name.
The fix is concrete and cheap to implement now. Any AI-derived judgement that attaches to a named person or company gets three things stored alongside the value: the policy version that authorised the decision, the human role accountable for that policy, and the written reasoning at the time. Stored once, immutably, at the moment of the call. Regenerated reasoning is a new judgement wearing an old date.
This is the same storage discipline that separates an auditable agent from an unaccountable one, applied to classifications instead of actions.
Separate the classifier from the consequence
For firms operating across jurisdictions, the same classifier can be an obligation in one regulator’s eyes and a liability in another’s. The EU’s Digital Services Act obliges large platforms to build risk-mitigation classifiers, with penalties up to 6% of global turnover. Other jurisdictions treat aggressive filtering itself as the exposure. A single global policy cannot satisfy both.
The firms navigating this well stopped hunting for the neutral setting and did something duller: they separated the classifier from the consequence. The model may score anything. What the score does — block, flag for review, or merely annotate — is a jurisdiction-scoped policy artefact with its own owner, change log, and legal review.
That split also answers the question that actually gets asked in a hearing. Not “why did your AI think that?” but “who decided what happens next, and under whose authority?”
The decision frame for a UAE firm
If your AI systems produce per-entity judgements that persist (scores, flags, classifications, rejections) and you cannot reconstruct full decision provenance for a case from 24 months ago, stop expanding coverage and fix retention first. The exposure compounds with volume. Every month of expansion buys more of it.
If your judgements are ephemeral or aggregate, such as forecasting, capacity planning, or routing with no durable per-person record, skip the heavy provenance apparatus and spend on the classifier-consequence separation instead. And when commissioning any new classification system, the vendor question is not accuracy. It is: when someone with authority asks who authorised this specific decision, what does your system return?
Provenance is an architecture decision, and it is cheapest before the volume arrives. The five-step consulting method maps which of your AI-derived judgements persist, what each must store to be defensible, and who owns the policy behind it, before a regulator or a claimant asks first. For a governed first implementation on a single process, a clinic day builds the record-keeping in from the start.
Frequently asked questions
What should be stored alongside every AI-derived decision?+
Three things, written once at decision time: the policy version that authorised the decision, the human role accountable for that policy, and the reasoning as it stood at the moment of the call. Reasoning regenerated later is a new judgement wearing an old date, and it will not survive questioning.
Why did GARM shut down, and what does it teach AI governance teams?+
GARM, the advertising industry's brand-safety body, discontinued operations within days of being sued by X in August 2024. Its parent federation said plainly that it lacked the resources to fight the case. No court ruled its taxonomy wrong; the cost of the discovery process ended it. The lesson: you do not need to lose a case to be destroyed by one, you only need to be unable to afford the reconstruction of your own decisions.
Can regenerated AI explanations be used as audit evidence?+
No. A regenerated explanation is produced by a different model state, under a different prompt, at a different time from the decision it claims to explain. Auditors and courts want the contemporaneous record: what was known, decided, and authorised at the moment of the call.

