ARKONE
Two mail slots side by side in an office door, one marked with a lock

Govern What Your AI Agents Send: Two Filters, Not One

August 9, 2026 · 5 min read

S
Sobin George Thomas

Producing text and code has become close to free; judging it has not. Firms that fuse integrity control with expressive judgement slow the first to protect the second, right as their own agents start sending more than their staff do.

Producing material, whether text or code, has become close to free. Judging it has not. That one asymmetry now sits underneath two problems that look unrelated: malware that rewrites itself on every execution, and the flood of agent-drafted emails, tickets, and commits leaving your own systems unread.

In 2025, ESET researchers identified PromptLock, the first known AI-powered ransomware: malware that calls a language model at runtime and generates its malicious scripts on the fly. It turned out to be an academic proof-of-concept, never observed in a live attack. Treating that as reassurance is the error. The relevant fact is the marginal cost of a new variant, which has fallen to roughly the price of an inference call.


The attacker’s code stopped being a durable object

Antivirus economics rested on a quiet assumption: writing genuinely novel malware is expensive, so attackers reuse code, so families of samples share signatures, so one analysis defends millions of endpoints. Detection scaled better than attack. That was the whole game.

Runtime-generated malware breaks the reuse. When a sample regenerates its own loader on each run, there is no family to signature. The first specimens are clumsy; the trajectory is what matters.

The defensive implication cuts against most current security budgets. If signatures decay, the money moves to behaviour: what a process actually does to the file system, the network, the credential store. Behavioural detection was always the more expensive, more false-positive-prone option, which is why it stayed a supplement. It is now the primary line, and in most firms the budget has not moved yet.


Your own agents are the volume problem arriving early

The same collapse in generation cost applies to material your own systems produce. A firm running agents that draft customer emails, file support tickets, or commit code has industrialised the production of artefacts nobody reads before dispatch. The volume argument that made universal human review impossible for social platforms is now arriving inside enterprises with 400 staff.

This is not a reason to slow the agents down. It is a reason to know exactly which checks run on everything automatically and which decisions still need a named human. Firms that cannot make that distinction end up with one committee reviewing both, which means neither gets reviewed well.


Two capabilities were fused for convenience

The strategic error is treating “filtering” as one capability. It is two, with opposite risk profiles.

Integrity control decides whether an artefact is authentic, safe to execute, and originated where it claims: blocking a self-modifying binary, rejecting a spoofed invoice, quarantining an agent’s outbound API call. This is engineering. It carries almost no political or legal exposure, and the case for expanding it has just strengthened considerably.

Expressive judgement decides whether a permitted, authentic message should be carried: what your firm is willing to say, to whom, in which jurisdiction. This is where the contested exposure lives.

A retail bank illustrates the split cleanly. Its policy on which words a relationship manager may put in a client email is expressive judgement — contestable, jurisdiction-specific, worth minimising. Its policy on whether an AI assistant may execute a payment instruction it drafted itself is integrity control — technical, absolute, worth expanding. Most institutions run both through the same governance committee, so the second gets slowed by the sensitivity of the first.


Separate them in the org chart, the audit trail, and the budget

The operating rule: integrity control should be able to ship a new rule in a day. Expressive judgement should not be able to ship one without a written record of who decided and on what basis.

That second requirement is the same provenance discipline that makes every stored AI judgement defensible: a policy version, an accountable role, contemporaneous reasoning. Integrity rules need none of that ceremony, which is exactly why they should not be queued behind it.


Where the first exposure actually sits for a UAE firm

The threat that reaches most firms first is not novel ransomware. It is their own automation acting on a poisoned instruction: an agent with payment access reading a spoofed email, an outbound integration pointed at an attacker’s endpoint. The DFSA’s finding that 21% of DIFC firms using AI in critical processes lack oversight describes exactly this gap: capability deployed ahead of the integrity layer around it.

If your security spend is still weighted toward signature detection, move the increment to behavioural and identity controls this budget cycle, and point it specifically at the outbound behaviour of your own agents. Scoped permissions, spend limits, destination allow-lists, and an immutable log cover most of the surface, and none of them slows the agent down.


An agent built with the integrity layer designed in is not slower to ship; it is faster to approve. A clinic day builds a working agent on one of your processes with scoped outbound, logged actions, and human sign-off on the irreversible steps, by evening. The governance thinking behind those controls is in why AI agents cheat, and the governance that actually holds.

Frequently asked questions

What is the difference between integrity control and expressive judgement?+

Integrity control decides whether an artefact is authentic, safe to execute, and originated where it claims: blocking a self-modifying binary, rejecting a spoofed invoice, quarantining an agent's outbound API call. Expressive judgement decides whether a permitted, authentic message should be carried. The first is engineering with little exposure; the second carries legal and reputational exposure in every jurisdiction. Fusing them slows the safe one to protect the contested one.

Why is signature-based malware detection losing ground to behavioural detection?+

Signature economics assumed writing novel malware was expensive, so attackers reused code and one analysis defended millions of endpoints. ESET's PromptLock proof-of-concept showed malware can now call a language model at runtime and regenerate its own scripts on each execution, leaving no stable family to signature. When variants cost an API call, the durable signal is behaviour: what a process does to files, networks, and credentials.

Do AI agent outputs need review before they are sent?+

Volume makes universal human review impossible, which is the point of splitting the filters. Integrity checks (authenticity, permissions, spend limits, destination allow-lists) run automatically on everything. Human judgement is reserved for the expressive layer and for actions that are expensive to reverse.

ഒരു സംഭാഷണം ആരംഭിക്കാൻ തയ്യാറോ?

അളക്കാവുന്ന AI ഫലങ്ങൾ ഡെലിവർ ചെയ്യാൻ ArkOne എങ്ങനെ ഭരണ സംവിധാനവും പ്രോഗ്രാം ആർക്കിടെക്ചറും നിർമ്മിക്കുന്നു എന്ന് കാണൂ.

ഡിസ്‌കവറി കോൾ ബുക്ക് ചെയ്യുക