The practical model for defensible AI use
The most common AI mistake in audit, risk, and controls is not usually the tool.
It is the standard.
Too many teams are treating a polished AI output as if it has already earned the right to influence a decision.
It has not.
Not until it can survive challenge.
That matters more now because boards and senior leaders do not need more prose. They need proof. At the same time, AI makes it much easier to generate volume. More summaries. More draft findings. More issue logs. More management wording. More slides. More “insight”.
So the problem changes.
It is no longer just, “Can we produce something quickly?”
It is, “Can we tell the difference between speed and substance before the output shapes a decision?”
That is the trust gap many teams are walking into.
AI increases volume faster than most governance processes can absorb it. If you do not tighten the mechanics, the result is predictable:
more polished wording
more hidden assumptions
more unsourced claims
more review pain later
less confidence in the output that actually matters
That is why this week is about AI proof, not AI prompts.
And why the operating model matters more than the tool.
Download my free AI defensibility pack here: AI Defensibility Pack.pdf
The operating model: The Evidence Chain
When an AI assisted output matters, I make sure to push it through six questions first:
1. Claim
What exactly is being said?
One sentence. Clear enough to challenge. Narrow enough to test.
2. Evidence
What supports the claim?
Numbers, records, logs, screenshots, documents, extracts, reconciliations, approvals.
3. Source
Where did that evidence come from?
Named systems. Dated reports. Policy versions. Workflow exports. Meeting records. Not “from the file” or “from the team”.
4. Assumption
What was inferred, estimated, or filled in?
This is where AI often creates false comfort. It bridges gaps unless you force those gaps into the open.
5. Risk
What could be wrong, missing, or overstated?
Data lag. Mapping errors. Missing populations. Weak reason codes. Excluded items. Bias in the prompt or in the reviewer.
6. Decision
What changes because of this?
Review, escalate, approve with conditions, monitor, redesign, hold, investigate. The output should move something.
That is the Evidence Chain.
And it matters because most weak AI use breaks in the same place. The wording looks better than the underlying proof.
When that happens, people argue about phrasing when they should be challenging the chain.
Why this works
The model is simple, but it does something useful.
It separates three things that often get mashed together:
what the AI drafted
what the evidence supports
what the human is prepared to stand behind
That separation is where defensibility lives.
Without it, the process becomes vague. The draft gets improved, the tone gets softened, the wording gets “made more strategic”, and suddenly nobody is clear on whether the underlying point was ever strong enough to begin with.
That is how weak output gets promoted.
A concrete example: O2C credit notes
Let’s take a realistic case.
An AI tool is used to summarise quarter end credit note activity and flag potential control concerns.
Before
The raw AI output says:
“Credit notes appear elevated in Q2 and may indicate pricing control weakness or inconsistent approval discipline.”
Plausible.
Polished.
Not decision ready.
There is no clear threshold.
No named source.
No boundary on what “elevated” means.
No clue what was excluded.
No clear owner or next step.
After
Now run the same topic through the Evidence Chain.
Claim
Credit notes above the agreed threshold increased in Q2 and were concentrated in three customers and two sales approvers.
Evidence
ERP extract of all Q2 credit notes above threshold.
Pricing override workflow log.
Approval history for the affected customers.
Policy threshold reference.
Source
SAP report run on 4 March.
Workflow export from approval tool.
Pricing policy version 6.
Customer account notes from CRM.
Assumption
Returned goods and freight adjustments were excluded based on reason codes.
Manual journal reclasses were not included.
Customer level concentration assumes consistent coding across business units.
Risk
If reason codes are miscoded, the volume may be overstated.
If manual adjustments sit outside the workflow log, the approver pattern may be incomplete.
Decision
Sales operations manager to review top three customer accounts and override practice by 31 March.
Financial controller to test reason code integrity on a sample of Q2 items.
Control owner / process owner decides whether control redesign is needed and owns the remediation plan. Risk & Controls team reviews whether the proposed response is adequate, challenges if needed, and decides whether further assurance or escalation is required.
Now you have something very different.
Not because the AI was banned.
Because the output was forced into a structure that exposes weak logic before it hits a real decision maker.
That is the move.
The defensibility checklist
Here is the checklist I would use before an AI assisted output is escalated, relied on, or recorded as a substantive view.
If 3 or more checks fail, do not escalate the output yet.
Rework the evidence chain first.
What to record
This is the part many teams skip because it feels administrative.
It is also the part that saves you later.
If AI is used for something that could influence judgement, record enough to support challenge, QA, and future review.
At a minimum, I would capture:
use case
tool used
prompt or task instruction version
source set used
date run
key assumptions
thresholds applied
reviewer name
decision owner
final action taken
link to supporting evidence
link to final output
That record does not need to be theatrical.
But it does need to exist.
Because the moment somebody asks, “What was this based on?” you need something better than memory.
Thresholds and materiality: how to avoid “everything is high risk”
This is where teams can accidentally create the opposite problem.
Once people get nervous about AI, they start treating everything as if it needs maximum challenge.
That is not mature governance either.
Use thresholds.
Three simple rules help:
Rule 1: Low impact plus high confidence can move quickly. Routine drafting. Early planning support. Low stakes summaries. Fine.
Rule 2: High impact plus low confidence should stop. Do not escalate it. Do not present it upward. Do not disguise uncertainty with better wording.
Rule 3: High impact plus high confidence still needs clear ownership. Even strong evidence does not remove judgement. It just improves it.
A useful threshold set usually includes:
financial threshold
control failure threshold
regulatory threshold
reputational threshold
stakeholder sensitivity threshold
The aim is not to pretend precision where none exists.
It is to stop “interesting” from being treated the same as “material”.
IRM guidance treats risk appetite and tolerance as board and management matters that should be clearly articulated and translated into qualitative and quantitative statements, so they can support meaningful decision making. The IIA’s current Global Internal Audit Standards similarly emphasise independence, objectivity, documented engagement work, and relevant, reliable, and sufficient information to support findings and conclusions. That is why threshold setting matters here. It gives AI assisted output a proper decision context instead of letting everything sit at the same level.
A short note on good practice
This is not about claiming that one checklist makes you compliant.
It does not.
What it does do is keep your operating model aligned to common good practice.
That means:
human review over material judgement
sufficient and traceable support for conclusions
clear documentation
independence and objectivity in review
quality assurance that is proportionate to impact
awareness of risk appetite and tolerance when decisions are made
That is a much more useful posture than pretending the answer lives in the prompt.
The real point
AI is useful.
But in governance work, usefulness is not the standard.
Defensibility is.
Because when the challenge comes, nobody gets protected by the fact that the summary was well written.
What matters is whether the logic, evidence, thresholds, ownership, and record are strong enough to hold.
That is the week’s point in one line:
AI drafts.
Humans validate.
Evidence earns the right to influence a decision.
If you want the full templates, worked examples, checklist, challenge questions, and governance workflow, download the AI Defensibility Pack here:
And if you want more people to receive practical tools and routines like this each week, share this edition and ask them to subscribe here:
Best,
Founder Beyond the Lines™ | Integral Assurance
Your Tax Data, Finally in One Place

Are you tired of hunting down data, fixing errors, and manually updating disconnected spreadsheets?
Tax reporting isn’t a simple as it used to be. You need real-time, flexible reporting so you can confidently make decisions backed by accurate, centralized data.
Learn how bringing all your tax information into one central system automates repetitive tasks, improves scenario planning, and frees your team to focus on strategy instead of data entry.
Whether you operate in one country or dozens, Longview Tax scales with you—reducing risk, speeding up your close process, and helping you optimize tax policies across all jurisdictions.




