Skip to content
CodelessOps
Go back

Citations aren't verification

Karim Lameer

Karim Lameer — Master Anaplanner, CIMA-qualified, 15 years in finance and FP&A. I build Grounded and published the Board Pack Test. LinkedIn · New here? Start here

The demo is going well. You ask the AI why consulting revenue fell in May, and it answers in three tidy sentences with a number in bold. Under the number sits a small grey footnote: the file name, the page. Everyone in the room relaxes a little. The footnote did that.

Then someone clicks it.

Most of the time, the source opens and the number is there, and the relief was justified. But notice what actually earned your trust in that moment. It wasn’t the answer. It was the click. Which raises an uncomfortable question about the other forty numbers you didn’t click.

A citation and a verification are different controls

Finance already has this distinction, it just hasn’t applied it to AI yet. A citation says the invoice is attached. A verification says the invoice was matched. One gives you something to check. The other is the check.

When an AI tool cites a source, it’s telling you where a number claims to come from. That’s genuinely useful, the way an attached invoice is useful. But in most tools the model writes the citation itself, in the same breath as the answer. The same statistical machine that produced the number produces the footnote under it. Nothing enforces the connection between the two. The reference can be real while the number beside it is rounded, stale, or from the wrong period entirely.

Verification is a different kind of promise. It means somebody, or something, actually did the matching. In practice that means code, because code is the only participant that can’t be talked into a plausible answer.

I think of it as a chain with four links: your documents, then extraction, then computation, then the answer. Call it a trust chain. A number is verified when every link between the source and the sentence held.

flowchart TB
  A[Your documents] --> B[Extraction<br/>the right cell, kept intact]
  B --> C[Computation<br/>done by code, formula kept]
  C --> D[Answer<br/>every figure checked]
  D -->|any link fails| E[Held back<br/>instead of stated]

The two questions that separate them

You can test any AI tool on this in one conversation, with two questions.

Who wrote the citation? If the model composed it as part of its answer, it’s decoration. It will usually be right, because the model usually read the source. But usually is doing a lot of work in that sentence. If the system captured the citation from what was actually retrieved, separately from the model’s prose, then the footnote is evidence rather than good manners.

Who did the arithmetic? This is the one nobody asks. Say the answer gives you a variance: revenue was 236.5 against a program total of 234.4, a gap of 2.1. The two inputs might both be perfectly cited. The 2.1 came from neither document. The model computed it, and language models compute the way a tired analyst does mental maths at 6pm, confidently and sometimes wrongly. A citation on the inputs does not cover the output. Unless the arithmetic ran through actual code, a calculator rather than a next-word predictor, the derived number is the model’s opinion.

I watched a talk recently by the founder of Kepler, a company that builds AI for hedge funds, and one line has stayed with me since. They trained a model that extracted financial figures with 94% accuracy, well above the general-purpose systems. His own verdict: who would trade off something that’s 94% accurate? The wrong number is still wrong if you’re in the unfortunate 6%. Anthropic later published a case study on how they rebuilt around that idea, and the fix was taking the numbers away from the model rather than training a better one.

What is verifiable AI, then?

It’s the name this discipline is starting to acquire. An AI system is verifiable when it can show you, for any figure it states, either the exact place in your documents the figure was traced to, or the calculation code performed to produce it from figures that were. Both halves of the trust chain, on demand, for every number. Citation tells you the story of a number. Verification lets you audit it.

What this looks like in your world

The reason this matters isn’t philosophical. It’s the board pack.

A cited-but-wrong number survives review precisely because it’s cited. The footnote lowers everyone’s guard, the deadline does the rest, and the figure travels: into the deck, into the commentary, into the model the next quarter gets built on. Nobody catches it, and that is the entire failure mode.

If you’ve read my piece on uploading spreadsheets to ChatGPT, this is the same story one level deeper. There the problem was whether the tool could find the right number. Here the problem is that it found the right numbers, did its own arithmetic on them, and nobody thought to ask who checked the sum. And because of how these systems read spreadsheets in the first place, the inputs themselves deserve more suspicion than a tidy footnote invites.

The controls, in your language

None of this means the technology is unusable. It means it needs the controls you’d apply to a new junior analyst, and they translate directly.

Click-through is your spot check. If a tool’s citations can’t be clicked through to the actual cell or passage, treat every number as unreviewed work. Ask in the demo: show me where that figure lives in the file.

Code-done arithmetic is your segregation of duties. The entity that writes the narrative should not be the entity that computes the numbers. Ask: when the answer contains a calculation, what performed it? “The model is very good at maths” is the wrong answer.

Refusal is your completeness control. A system that can’t say “I can’t confirm that” will fill every gap with something fluent. The most trustworthy behaviour an AI can show you is declining to answer, visibly, with a reason. If you never see it refuse, no one has tested where its knowledge ends.

And a standing test set is your reconciliation. Write down the ten questions your team must never get wrong, with the answers, and re-run them whenever the tool changes. Quality becomes a number you track rather than a feeling you have.

What I’d tell your CFO

A citation tells you where a number says it came from. A verification tells you the number is right. The first is a footnote; the second is a control. Any AI tool you put near the board pack should be able to show you, for every figure, either the cell it was traced to or the formula code used to compute it, and it should refuse to answer rather than state a number that has neither. Buy the second thing. Enjoy the first.


This is part of a series explaining AI and the systems around it for finance people, in their own language. I build AI systems for finance teams; the series is what I’ve learned doing it.


Share this post:

Keep reading

All posts →