Someone asks why R&D moved against plan. You open the pack. Then the model behind the pack. Then the CRO contract behind the accrual line. Then you check that the three agree. That’s the afternoon. And it’s why, most months, the question doesn’t get asked at all.
Almost half of a finance team’s effort goes on this. The FP&A Trends Survey 2026 puts it at 47% of effort spent collecting and validating data, before any of it gets analysed.
AI hasn’t changed that, and the reason is simple. A number you can’t trace is a number you have to redo. If the tool gives you a figure with no page or cell behind it, you’re back in the pack doing the afternoon anyway, plus five minutes to read the answer you didn’t trust.
What follows is the entire pitch for what I’ve built, asterisk included, so you can pick it apart.
Whose afternoon it is
Finance teams whose numbers face a board, an auditor or a regulator every month, and who are small compared with the reporting they produce. The CFO or FD owns the risk of a wrong number in front of the board. The FP&A manager or analyst owns the afternoon.
The sharpest version I know is clinical-stage biotech, because I built the planning models for one. The month is spent chasing numbers that live in different places: the CRO accrual in a spreadsheet, spend by programme and phase in the model, cash and runway in the plan, and three versions of the forecast where only one is the one the board approved. Reforecasts are constant, scrutiny is intense, and the finance team is a handful of people.
The cost isn’t the afternoon
Time is the obvious cost. Half of a small team’s month is, in effect, one full-time head doing lookups. The worse cost is quieter.
When an answer costs an afternoon, questions get rationed. The board pack questions get asked. The “hang on, why did that move?” ones wait for a quieter week, and the quieter week doesn’t come. Nobody’s at fault here, that is what retrieval by hand costs. The number that hurts is the one that looked fine and nobody had the afternoon to open.
In biotech the two numbers that matter most are the trial accrual and the runway, and they’re exactly the ones nobody will take on trust. Most of the “why did that move” questions wait until the reforecast, because each one costs someone a day. So the runway on the slide is last month’s, and the accrual rolls forward. A runway figure that’s wrong in front of a board or an investor is not a rounding error.
flowchart TB
subgraph T[Traced: a minute]
direction TB
B1[Same question] --> B2[The answer, with the cell it came from]
B2 --> B3[One click to check]
end
subgraph H[By hand: an afternoon]
direction TB
A1[Someone asks why a number moved] --> A2[Open the pack]
A2 --> A3[Open the model behind it]
A3 --> A4[Open the contract behind that]
A4 --> A5[Check all three agree]
end
What people do instead, and where each one breaks
Four things, in the order most teams try them.
By hand. Open, cross-check, repeat. It works, it’s slow, and it lives in one analyst’s head. It’s also why the questions get rationed.
Search and shared drives. They find files. You need figures.
Planning tools and BI. Anaplan, the ERP reports, the dashboards. Excellent for the structured data they were built around, and I’ve spent years building them. But the retrieval pain lives in everything around them: the pack, the contracts, the board deck, last quarter’s reforecast memo.
Upload it to ChatGPT, Claude or Copilot. Brilliant for one file, one question, once. A finance function isn’t that. It’s hundreds of workbooks, questions that recur every month, and answers that have to be right, checkable and available to the whole team. Drop a workbook into a chat and it reads it as flattened text, so you get a plausible number rather than the number, with no cell to click through to, no way to know if it read the March forecast or the budget, and your board deck now sitting in someone’s personal chat account. You find out whether to trust it by getting burned.
For a one-off question on one spreadsheet, honestly, just use Claude. Same brain, different job. The model isn’t the problem. The harness around it is.
Put the four side by side and the pattern is the same each time: the tool finds something, and the checking is left to you.
| What teams try | Where it breaks |
|---|---|
| By hand | Slow, lives in one head, questions get rationed |
| Search, shared drives | Finds files, not figures |
| Planning tools, BI | Structured data only; the pack, contracts and decks sit outside |
| Upload to a chatbot | One file, once; flattened numbers; no cell to check; version-blind; ungoverned |
What changed this year
The models got genuinely good at reading finance documents. I tested this rather than assumed it. On a realistic board pack, raw Claude and raw GPT get most of the questions right and catch every trap I buried, and anyone telling you otherwise is selling something.
So the model is no longer the problem. What became possible is the layer around it: reading spreadsheets cell by cell rather than as prose, keeping versions apart, and tying every figure in an answer back to the exact page, cell or clause it came from. The checking, which used to be the analyst’s afternoon, can now be built into the tool. And it can run inside your walls, on your enterprise keys or a private single-tenant instance, at a price a mid-market finance team can say yes to.
What I built, and how it behaves
Grounded reads the packs, models, contracts and decks your team already produces, and keeps them current each month-end. The files stay the source of truth. It reads spreadsheets positionally, keeps Budget, Actuals and the March forecast apart, and when two files carry the same figure it prefers the authoritative one: the reissued FINAL over the superseded original.
You ask in plain language. Every figure in the answer links to the document page, spreadsheet cell or contract clause it came from, one click to check. If a number isn’t in your documents it says so, and it doesn’t quote the numbers that are. It ships behind a golden question set your team writes and signs off, on your instance, with everything logged.
Then, because every answer is traceable, you can build on it. Checks that run every close: variance flags, accrued against invoiced by study, vendor invoices over the contracted rate, open POs not yet in the ledger, runway refreshed on the latest forecast against the version the board saw. Each one shows its working. Your team stops chasing numbers and starts signing them off.
How I know, and the asterisk
I published a test. On the seventh of August I built a company that doesn’t exist, Caldergate Distribution Group, and gave it 34 documents of finance reporting: workbooks, PDFs, two board decks, two CSV extracts, with three problems buried on purpose. Then 25 questions, a cell-level answer key and a script that grades. It’s called the Board Pack Test and it’s public, answer key included.
81 of the 100 points can be verified by a machine with no human in the loop. Raw Claude Opus 5 banked 66 of those 81. Raw GPT-5.5, 71. The first version of Grounded banked 66 and 68, roughly raw-model territory. After three changes (refusals carry no figures, prefer the authoritative file, cite the exact cell) it banked all 81, three runs out of three, on the same underlying models.
The asterisk. If a human grades everything and reads every refusal charitably, raw Claude scores 100 and my tool scores 99. The models are equally smart. The gap is whether the answer can be proven right without a marker, and at your desk there is no marker. Nobody re-grades the AI’s answer against a key before it goes in the board pack. (If you remember one thing from this post, I’d want it to be that.)
Everything is in the repo: documents, questions, grader, every transcript from every run, including the runs where I lost. Then the version that counts: your pack, your 25 questions, success criteria your team writes. That’s the pilot.
The qualitative part is shorter. I spent about fifteen years in finance and FP&A operations and built the planning models for a clinical-stage biotech, so the afternoon in question was mine. I publish where the tool loses as well as where it wins. And it’s the thing your team is already doing ungoverned, made trustworthy, permanent and safe enough for the board pack.
What it costs
A fraction of one analyst’s year, and I’d rather you tested that than took it. The dedicated finance AI tools are priced for listed companies. This is the bootstrapped version, and it starts with a pilot: your pack in your own instance, your golden question set validated on your real numbers, success criteria you write before the pilot starts, around thirty days, fixed fee, credited in full if you continue. Miss the criteria and you hear it from me first.
What I’d tell your CFO: almost half the month goes on finding and checking numbers, and AI only fixes that when every figure shows the cell it came from. A number you can’t trace is a number you have to redo. The chasing goes, the sign-off stays yours.
Bring your own pack
Don’t take any of this from me. The Board Pack Test has the documents, the questions, the answer key, the grader and every transcript, so run it against whatever you’re being sold. Better still, ignore my fictitious company and bring your own pack with your own 25 questions, the ones your team already knows the answers to. That’s how I onboard pilots, and it’s a better exam than mine, because it’s yours.