These are the questions I kept asking while designing this, with the answers I settled on. Some are basic on purpose: the basic ones are the ones that decide whether the design holds. The page grows. Each question carries the date it was added, and notes go at the end.
Questions
What is “a forecast that updates itself”, and where does it live?
asked 6 October 2026 · Where does it live, can several people work on it at once, and how do forecasts for different parts of the company consolidate into one without Anaplan, in Excel or Google Sheets?
The phrase is Friar’s: one of OpenAI’s two ambitions is “a forecast that updates itself”.1 Nothing in this design updates by magic. Three things move without a person touching them, and the fourth is the part a person owns.
- Closed months. When a month’s actuals post, they replace the forecast for that month in the latest-estimate view. Nobody pastes actuals into a workbook.
- Driver lines. Lines whose logic is a rule rather than a judgement are recomputed from records every night: a contract schedule, a run rate from the last three months, headcount times cost, payroll from the last run adjusted for joiners and leavers. The rule lives in a table, not in a cell.
- Submissions. When an owner marks a workbook submitted, that night’s run reads it, stores a version and recalculates the consolidation.
- Judgement lines stay in the owner’s workbook. A sales forecast is a person’s view of the pipeline. The design makes it count the night it is submitted. It never writes it.
Where it lives. Two layers. The workbooks are where people think; the store is where the forecast is kept. Each area has its own workbook with the standard tab: sales, operations, marketing, each cost centre, twelve in the use case. The reader takes the tab and nothing else, so the owner’s model behind it can be as personal as it likes. The store holds every submitted version of every area, keyed by area, version, line and period. The consolidated forecast is a query: the latest submitted version of each area, summed by reporting line and period, with closed months swapped for actuals. The approved baseline is a record in the store, a named set of versions, not a file. The page shows the result, and nobody edits it by hand.
Several people at once. Each owner has their own workbook, so there is no fight over one file. Where two people share an area, Excel in OneDrive or SharePoint lets them edit the same workbook at the same time: “when you co-author, you can see each other’s changes quickly”.2 Google Sheets does the same natively, up to its limit of 100 open tabs or devices on one file.3 The consolidation never has two editors because it has none.
Without Anaplan. In Anaplan the model holds the dimensions (version, line item, time, entity), the import, the calculation and the dashboard in one place. Here the same four jobs are split. The standard tab is the import specification. The mapping table is the line-item hierarchy. The versions table and the compute step are the calculation. The page is the dashboard. That split is what makes it cheap, and it’s why the controls in part 3 matter: Anaplan sells you the model’s discipline with the licence, and here the finance team supplies it. Excel, Google Sheets or a web form all work as the owner’s surface. The only requirement is a fixed layout and a submit flag the reader can see.
The accounting close and the FP&A close: which one is this?
asked 6 October 2026 · This is an FP&A close activity. It can’t be an accounting close. What does each one own?
This is an FP&A close. I call it a close because it has a cut-off and a sign-off too: the point at which the management numbers for the month are agreed and the explanations are signed. It is not the accounting close and it cannot be. The accounting close produces the books; this consumes them.
| The accounting close | The FP&A close | |
|---|---|---|
| Purpose | The books: a true record of the period | Run the business: budget, forecast, variances, explanations |
| Who relies on it | Outsiders: auditors, lenders, the tax authority, the board as a legal body | Insiders: the CFO, budget holders, the planning cycle |
| What counts as a number | Posted, reconciled, final | Posted, or estimated and labelled as such |
| Cadence | Monthly, with a quarterly and annual hard stop | Every night, agreed weekly, closed with the accounting month |
| Who signs | The controller; the CFO; the auditors on the year | The FP&A lead on explanations; the controller on the baseline |
| What it produces | Journals, reconciliations, the trial balance, the statements | Variances, the latest estimate, the forecast versions, the memos |
| What it never does | Forecast | Post a journal |
The two meet twice a day and once a month. Every night the FP&A side reads the ledger as it stands and adds its labelled estimates on top. At month end the accounting side posts the accruals, the allocations, the revaluation and the recognition run, and those finals replace the estimates. The gap between estimate and final is the FP&A side’s trust score, and the things the nightly checks turn up (an unmapped account, a closed period that moved, a trial balance that doesn’t tie) are the accounting side’s early-warning list. Neither side does the other’s job. Part 1 says this once; part 2 shows what lands at month end and why.
If the budget is monthly, why load the actuals every night?
asked 6 October 2026, sharpened 7 October · The budget has month-level granularity. What does a daily load of actuals add?
A monthly budget doesn’t ask for daily actuals, and nothing in the design needs the actuals more often than the decisions do. Nightly is a choice. Here is what it buys when the budget is one number a month, and where it buys nothing.
The comparison is month to date against the budget phased to date. The budget stays a monthly number. For the comparison on the 12th, that number is phased: straight-line by calendar day for lines that accrue evenly, by working day for lines that follow activity, by a profile for lines with a known shape (last year’s pattern, a known event), and by the schedule for lines that land on a date, like rent on the 1st or a quarterly licence. A line is flagged only when its month-to-date figure, posted plus estimated, sits outside a threshold around the phased figure. Lines that are lumpy by nature are marked “compare at month end” and never flag mid-month. Comparing month-to-date actuals with the full-month budget is the mistake this avoids: do that and everything looks under budget until the last week.
So why nightly rather than weekly, or at month end? Not for the budget’s sake. For the state of the month: what’s posted, what’s estimated, what’s still waiting. That state changes every day because invoices post every day, and each night an estimate either becomes a fact or doesn’t.
- An error is a day old when you find it. A duplicate posting, a mis-coded invoice, an export that dropped rows: the checks catch it the first night it appears, and the person who knows why is still reachable.
- The forecast change on the 7th is on the page on the 8th. Submissions are picked up when the folder is read. A weekly cadence would make the contract wait a week.
- Exceptions arrive one or two a night, not thirty on close day. Unmapped accounts, rejected workbooks and moved periods get decided while they’re small.
- The estimates sharpen with each invoice. By the last night, most of what accounting will accrue has already been estimated from receipts.
- It costs nothing more. A scheduled job runs at 02:00 whether it runs once a month or every night, and the controls are the same either way.
Where nightly adds nothing. A line whose actuals land in one posting gains nothing from being loaded on the 12th: payroll on the 30th, rent on the 1st, a quarterly invoice. It shows “waiting” until its date, and that’s correct. The variance view earns its keep on the lines that accrue through the month: cost of sales, freight, contractors, travel, anything with a purchase order behind it. If every line in your P&L were lumpy, weekly would do and the design would work unchanged.
The number that matters is still the monthly one. The mid-month page is provisional by construction. The month-end comparison, after the finals replace the estimates, is the one the FP&A lead signs. Nightly actuals make that signature faster and the month quieter; they don’t replace it. The figure below is what the month looks like when the posted share grows every night.
What is an exception queue, and what does its owner do?
asked 6 October 2026 · How does someone know an exception they’re responsible for has been raised: the web app, email or Slack? And then what?
An exception is something a check or a rule found that the schedule can’t decide for itself. An account with no mapping row. A workbook whose layout the reader doesn’t recognise. A closed period whose total moved after a late journal. A line whose estimate is weak because the purchase orders have no receipts. A bank item unmatched for more than seven days. A drafted explanation whose citation didn’t check out. The queue is the list of open exceptions in the store, and each carries what it is, when it was raised, the owner (a role, set by a routing table), its age, the evidence, a suggested action where the model has one, and its state: open, decided, cleared.
How the owner finds out. Three ways, and they’re the same for every exception.
- The queue on the console. The “Your queue” tab, by role, with the actions on the row (try it).
- A morning digest. At 06:30 the owner gets one message, by email or as a card in Slack or Teams: how many items, the oldest, each with a link to the item and its evidence. One message a day, not one per exception.
- Escalation. A reminder at two days. At five, the item appears on the controller’s queue as well. Nothing crosses a month boundary without a waiver and a reason. A failed run is the one thing that alerts immediately, because the page is stale until it’s fixed.
An example, end to end. Monday, 02:14: the mapping step finds £3,120 posted to account 61255, “Data subscriptions”, which has no row in the mapping table. The amount goes on the Unmapped line so the totals still tie, and an exception is raised. The routing table says new accounts go to the analyst. The model looks at the three accounts with similar names and suggests Software; the suggestion is attached to the item, not acted on. 06:30: the analyst’s digest lists three items, this one newest, with the suggestion and the three transactions behind the £3,120. Tuesday, 10:20: the analyst opens the item, agrees with Software, and records the reason (“matches 61240, cloud hosting”). That click writes a mapping row effective 1 October, mapping version 24, and a log entry with the actor, the time and the reason. Wednesday, 02:14: the run applies version 24, the £3,120 leaves the Unmapped line, and the exception clears on evidence: nobody ticks it, the next run sees it resolved. If the analyst had done nothing, Wednesday’s digest would have carried a reminder and Saturday’s would have put it on the controller’s list.
Part 3 draws the three ways an item clears and the gates. The appendix has the sixteen failure modes that feed the queue.
Posted, estimated, final: deterministic or agentic?
asked 6 October 2026 · Does someone have to go through every transaction and tag it?
Deterministic, all three, and nobody tags transactions.
Posted is whatever the ledger holds for the period, mapped to a reporting line by the account and department it was posted to. The mapping is a row per account, set once by finance, not a label per transaction. A transaction posted to the wrong account is accounting’s problem, and it shows up as a variance, which is the right place for it to show up.
Estimated is computed by a rule per line from records finance already keeps. Goods received and not invoiced is a query over the purchase order and receipt tables. A contract line is the monthly amount from the contract register. Payroll is the last run adjusted for joiners and leavers from the HR extract. A run-rate line is the average of the last three months’ postings, if the owner chose that rule. The rules live in a small table: line, method, source records, parameters, what it is checked against at month end. Finance sets it, versions it, and changes it the way it changes the mapping, with a second reviewer.
Final is what accounting posts at month end. It replaces the estimate, and the gap between the two is recorded line by line.
The model has no part in any of the three. It may suggest a method for a new line, or notice in a draft explanation that a line’s estimate gap has been growing, but the numbers come from the rules and the rules change only by a person’s hand. That is the whole reason the layers can be trusted: an estimate built by a rule can be re-performed by an auditor, compared with the final, and improved when the gap says so. An estimate built by a model’s judgement each night can’t be any of those things.
The one place that looks like tagging isn’t. Purchase orders and receipts already carry an account, so the estimate for a line is the sum over the receipts whose account maps to that line. The work was done when the order was raised. Part 2 has the chain from purchase order to accrual.
What does “idempotent” mean here?
asked 6 October 2026 · I keep hearing the word. How and where does it apply in this design?
A step is idempotent when running it twice with the same inputs leaves the same state as running it once. Stripe’s API reference puts it plainly: idempotency is “for safely retrying requests without accidentally performing the same operation twice”.4 It is a cousin of determinism, which is about getting the same answer. Idempotence is about a retry leaving nothing extra behind: no duplicate rows, no double count, no second approval.
In this design it applies in five places, and each has a key that makes it true.
- The ledger load. Every open period is pulled again each night. The load replaces that period’s rows, keyed by transaction id, rather than appending them, so a re-pull after a timeout can’t double-count a day.
- The workbook reader. A submitted file is fingerprinted. The same fingerprint stores no new version; a changed one does. Opening and saving a workbook without changing a cell changes the file’s fingerprint but not the values’, which is why there are two.
- The computed tables. Estimates, variances and version differences are rebuilt from scratch for each run and keyed by run id. Rerunning Tuesday gives Tuesday’s numbers, which is what an auditor’s reperformance needs.
- Approvals and alerts. An approval click can arrive twice (a double tap, a retried webhook). It carries an event id and a one-time token, so the baseline moves once. The email link works once. The “no run for a day and a half” alert fires once, not every minute.
- The log. Events are the one thing that isn’t naturally idempotent, so each carries a key made of the run, the step and the object, and a retried step can’t log the same event twice.
Why it matters for a finance team more than for a software team: retries are the normal case, not the edge case. Extracts time out, files are locked, a chat platform redelivers a message. If every step is safe to repeat, the schedule can retry plumbing failures on its own and a person is only called when a check fails. And the audit trail stays clean, because a rerun doesn’t leave a trail of duplicates for someone to explain.
How to know it’s true: the twice test. In the test month, every step runs twice, and the second run must change nothing. It’s one of the cheapest tests to write and the one that catches the most. Part 3 has the processes this applies to.
Notes
Later notes go here, each with its date.
This is part of a series explaining AI and the systems around it for finance people, in their own language. I build AI systems for finance teams; the series is what I’ve learned doing it. This one is a design on paper, not a system I have run.
Sources
Footnotes
-
Sarah Friar, “What building an AI-native finance function taught me”, OpenAI, 10 August 2026: https://openai.com/index/building-an-ai-native-finance-function/ ↩
-
Google Docs Editors Help, “Share files from Google Drive”: https://support.google.com/docs/answer/2494822. “A Google Docs, Sheets, Slides or Vids file can only be edited on up to 100 open tabs or devices.” ↩
-
Stripe API reference, “Idempotent requests”: https://docs.stripe.com/api/idempotent_requests. Used for the definition only; nothing in the design uses Stripe. ↩
