02:14. Nobody is in the office, and nine things are happening in the close.
The ledger has been pulled and tied to the trial balance, account by account. Twelve forecast workbooks have been read; one was rejected because its owner saved it without recalculating, and the page will keep last week’s version of that one and say so. The estimates for what hasn’t posted have been recomputed from this week’s receipts. Variances have been flagged on two lines, and a model has drafted an explanation for each, with citations a program has already checked. The Sales file has been stored as version 9 and compared with the baseline. Seven audience files have been written, one per department. One alert has cleared. By 02:20 the log has 140 new lines and the page is current.
Everything in this design is either a small program or a decision. The architecture is how the two meet. This part draws it: what it runs on, how the processes are orchestrated, the log nothing is ever deleted from, where people sit, how approvals reach them, and what the console is. Then what it costs to own.
What it runs on
It runs on six ordinary things: a scheduler, one database, the shared folder you have already, a small web page behind your company sign-in, a language model reached over the internet, and Slack, Teams or email, whichever your team lives in. None of it needs a licence beyond what the company has, and nothing needs the ERP vendor’s approval except a read-only role with sight of every account in scope.
An AI coding assistant writes the small programs. A finance person who knows the process directs it, and tests the result against a month with known answers. IT owns where the database and the page live, the sign-in, and the backups. Those three sentences are the whole division of labour.
Data goes out and events come back. Out: ledger extracts nightly, the workbooks on submission, purchase orders and receipts, the bank feed, the pay run. Back: every approval, signature and decision, as an event in the log. The products named in the appendix, NetSuite’s query interface and its 100,000-row cap among them, are examples of the plumbing, not endorsements.1
Small processes, orchestrated centrally
One process, anatomised
Every process has the same shape: inputs, each with a fingerprint; checks; outputs; and a run record holding the run number, the code version, the mapping version, the inputs, the checks and the decision to publish. The ledger load has that shape, and so do the workbook reader, the estimate of goods received and not invoiced, the bank match, the split of outputs by audience and the drafting step. A new process is a new instance of the shape, not a new design.
Fixed, tested code at every step except where the model is called, and the model’s output is an input to a check, never to the ledger.
The night’s order
The scheduler is the one thing that knows what depends on what: extracts first, then the loads and their checks, mapping, the estimates and the arithmetic, the explanations, publish, and finally the audience split and the alerts. The bank match runs hourly on its own clock, the workbook reader watches the folder and runs on submission, and the rest is nightly.
Retries are for plumbing: a timed-out extract, a locked file. A failed check is never retried, because a failed check is a decision for a person. Independent processes write to the same log, and nothing waits for the whole night to finish.
One bad input never stops the rest
A failed check holds back that input only. The page keeps the last good run for it and says so. On the 27th in the console sketch, the Sales workbook was rejected and the actuals published anyway. A named owner gets the alert when a run fails, and when a day and a half passes with no good run, since a job that stops quietly is the failure nobody sees.
Everything is appended, nothing overwritten
The four layers
The store keeps four layers: the raw data as it arrived, the mapped version, the computed tables, and the narrative. Data engineers call the first three a medallion architecture, with the raw layer preserved so anything can be reprocessed or audited;2 the narrative layer is the addition here. Nobody edits the raw layer, the database refuses updates to it, and corrections are new runs. Each input carries two fingerprints, for the file and for the values inside it, since merely saving a workbook alters the file’s fingerprint while every cell stays the same.
Anatomy of a log entry
Every event has the same fields: the time, the run, the actor (the schedule, the model, or a role), the process, the object (a line, an account, a file, a version, an item), the action, what changed before and after, the evidence (a transaction id, a file and cell, a fingerprint), and the reason, for decisions and waivers. For model events it also holds the model version and what it was given, which is what COSO asks for: the prompts, outputs, model version and parameters kept for every call.3
The log is append-only and readable by role, and it inherits the access rules of the data it describes. What this buys is reperformance. Auditing standards name two procedures that fit exactly: recalculating a figure, and re-performing a control the company ran.4 Fixed code and kept inputs make both possible for anything the schedule did.
What needs human eyes
Automated, decided, or waived
An item clears one of three ways, and each is a log event. The evidence lands: the next run sees the journal, the receipt, the match or the approval, and closes it. A person decides, on the page or in the chat card: approve, send back, sign, map, acknowledge. Or a person waives it, with a reason that stays in the log. Most items clear on evidence. The decisions are few and nameable. Nobody ticks a box without one of the three behind it.
The gates
There are six. The analyst decides where a new account maps, with the model’s suggestion in front of them. The controller approves a change to the baseline, sets a provision from the evidence the model assembled, and acknowledges a closed period that moved after a late journal. The FP&A lead signs an explanation, with the model’s draft in front of them. And the CFO signs the month off, with a button that is live only when the list is empty.
What the model may do up to each gate: suggest, draft, assemble, answer. What it may never do: post, approve, sign, or the arithmetic.
Escalation
Owners are alerted at two days. Anything older than five also appears on the controller’s queue. Nothing crosses a month boundary without a waiver and a reason. The queue itself is generated, never typed: from the close calendar when the period rolls, from the checks, from the thresholds, from submissions.
Approvals where people already are
The card in Slack or Teams
The request arrives where the controller already is. What’s being approved, the evidence in three lines (who submitted it, the checks it passed, what it changes), and two buttons. The click is an event with who, when and what. The card updates to say approved and by whom, and links to the log entry.
Identity comes from the workspace sign-in, mapped to a role, never from the text of a message. Both platforms deliver a signed payload to the bot naming the user who clicked; the design verifies the signature and reads the user from it.5 6 Typing “approve” in a thread is a comment. The button is the approval.
The email twin
The same request by email, for teams that live in Outlook: the evidence lines and one link, signed, single-use, for that approver alone. The click is the event. A reply is not an approval, and the bot doesn’t read replies. Slack or Teams is a convenience. Email is the floor.
The question box in the channel
“@close why is freight over budget?” gets an answer with citations, read as the person asking, with every number checked by code before it’s shown. Every question, every answer and every tool call goes in the log, and each person has a daily limit on questions and on cost. A worked question still ends in a memo and a reviewer. The channel is where it’s asked, not where it’s decided.
What the bot can’t do
It can’t write to the ledger, move the baseline except through an approval event, send data outside the company, or render links and images out of file contents. It can’t act on instructions found in a document either: text inside files is data, and the EchoLeak vulnerability in the appendix is the reminder of why.7 The question box’s ten rules are in the appendix; two of them belong here. The tools act as the person asking. And a fixed set of test questions, including ones the agent must refuse, runs before any change of model or data layout.
The console: dashboards as pages over the log
The console is a web page that reads the computed tables and the log, nothing else. It shows close debt, meaning what only a person can still do before the month can be signed and how that burns down through the month; what the month is still waiting for, line by line; tonight’s tie-outs; the queue by role, with the actions on the row; the forecast versions; the log itself, filterable; and a thirty-day scorecard.
Per-audience pages are built from per-audience files, so a department sees its own lines because that’s all its file contains. Access is enforced where the data is fetched, and a nightly test confirms no row leaked. OWASP’s rule for applications built on language models is the one to follow: “implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not”.8
Everything on the page sits one click from its source. Provisional figures look provisional. The time of the last good run is always visible. Try the sketch below: approve v9 on the controller’s tab and watch the log take the event.
The seven controls, and the gate before anyone outside finance logs in
The controls already on the page: the three layers, the approved baseline, the signed explanation, the test month, the checked answer. Seven more sit underneath.
| Requirement | What it means |
|---|---|
| Nothing is overwritten | Raw data is kept as received. A mistake is fixed with a new run. |
| Every run leaves a record | Each night’s inputs, code version and mapping version are stored, so any figure can be traced and any run repeated. |
| Checks block publishing | Rows fetched equal rows stored. Amounts tie to the trial balance. Each forecast file adds up to the total it states. An amount is either mapped or shown as Unmapped. A failed check leaves the last good run on the page, and the page says so. |
| Nobody edits by hand | Tables are never changed directly. Every change to code or mapping gets a second reviewer, and nobody approves a number they prepared. |
| Access is enforced in the data | Each person’s rights apply where the data is fetched, never through instructions to a model. |
| The model’s work is logged | Its inputs, its draft, the reviewer’s edits and the signature are kept. |
| Someone owns the alarm | A named person hears about a failed run, and about thirty-six hours without a good one. |
Nobody outside the finance team gets a login until all of it has held for a month of nightly runs, including a failed check that visibly blocked publishing. Then one test. Pick any number on the page. You should be able to trace it to its source, and see who touched it, inside a minute.
The appendix lists sixteen failure modes, each with its control. Two are worth repeating. When the model is upgraded, its version is on every draft and the test questions rerun before the switch. When the builder leaves, the mapping, the checks and the tests are documented and owned by finance.
Build it in stages
Each stage is usable without the one after it, and most of the value of the first two has nothing to do with AI.
| Stage | What you add | What you can stop doing |
|---|---|---|
| 1. Data moves on a schedule | The nightly load, its checks and its run record. Forecast workbooks read and versioned. | Exporting, tidying and uploading by hand |
| 2. The numbers compute and tie out | The mapping table, the exception queue, tested code for estimates and variances, a page for the finance team | Rebuilding the variance workbook each month |
| 3. Explanations are drafted | A model drafts from the evidence, code checks its citations, a reviewer signs weekly | Writing first drafts after close |
| 4. Other people get a login | One output per audience, with access enforced and tested | Emailing packs |
| 5. Questions are answered | The question box: quick answers first, worked memos later | Saving questions for a quiet week |
Spend week one without a model. Tie one closed month’s ledger to the trial balance by account, to the penny. Agree the standard tab for the forecast workbooks. Write the mapping by hand and list what doesn’t map. Work one month’s variances by hand and keep the answers: that’s the test month. Note, for each line, what posts late and what an estimate could be built from.
Then measure four things from the first night: how long a change takes to become a reviewed number, what share of rows tie with nobody touching them, how often a reviewer rewrites a draft, and how far the mid-month estimates were from the final postings.
What it costs to own
Run it like a small development team. A month of known answers to reproduce before every release. A second reviewer on every change to code or mapping. A release record. The run record, every night. Access enforced in the data and reviewed quarterly. A named owner on the alert.
What that costs in hours isn’t claimed here; I have no figure that would survive checking. Part 1 gives the evidence that it isn’t zero, and the reason it’s still cheap: no licence, no team, and the discipline is the price.
Where the series ends is where it started. A finance team that hears about the contract the week it’s signed, can afford to ask what it means, and can show an auditor every number’s path.
This is part of a series explaining AI and the systems around it for finance people, in their own language. I build AI systems for finance teams; the series is what I’ve learned doing it. This one is a design on paper, not a system I have run.
Sources
Footnotes
-
Oracle NetSuite help, “Executing SuiteQL Queries Through REST Web Services”: https://docs.oracle.com/en/cloud/saas/netsuite/ns-online-help/section_157909186990.html (undated; read October 2026). ↩
-
Databricks documentation, “What is the medallion lakehouse architecture?”: https://docs.databricks.com/aws/en/lakehouse/medallion (read October 2026). Used here as a definition. ↩
-
COSO, “Achieving Effective Internal Control Over Generative AI”, February 2026. Read in a hosted copy: https://auditoresinternos.es/wp-content/uploads/2026/02/Achieving-Effective-Internal-Control-Over-Generative-AI_COSO_compressed.pdf. Deloitte’s summary: https://dart.deloitte.com/USDART/home/publications/deloitte/heads-up/2026/coso-internal-controls-generative-ai ↩
-
International Standard on Auditing 500, “Audit Evidence”, paragraphs A19 and A20, IAASB. Read in a hosted copy of the standard: https://ktkt.uel.edu.vn/Resources/Docs/SubDomain/ktkt/VanBan/ISA/ISA%20500.pdf ↩
-
Slack developer documentation, “block_actions payload” and “Verifying requests from Slack”: https://docs.slack.dev/reference/interaction-payloads/block_actions-payload and https://docs.slack.dev/authentication/verifying-requests-from-slack. The payload’s
useris “The user who interacted to trigger this request”; requests carry a signature header computed with the app’s signing secret. ↩ -
Microsoft Learn, “Universal Action Model” and “Use Universal Actions for Adaptive Card”: https://learn.microsoft.com/en-us/adaptive-cards/authoring-cards/universal-action-model and https://learn.microsoft.com/en-us/microsoftteams/platform/task-modules-and-cards/cards/universal-actions-for-adaptive-cards/work-with-universal-actions-for-adaptive-cards. A button’s Action.Execute reaches the bot as an invoke activity whose
fromfield names the acting user (Bot Framework activity schema). ↩ -
The Hacker News, on the EchoLeak vulnerability (CVE-2025-32711), June 2025: https://thehackernews.com/2025/06/zero-click-ai-vulnerability-exposes.html ↩
-
OWASP Top 10 for LLM Applications 2025, “LLM06: Excessive Agency”: https://genai.owasp.org/llmrisk/llm062025-excessive-agency/ ↩
