This is the appendix to the continuous close series. The three parts make the argument and set out the design; this page holds the detail they leave out, for whoever builds or reviews the thing: a finance systems analyst, a controller, an IT lead. Read the series first.
The same limits apply. One FP&A cycle, for one entity in one currency, P&L only. It’s a design on paper. It hasn’t been run, it claims no results, and it gives no costs or build times.
What’s here
- The design at a glance, and who does what
- The nightly load in detail
- What each mid-month line is waiting on
- The record one night’s run leaves behind
- A language model against fixed code: the benchmarks
- What people see when they log in
- The question box: evidence and rules
- Who sees what, and the ways round it
- Sixteen ways it goes wrong, and the go-live gate
- The longer starting lists
Who does what, by role
| Role | Every night | Every week | At month end |
|---|---|---|---|
| The schedule | Loads, checks, variances, drafts | ||
| Analyst | Clears the exception queue, fixes mappings | Supplies accrual estimates | |
| FP&A lead | Reads, corrects and signs explanations | Signs the commentary | |
| Controller | Approves baseline and mapping changes | Owns the close and the final sign-off | |
| Budget holder | Reads their own lines, asks questions | Confirms what they owe |
The nightly load, in detail
It runs on a scheduled job, one database, the shared folder you already have, a small site behind company sign-in, and a language model called over the internet.
Actuals in. Nightly is enough. Reporting needs the numbers there every morning, and doesn’t need them every second.
Transactions in an open period get edited and deleted, and late journals land in earlier periods. So the job pulls every unlocked period again each night, and a closed period whose total moves raises an exception. One detail for whoever builds it: NetSuite’s query interface returns at most 100,000 rows per query, so the pull goes in chunks under that cap, and a chunk that comes back with exactly 100,000 rows fails the run.1
Forecasts in. Standardising the workbooks is the real work of this stage, because a dozen files built by different people won’t share a layout. The standard tab has a fixed name and fixed columns, with the workbook’s ID and layout version in fixed cells. Submission is a status cell, and half-finished work stays unsubmitted. The budget’s locked version carries its monthly phasing.
The reader rejects a file whose layout it doesn’t recognise and never guesses. It also rejects a workbook that wasn’t recalculated before saving, because a program reading the file sees only the last values Excel worked out. And it sums the lines it reads and compares the result with the total the workbook states.
Mapping. Each row of the mapping table says from what date it applies. A second table does the same job for forecast lines.
Compute. The same inputs give the same answer every time. (It’s the dullest sentence in this guide, and everything else leans on it.) A line is flagged when its variance passes the greater of an amount and a percentage, set per line by its owner, so small lines don’t flood the queue.
Explain. The evidence for a flagged line is the transactions behind it, summarised first if there are thousands, the purchase orders, the analysis files in the folder and last month’s commentary. Without that material a model does what one analyst described after testing AI on budget variances: “it literally just puts into words what the data is showing”.2
What each mid-month line is waiting on
| What’s missing mid-month | When the real number lands | What the page shows until then |
|---|---|---|
| Supplier invoices not yet received | Through the month, and after it ends | Open purchase orders and goods received, as an estimate |
| Accruals | Posted by accounting at month end | A running estimate from purchase orders and contracts, replaced at close |
| Payroll | On the pay date | Last month’s payroll, adjusted for joiners and leavers |
| Revenue recognised by a month-end run | At close | Billings to date, labelled as billings |
| Depreciation, allocations, currency revaluation, intercompany | When accounting runs them at close | Last month’s figure, marked as carried forward |
A line with an accrual reversal still open isn’t flagged until the reversal is matched, or the first week of every month would light up.
People already make these estimates by hand. One FP&A practitioner on Reddit, in a thread about how long teams get for analysis after close, wrote: “Sometimes the month’s payroll is not even out yet and I have to estimate it for every cost centre.”3 The design makes the same estimate every night and labels it.
The view also depends on habits upstream. Invoices coded and posted as they arrive give a truer mid-month number than invoices batched in the last week. That part is accounting’s process to change, and it’s worth asking for.
The record one night’s run leaves behind
Nothing is overwritten. Data is kept in layers: raw as received, mapped, computed, and the narrative written about it. The first three follow what data engineers call a medallion design, where the raw layer is kept for reprocessing and audit. The narrative layer is added here.4 The raw layer is never edited, so any run can be repeated later. A correction is a new run and never an edit to an old one.
Every run is stamped. Each night’s run records a run number, the version of the code, the version of the mapping table and two fingerprints for every input. A fingerprint is a short code calculated from content: change the content and the fingerprint changes. One is taken of the file and one of the values read from it, because saving a workbook again changes the file’s fingerprint even when no cell has moved.
Checks block publishing. Four checks run before anything reaches the page.
- The rows fetched from the ledger equal the rows stored.
- For each account and period, the amounts loaded sum to the ledger’s own trial balance movement, taken at the same moment and through a role that can see everything in scope.
- The lines read from each forecast file sum to its stated total.
- Every amount is mapped or sits on the Unmapped line.
If a check fails for one input, the page keeps showing the last good run for that input and says so. One bad workbook shouldn’t freeze the actuals.
The case for the first check is Public Health England in 2020. Test results were stored in an old spreadsheet format that holds 65,536 rows per sheet, and 15,841 Covid cases from an eight-day stretch were left out of the daily figures.5 A count of rows received against rows loaded would have caught it on the first night.
The formulas are tested. Totals don’t catch a wrong formula. The fix is a test month with known answers. The code has to reproduce them before every release.
Nobody edits by hand. No one changes a table directly. Changes to the code are reviewed by a second person. The database refuses updates to the raw layer. And a named owner is alerted when no run has succeeded for a day and a half, because a job that quietly stops is the failure nobody notices.
The model’s work is logged. COSO asks for “logs of prompts, outputs, system messages, model version, parameters, and plugins used”.6 For each drafted explanation, keep what the model was given, what it wrote, what the reviewer changed, and who signed.
Together these let anyone run the work again. Two of the procedures auditors use apply directly. Recalculation is “checking the mathematical accuracy of documents or records”. Reperformance is the auditor’s “independent execution of procedures or controls” that the company originally performed.7 Fixed code and kept inputs make both possible. A model’s free-text answer doesn’t.
A citation shows where an answer points and doesn’t prove the answer is right, which is the subject of an earlier essay on verification.
A language model against fixed code: the benchmarks
The principle is already in print. F3 Insights, a fractional-CFO firm, put it this way in June 2026: “Compute every figure directly from the ledger with ordinary code. No model touches the arithmetic.”8 The evidence for that rule is strong, and it has limits.
| A language model | Fixed, tested code | |
|---|---|---|
| As the calculator | With no formula given, the best model got 38.85% of table calculations right on a 2026 benchmark of full financial statements.9 Writing and running its own code, models reached 89.1% on one 2025 test and 75.32% on another.10 11 | Same inputs, same answer |
| As the reader of your files | 96.5% on clean invoices and 87.5% on scanned receipts for the best model in one 2025 test.12 On questions over financial spreadsheets the best model scored 82.4%, and ten models averaged 48.6% on the largest workbook.13 | Reads an agreed layout exactly, and rejects what it doesn’t recognise |
| As the checker | Under one study’s checklist prompt, nine of fourteen runs flagged 95% to 100% of clean statements as wrong. One run caught every planted error with no false alarms, on unrounded figures.14 | No false alarms, and about half the planted errors found14 |
Three things follow from that table.
A model writing fresh code for each question is a different thing from code written once, tested and run every night. Only the second gives the same answer twice. In the paper where the model’s code was run and its mistakes reviewed, most were errors of logic in the program it had written.11
A model shouldn’t read your forecast files unaided either. That’s why the workbooks here go through a reader against an agreed layout, and why anything a model does read is checked like any other input.
And fixed rules have a blind spot of their own. They catch what someone thought to write a rule for. A model makes a useful second reader, but its verdict swings with how it’s asked, so it should never be the control.
The essay’s design is the conservative one. COSO’s February 2026 guidance on generative AI allows more, including an agent that posts reconciliations by itself above a validated confidence threshold.6
The model does touch the numbers in one place: a suggested mapping decides which line an amount lands on. So the person makes the mapping decision themselves, with the suggestion as a prompt, and records it. Where a control’s owner re-performs the model’s work like that, COSO says evidence expectations “may be proportionately lower”.6
What people see when they log in
The page has four views.
Variances. Flagged lines come first. Each line opens to show the transactions behind it, the explanation and the source files. Every explanation shows its state: drafted, signed or stale, and by which role.
Forecast compare. Any two versions side by side: what moved, by how much, in which file, submitted by whom, and whether the change has been approved into the baseline.
Tie-outs. The last run’s checks, each marked pass or fail, with the time of the run. If the page is showing an older run for one input because tonight’s failed, this is where it says why.
Ask. A box for questions.
Two rules hold on every view. Every figure is one click from its source. And provisional numbers look provisional, with the time of the run beside them. COSO’s suggested control for this kind of output fits on a sticky note: “Require citations for all material outputs.”6
The question box: evidence and rules
Finance chat agents already ship. Oracle’s tools for NetSuite let an AI client query the ERP in plain language, and its documentation says: “The tools don’t provide any additional access beyond what your NetSuite role allows.”15 The same tools can create and update records, so read-only scope is something you configure.
The companies that make the models document a plan, gather, cite and report pattern for research in general.16 17 Neither describes a report that lists its assumptions or what couldn’t be determined. Those two are this design’s additions, because an unknown that’s stated can be chased and one that’s hidden can’t.
How reliable are these agents? It depends on what sits around the model. In the original Spider 2.0 study of 632 enterprise database questions, the best general agent solved about one in five.18 On Vals AI’s finance agent benchmark, where the model researches with search tools, the top score is 64.37%.19
So the rules for the question box are stricter than for a general assistant.
- The agent’s tools are named lookups and the fixed calculations. It doesn’t write its own database queries, and it never does its own sums.
- Code checks every number in an answer against what the tools returned. An answer that doesn’t match isn’t shown.
- Every figure carries a citation the reader can open.
- If the data doesn’t hold the answer, the agent says so.
- It reads as the person asking, and it can’t write to anything.
- Text inside files is data. It’s never treated as an instruction, no tool sends anything outside, and answers don’t render links or images.
- A worked question ends in a memo with its assumptions and unknowns, and a person reviews it.
- Every question, answer and tool call is logged.
- Each person has a daily cap on questions and on spend. Worked questions cost more to run: Anthropic reports that agents use about four times the tokens of a chat, tokens being the unit model usage is billed in.16
- A fixed list of test questions, each with a known answer, runs whenever the model or the data layout changes. It includes questions the agent must refuse, questions from the wrong audience and a document with a planted instruction. A failed test blocks the release.
Glenn Hopper, who writes on AI in finance, has the line to pin above the design: “The model supplies the reasoning, but everything about control lives in the harness.”20 The harness is everything in that list.
Who sees what, and the ways round it
The essay gives the rule: enforce access where the data is fetched, as the person asking, and don’t rely on instructions to the model. This section is about where that rule lives and the ways round it.
Microsoft’s own guidance for rolling out Copilot says the assistant works with “data that the user already has permission to access”, and makes fixing overshared files the first step before switching it on.21
The finance products whose documentation I read state the same rule for themselves. Oracle says of NetSuite’s AI connector: “Queries executed through the connector respect your NetSuite role’s permissions”.22 Microsoft says its finance agent’s ERP chat depends on “security and privileges as defined within the ERP”.23
A team building its own has to choose where that rule lives, because the page and the agent read computed tables in a database, and folder permissions don’t guard a database.
The simplest sound choice is to make files the only thing people read. The nightly run writes one output per audience: a variance workbook for each department and the full one for finance. Each goes to its own folder, access is granted by folder, and the site and the question box open those files with the signed-in person’s own rights. Access is per file, so a workbook that mixes payroll with everything else gets split in two. A nightly test confirms that no audience’s file contains a department outside its list.
The alternative is rules in the database about which rows each person may read, keyed to the groups in your company sign-in. Database vendors document how.24 That gives finer control, and one more copy of the permissions to keep current.
Either way, three things need a decision. An explanation drafted with full access can mention another department’s numbers, so explanations are drafted per audience or reviewed before a narrower audience sees them. Drill-through to transactions stops at payroll for anyone outside the payroll audience. And a manager who spans departments gets the folders for each, with no merged view built specially for them.
Then there are the ways round the door. Each is worth a line in your design review.
| The way round | What it rests on | What to do |
|---|---|---|
| Output built with one person’s rights, then shared | Pigment’s documentation: “Access rights are not re-checked when other Members select View output.”25 | Build outputs per audience |
| A stale copy of the permissions | Anaplan: access changes “take effect the next time your data syncs”.26 Microsoft documents the same lag for its search index.27 | Read permissions live, or refresh them nightly and when someone leaves |
| A tool that runs as one shared system login | OWASP: tools should act in the user’s own security context.28 | Tools run as the person asking |
| The workbook itself | Reasoning: whoever holds a file holds everything in it | Split by audience before writing |
| Data that has left the system | Oracle says it can’t control what an AI client does with data once it leaves NetSuite.22 | No tool sends data out; exports are logged |
| Totals that reveal a value by subtraction | Reasoning, borrowed from official statistics: a team’s payroll total, minus the same total without one person, is a salary.29 | No totals over very small groups |
| The logs | Reasoning: the log holds everything anyone was shown | Same access rules as the data |
| Instructions hidden in a document | EchoLeak, 2025: a flaw that could let a crafted email make Microsoft 365 Copilot send out data from its context. Patched, with no use reported in the wild.30 | File text is data; the agent can’t send anything out |
I found no documented case of payroll leaking through a finance chatbot, and none is implied here. The risk is argued from how these systems work. The old rule still applies as well. The person who prepares a number isn’t the person who approves it, and a change to the mapping needs a second pair of eyes.
The other half of this subject is data leaving the company altogether, which is covered in how to use AI at work without leaking data.
Sixteen ways it goes wrong, and the control for each
| What goes wrong | The control |
|---|---|
| A provisional number gets read as final | Provisional figures are marked on every view and export. The statutory close keeps its own timetable. |
| A line looks healthy because its invoices haven’t arrived | Posted, estimated and final layers on every line. Flags run against what should have landed by now. |
| One event whipsaws the forecast | A new version is a proposal. The baseline moves only when a named role approves it. |
| A signed explanation goes stale | A signature belongs to a run and an amount. If the number moves past its threshold again, the explanation shows as stale. |
| The exception queue outgrows its reviewers | Measure its size and age. A recurring exception becomes a mapping row or a rule that week. |
| A workbook’s layout drifts | The reader rejects what it doesn’t recognise. It never guesses. |
| The nightly run quietly stops | A named owner is alerted after a day and a half with no good run. |
| An error spreads faster than it used to | Checks block publishing per input, and the page falls back to the last good run. |
| A wrong formula runs every night, whoever or whatever wrote it | A test month with known answers, run before every release. A second person reviews every change. |
| Someone edits the database directly | No hand edits, reviewed code changes, a raw layer that can’t be updated. |
| The model is upgraded and behaves differently | The model version is recorded on every draft. The test questions run before switching. |
| Reviewers stop reading the drafts | Sample signed explanations each month. Track how often reviewers change a draft. |
| The agent gives a wrong answer with confidence | Code checks every number against the tool outputs before the answer is shown. |
| A document carries hidden instructions | File text is data. The agent has no tool that sends anything out. |
| Someone sees what they shouldn’t | Outputs per audience, access enforced where data is fetched, a nightly test of the split, a quarterly access review. |
| The person who built it leaves | The mapping, checks and tests are written down and owned by finance. |
Before anyone else gets a login. The controls above reduce to a gate. Nobody outside the finance team is let in until every line is true.
- The four checks have passed for a full month of nightly runs, and a failed check has been seen to block publishing.
- The test month reproduces its known answers.
- Provisional numbers are marked on every view and every export.
- Every explanation on the page is signed, or shows as drafted or stale.
- Outputs are split per audience, and the nightly test of the split passes.
- If the question box is on, it passes its test questions, including the ones it must refuse.
- The logs carry the same access rules as the data.
- A named owner gets the alert when a run fails, or when none has succeeded for a day and a half.
Starting: the longer lists
The essay has the stages. These are the lists behind them.
First, check what you already own. Microsoft’s documentation for its Finance Agent add-in for Excel sets out this variance step for a single workbook: a calculation by a stated formula, a written summary, and reference sheets showing the rows used.31 It’s a preview, and Microsoft’s notes for the sister reconciliation feature say preview features “aren’t meant for production use”.32 Planning vendors sell AI analysis on data you move into their platforms.33 34 35 For some of these products I read the documentation, and for others only the product page.
- One workbook and one team: use the feature you already have, and watch Microsoft’s preview.
- Your planning platform already holds the data: use its AI features, and ask how they compute their numbers and who can see what.
- Many files, versions to track, a mapping to govern and other people to let in: build.
Pick one process. Friar’s advice is to “begin with a consequential decision and work backward”.36 PYMNTS, writing for mid-market finance chiefs, suggests one recurring process with an owner, a measurable cycle time and a clear output.37 The reason to stay narrow is in Gartner’s 2025 survey. Of the finance functions using AI, 91% reported a low or moderate impact.38
Decide who builds it, and who owns it. Three kinds of skill are needed, and they needn’t be three people. Someone who knows the ERP’s data well enough to pull the ledger and explain every column. Someone who can write, test and schedule the code: a finance systems analyst, a data engineer in IT, or a contractor. And an owner in finance for the mapping, the thresholds and the test month.
That last role can’t be handed to anyone else. Look at what’s left once the code runs: which formula is right for a line, how a rate should be aggregated, which checks deserve a rule, who may change the mapping. One paper’s error review fits. With a program doing the sums, most of the mistakes reviewed were errors of logic.11 Reading those as finance judgements is my inference, and it’s why finance has to own the design even when IT builds it.
The second skill is the one AI has changed. With an AI coding assistant, a finance systems analyst or a determined FP&A analyst can write the reader, the checks and the test month, with a second person reviewing every change. Leave IT what it’s best placed to own: where the database and the site run, company sign-in and backups. The controls in the essay are the things spreadsheets never had, and they apply whoever, or whatever, wrote the code.
Take a list to IT.
- A read-only role for the ERP’s query interface that can see every account in scope.
- Somewhere for the database and the site to live, and someone to patch them.
- Company sign-in for the site.
- Which model provider is approved and on what terms, and whether ledger detail may be sent to it.
- Read access to the forecast folder for the scheduled job.
- Who is alerted when a night’s run fails.
Week one needs no model and almost no code.
- Pick the process, the entity and the owner.
- Pull one closed month from the ledger and tie it to the trial balance report by account, to the penny. Pull it again the next day and compare.
- List every forecast workbook and how each is laid out.
- Agree the standard tab the reader will look for.
- Write the mapping table by hand for that month, and list what doesn’t map.
- Work out the variances by hand once and keep the answers. That’s the test month.
- For each line, write down what posts late and where an estimate could come from.
- Set the threshold for a flagged line.
Measure it. Friar’s test for AI spend is the volume of completed work that meets a defined quality bar, and what each successful task costs.39 COSO’s guidance suggests measures such as forecast error and the share of outputs with a reviewer’s sign-off.6 For this process that becomes a short scorecard:
- Cycle time, from a change in the ledger or a submitted forecast to a reviewed number on the page.
- The share of rows that tie without a person.
- Exceptions per run, and the age of the oldest.
- How often reviewers change a drafted explanation, and by how much.
- The reviewer’s hours each week.
- Questions asked, the share answered with every number verified, and what they cost to run.
- The gap between mid-month estimates and what was finally posted.
- Forecast error against actuals.
Then improve it. Two loops do the work. The exception queue is the backlog: anything that recurs becomes a mapping row or a rule. The reviewers’ edits are the feedback: where they keep rewriting a draft, the model was missing a source, so add the source before reaching for a cleverer prompt.
This is part of a series explaining AI and the systems around it for finance people, in their own language. I build AI systems for finance teams; the series is what I’ve learned doing it. This one is a design on paper, not a system I have run.
Sources
Footnotes
-
Oracle NetSuite help, “Executing SuiteQL Queries Through REST Web Services”: https://docs.oracle.com/en/cloud/saas/netsuite/ns-online-help/section_157909186990.html (undated; read October 2026). ↩
-
u/echobik in r/FPandA, “Prompts for variance analysis?”, 3 August 2026: https://www.reddit.com/r/FPandA/comments/1ve9pri/prompts_for_variance_analysis/. A pseudonymous practitioner describing their own test. ↩
-
u/sand_snow and commenters in r/FPandA, “How much time do you usually have for analysis after close?”, 7 March 2026: https://www.reddit.com/r/FPandA/comments/1rnh2wt/how_much_time_do_you_usually_have_for_analysis/. Pseudonymous practitioners, quoted only as testimony about their own work. ↩
-
Databricks documentation, “What is the medallion lakehouse architecture?”: https://docs.databricks.com/aws/en/lakehouse/medallion (read October 2026). Used here as a definition. ↩
-
The Register, “What a Hancock-up: Excel spreadsheet blunder blamed after England under-reports 16,000 COVID-19 cases”, 5 October 2020: https://www.theregister.com/2020/10/05/test_and_trace/ ↩
-
COSO, “Achieving Effective Internal Control Over Generative AI”, February 2026. Read in a hosted copy: https://auditoresinternos.es/wp-content/uploads/2026/02/Achieving-Effective-Internal-Control-Over-Generative-AI_COSO_compressed.pdf. Deloitte’s summary: https://dart.deloitte.com/USDART/home/publications/deloitte/heads-up/2026/coso-internal-controls-generative-ai ↩ ↩2 ↩3 ↩4 ↩5
-
International Standard on Auditing 500, “Audit Evidence”, paragraphs A19 and A20, IAASB. Read in a hosted copy of the standard: https://ktkt.uel.edu.vn/Resources/Docs/SubDomain/ktkt/VanBan/ISA/ISA%20500.pdf ↩
-
F3 Insights, on AI variance analysis with an audit trail, June 2026: https://f3insights.com/blog/articles/ai-variance-analysis-audit-trail ↩
-
“FININDICES” benchmark paper, arXiv 2607.28661, July 2026: https://arxiv.org/abs/2607.28661. The paper mentions no tools or code execution for the models. ↩
-
“FinanceReasoning” benchmark paper, arXiv 2506.05828, June 2025: https://arxiv.org/abs/2506.05828. OpenAI o1 with executed programs: 89.1% on the hard subset. ↩
-
“FINDER” paper, arXiv 2510.13157, October 2025: https://arxiv.org/abs/2510.13157. 75.32% on FinQA with executed code; of 100 errors reviewed, most were logical. ↩ ↩2 ↩3
-
Fraunhofer IAIS benchmark of eight models on invoice and receipt extraction, arXiv 2509.04469, 2025: https://arxiv.org/abs/2509.04469 ↩
-
“FinSheet-Bench” paper, arXiv 2603.07316, March 2026: https://arxiv.org/abs/2603.07316. Ten model configurations on questions over financial spreadsheets. ↩
-
“FinVerBench” paper, arXiv 2605.29586, May 2026: https://arxiv.org/abs/2605.29586. The rule-based verifier’s recall was 51.6%. One model run found all 62 injected errors with no false positives, on unrounded figures. ↩ ↩2
-
Oracle NetSuite help, MCP Standard Tools SuiteApp: https://docs.oracle.com/en/cloud/saas/netsuite/ns-online-help/article_143403258.html (undated; read October 2026). ↩
-
Anthropic, “How we built our multi-agent research system”, June 2025: https://www.anthropic.com/engineering/built-multi-agent-research-system ↩ ↩2
-
OpenAI, deep research API guide: https://developers.openai.com/api/docs/guides/deep-research (undated; read October 2026). ↩
-
Spider 2.0, a benchmark of 632 enterprise text-to-SQL problems, ICLR 2025: https://spider2-sql.github.io/. The project page gives 17.1% for the paper’s o1-preview agent and the revised paper 21.3%. Later purpose-built agents score far higher on its self-reported leaderboard. ↩
-
Vals AI, Finance Agent benchmark, updated 4 June 2026: https://www.vals.ai/benchmarks/finance_agent. 537 analyst questions; top model Claude Opus 4.7. ↩
-
Glenn Hopper, “AI Agents for Finance: What They Are, Why They Got Better, and How They Work”, 21 August 2026: https://glennhopper.substack.com/p/ai-agents-for-finance. Hopper is the author of Deep Finance. ↩
-
Microsoft Learn, guidance on a secure and governed data foundation for Microsoft 365 Copilot, April 2026: https://learn.microsoft.com/en-us/microsoft-365/copilot/configure-secure-governed-data-foundation-microsoft-365-copilot ↩
-
Oracle NetSuite help, AI Connector Service frequently asked questions: https://docs.oracle.com/en/cloud/saas/netsuite/ns-online-help/article_4160616848.html (undated; read October 2026). ↩ ↩2
-
Microsoft Learn, Finance Agent, questions and answers over ERP data (preview), July 2026: https://learn.microsoft.com/en-us/copilot/finance/agent-in-copilot-chat/erp-qa ↩
-
Supabase documentation, “RAG with Permissions”: https://supabase.com/docs/guides/ai/rag-with-permissions (read October 2026). ↩
-
Pigment documentation, “Set up Analyst Agent”, July 2026: https://kb.pigment.com/docs/set-up-analyst-agent ↩
-
Anaplan help, CoPlanner access controls, May 2025: https://help.anaplan.com/68ce92c0-600e-432a-aed5-b0ae395697e3 ↩
-
Microsoft Learn, “Document-level access control in Azure AI Search”, August 2026: https://learn.microsoft.com/en-us/azure/search/search-document-level-access-overview ↩
-
OWASP Top 10 for LLM Applications 2025, “LLM06: Excessive Agency”: https://genai.owasp.org/llmrisk/llm062025-excessive-agency/ ↩
-
Immigration and Refugee Board of Canada, “Small value suppression”: https://www.irb-cisr.gc.ca/en/statistics/Pages/small-value-suppression.aspx (read October 2026). An analogy from official statistics, not finance guidance. ↩
-
The Hacker News, on the EchoLeak vulnerability (CVE-2025-32711), June 2025: https://thehackernews.com/2025/06/zero-click-ai-vulnerability-exposes.html ↩
-
Microsoft Learn, Finance Agent, “Analyze variances” (preview), May 2026: https://learn.microsoft.com/en-us/copilot/finance/variance/analyze-variances ↩
-
Microsoft Learn, “Responsible AI FAQ for reconciliation”, October 2025: https://learn.microsoft.com/en-us/copilot/finance/responsible-ai/responsible-ai-faq-for-reconciliation ↩
-
Aleph, guide to a live, drillable budget-versus-actuals report in Excel, updated October 2026: https://www.getaleph.com/answers/live-drillable-bva-excel-ai ↩
-
Cube, AI product page: https://www.cubesoftware.com/ai (read October 2026). A product page, not documentation. ↩
-
Planful, Analyst Assistant product page: https://planful.com/ai/analyst/ (read October 2026). A product page, not documentation. ↩
-
Sarah Friar, “What building an AI-native finance function taught me”, OpenAI, 10 August 2026: https://openai.com/index/building-an-ai-native-finance-function/ ↩
-
PYMNTS, on what mid-market finance chiefs can take from Friar’s article, 13 August 2026: https://www.pymnts.com/news/artificial-intelligence/2026/openai-cfo-sarah-friar-built-an-ai-native-department-mid-market-cfos-can-start-smaller/ ↩
-
Gartner survey of 183 CFOs and senior finance leaders, May to June 2025, read in a reprint of its release: https://insurance-canada.ca/2025/12/10/gartner-finance-ai-adoption-steady/ ↩
-
Fortune, on Sarah Friar’s four questions for AI spend, 17 July 2026: https://fortune.com/2026/07/17/openai-cfo-4-questions-reveal-your-ai-spend-paying-off/ ↩
