I built my first AI agent without understanding what an agent was.
It was about eighteen months ago, in n8n, a tool where you build automations by joining boxes on a canvas. The agent was one box with three things hanging off it: a chat model, a memory and a list of tools. I joined them up and it ran. I couldn’t have told you what happened inside that box between my question going in and the answer coming out.
What happens inside is a loop about ten lines long. I learned it roughly a year later, and once I’d seen it written down, agents stopped being a mystery and became something I could design and check. This post opens the box. By the end you’ll know what an agent is, how it differs from a normal chat, and the one algorithm underneath, well enough to draw it yourself and to judge anyone who tells you their product is “agentic”.
I built an agent before I understood one
It’s a good picture of what an agent is made of. The model does the thinking, the memory holds the conversation, and the tools are what it can reach for. n8n’s documentation insists you attach at least one tool,1 which turns out to be the whole point.
What the picture leaves out is what the box does with those parts. n8n says as much itself: “In n8n, the AI Agent node handles this loop automatically.”2 So the loop is in there, and the canvas doesn’t draw it. I could use an agent. I couldn’t explain one.
I had company. In a 2025 Hacker News thread on what an agent is, Simon Willison stopped to count: “I count SIX definitions of agents in this thread already”.3
A chat answers once. An agent keeps going.
A normal chat turn is one pass. Your message goes to the model with the conversation so far, and one message comes back. The model can’t look anything up along the way. It can only write.
An agent adds one option. The model can still answer, or it can ask for something to be done: search these documents, run this sum, open that file. It doesn’t do the thing itself. Anthropic’s documentation is blunt about that: “The model never executes anything on its own.”4 The application around the model does the work, hands back the result and asks again. And again, until it has what it needs.
Later that year Willison settled on a one-line definition: “An LLM agent runs tools in a loop to achieve a goal.”5 (An LLM is the language model.) Everything is in that sentence.
Finance readers rarely get told this. The usual pitch is outcomes: one consultancy piece promises agents that “will execute, monitor, and optimize finance activities in real time”,6 and says nothing about how. Glenn Hopper’s explainer for finance teams is the exception I know of.7 I want to go one level below it, to the mechanism.
The agentic loop, explained in about ten lines
Here is the whole algorithm, in plain words where a programmer would write code.
conversation = [the system prompt, your question]
repeat, up to 25 times:
reply = ask the model(conversation, list of tools)
add the reply to the conversation
if the reply asks for tools:
run each tool
add each result to the conversation
go round again
otherwise:
stop. The reply is the answer.
That’s all of it. Anthropic’s docs describe the same thing: “The canonical shape is a while loop keyed on stop_reason”.4 OpenAI’s say your application can continue the flow “for as many tool calls as the task requires”.8 Braintrust’s engineers write it in under a dozen lines of real code and note that the same loop sits under Claude Code.9 Thorsten Ball, who built one from scratch, put it this way: “It’s an LLM, a loop, and enough tokens.”10
People who build these for a living will tell you the loop is the easy part, and they’re right.11 You don’t even have to write it. A library will run it for you, and so will a box on a canvas, which is how it got hidden from me. The mystery was never the loop. The picture I’d been given just didn’t draw it.
The six parts of the loop
Six things move through that loop.
The system prompt is the standing instruction: the job description and the house rules, written before the conversation starts. Mine tells the model who it is, how to go about a search, to cite a source for every figure and to prefer a total the document states over one it could add up.
The messages are the conversation so far, kept as a list. Think of it as the working-paper file, everything asked and everything found, in order. The model has no memory of its own, so every time round it gets handed the whole file again.
The tools are a list of things the application offers to do, each with a name, a description and the inputs it needs. If you’ve read my post on APIs, a tool is usually an API call with a label on it. My own system offers about thirty, from search to a calculator.
A tool call is the model’s request, and only a request. It’s a structured note that says “run search with these words”. Nothing has happened yet.
A tool result is what came back, added to the file. An error counts as a result. The model reads it like anything else and decides what to do next.
The stop reason is the model saying why it stopped writing. Two answers matter. “I want a tool” means go round again. “I’m finished” means leave the loop.12
AWS and Google run the same loop under other names
Every platform I’ve used has the same six parts. Only the labels change.
| The part | Anthropic | OpenAI | AWS Bedrock | |
|---|---|---|---|---|
| Standing instructions | system | instructions | system | system_instruction |
| The conversation | messages | input | messages | contents |
| The tool list | tools | tools | toolConfig | tools |
| A tool request | tool_use | function_call | toolUse | function call |
| A tool result | tool_result | function_call_output | toolResult | function response |
| ”Go round again” | stop_reason: tool_use | function calls in the reply | stopReason: tool_use | function calls in the reply |
| ”Finished” | stop_reason: end_turn | a final response | stopReason: end_turn | no function calls left |
The same loop in four dialects, from each vendor’s own documentation.4 8 12 13 14
On AWS you call Bedrock’s Converse API through boto3, Amazon’s Python library. Each call is one visit to the model, so the loop is yours to write. Amazon’s own sample writes it by hand and stops after five rounds.15
On Google Cloud the current library is google-genai. Google’s notice said the older vertexai module would “no longer be available after June 24, 2026”.16 The round trip is the same, with one difference worth knowing. Hand the library plain Python functions as tools and it runs the loop for you, up to ten calls unless you change that.14
n8n does the same job with a box. Once a loop is standard enough for a library to absorb it, most people never see it. (I’ve written about using Claude through both clouds before.)
The inner harness and the outer harness
I think of an agent as two layers. These are my labels, not the industry’s. The nearest published line I’ve found is LangChain’s: “If you’re not the model, you’re the harness.”17
The inner harness is the loop. It’s small, and it’s the same everywhere.
The outer harness is everything I wrap around it, and that’s where the decisions are. In Grounded, the system I built to answer finance questions from a team’s own documents, it means the prompt and the choice of tools. It means keeping the working-paper file from overflowing: once the conversation reaches three quarters of what the model can hold, old tool results get cleared out first. It means a cap of 25 rounds, and a 60-second limit on each simple tool, after which the tool hands the model an error instead of hanging. And it means checks that run after the loop has finished.
I got this wrong at first. When I started building, I still didn’t really understand the loop, so I treated the whole thing like a simple chat and put everything in the prompt. The loop is there so you don’t have to: the model can fetch what it needs, one tool call at a time. The other hard part was proper ingestion, getting the documents in well, because a search tool can only find what was loaded properly.
If you’re hiring someone to build this, or buying it, the split is useful. “Agentic” tells you the inner harness exists. What makes it safe or useful lives in the outer one, so that’s where to look at the work. The model is rented and the loop is standard. The outer harness is the job.
The same loop reconciles a bank account
Once you know the algorithm you can point it at work that has nothing to do with chat. Anthropic’s rule of thumb for when: if the steps are fixed, write an ordinary workflow, and keep agents for jobs where you can’t predict how many steps it will take.18
A bank reconciliation is a fair test. I have a skill for Claude that does the first step of a month-end close for Caldergate, a fictional company. This is one run of it, captured for this post.19
Round one pulls the bank statement and the cashbook. Round two reads last month’s signed reconciliation. Round three runs the matching script. Round four has no tool call.
The script explained all 48 statement lines, passed its five checks and agreed the reconciliation to the cashbook to the penny. It flagged one item: cheque 004182 for £1,850, outstanding for 47 days.
Two choices in that design are worth copying. The model never matches or adds anything. A script does, and the model only decides what to run next.
And the loop stops by asking. Reissue the cheque, write it back, or carry it forward? The agent isn’t allowed to decide that, so it goes to a named reviewer.
A clean month like this takes a predictable four rounds. The loop earns its keep in the messy month, when every unmatched line wants a different lookup, and I haven’t shown that here.
Twenty-nine actions I never scripted
Grounded runs the same loop over a finance team’s documents. Take a small question it was asked in testing, again on the Caldergate data: what was total revenue for May 2026?20
It took two rounds and 17 seconds. In round one the model asked for three searches at once, and every passage that came back carried a numbered stamp. In round two it asked the calculator to divide one stamped figure by a thousand. Then it answered: £21.063m, with the stamp beside it.
That stamp is the citation, and the tool issues it. The model can only copy it. After the loop ends, any stamp the tools never issued is stripped out, and every figure in the answer is checked against the stamped sources and the calculations shown. The result of that check is stored with the answer. It’s what I mean by verifiable.
The bigger question was a reconciliation one: where does a £285k gap between two reports come from? This time the loop ran for 12 rounds and 59 seconds, and made 29 tool calls across nine different tools.20
It didn’t go smoothly, which is why I like it. The model went looking in the wrong month. Then it reached for the ERP, and the ERP was switched off. The tool handed back that error as text, the model read it and went back to the documents. About half the rounds went on the detour.
It finished on a subtraction it asked the calculator to do, 4,832 less 4,547, and on one journal: JNL-508295, a £285,000 carrier accrual posted on 10 June into May, for an invoice that arrived after the ledger cut.
Nobody scripted those 29 steps. The path was chosen one round at a time, and a fixed pipeline couldn’t have got there.
Where it stops, and where it runs away
My loop has three ways out. The model says it’s finished. The round cap is reached. Or the person cancels.
A tool that times out isn’t an exit. Its error goes back into the file and the loop carries on. The check that can replace an answer with a refusal runs after the loop has ended, and in my local testing it has never had to.
The stop I like best is the one where it finds nothing. Asked for one person’s salary, the model searched, listed files, found no payroll document and said so. Three rounds, 14 tool calls, and no number in the answer.20
Loops do run away. A 2026 study found 68 confirmed infinite loops across 47 agent projects.21 One public bug report describes an agent making the same search call 364 times in about 50 minutes.22 The toolkits I looked at all offer a cap, and their defaults range from ten rounds to none at all.23 24
Mine: across 652 threads on my local machine, most of them automated test runs, the typical deepest turn is 5 rounds and the deepest ever is 21. The cap of 25 has never been hit. There is no spend limit in the loop.20
What I’d tell your CFO
An agent is a model in a loop with a list of tools. When someone says “agentic”, ask three things: what tools does it have, what makes it stop, and who checks the answer after it stops. I’d hand that loop a reconciliation or a review of an Anaplan model next. I wouldn’t let it post a journal, approve a payment or sign off the close. Knowing the loop gives me a framework for defining what has to be done, in the context of the best way for an agent to solve the problem.
This is part of a series explaining AI and the systems around it for finance people, in their own language. I build AI systems for finance teams; the series is what I’ve learned doing it.
Sources
Footnotes
-
n8n, “AI Agent node” documentation: https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.agent/ (undated; read 1 October 2026). “You must connect at least one tool sub-node to an AI Agent node.” ↩
-
n8n blog, post on ReAct agents, 30 April 2026: https://blog.n8n.io/react-agent/ ↩
-
simonw on Hacker News, 2 April 2025: https://news.ycombinator.com/item?id=43560849. Simon Willison is a co-creator of the Django web framework and writes at simonwillison.net. ↩
-
Anthropic, “How tool use works”: https://platform.claude.com/docs/en/agents-and-tools/tool-use/how-tool-use-works (undated; read 1 October 2026). ↩ ↩2 ↩3
-
Simon Willison, “I think ‘agent’ may finally have a widely enough agreed upon definition to be useful jargon now”, 18 September 2025: https://simonwillison.net/2025/Sep/18/agents/ ↩
-
BCG, publication on the AI-first finance function, June 2026: https://www.bcg.com/publications/2026/the-artificial-intelligence-first-finance-function ↩
-
Glenn Hopper, “AI Agents for Finance”, 21 August 2026: https://glennhopper.substack.com/p/ai-agents-for-finance ↩
-
OpenAI, “Function calling” guide: https://developers.openai.com/api/docs/guides/function-calling (undated; read 1 October 2026). ↩ ↩2
-
Ankur Goyal, “The canonical agent architecture: a while loop with tools”, Braintrust, August 2025: https://www.braintrust.dev/blog/agent-while-loop ↩
-
Thorsten Ball, “How to Build an Agent”, April 2025: https://ampcode.com/how-to-build-an-agent ↩
-
“The Loop Was the Easy Part”, Towards AI: https://pub.towardsai.net/the-loop-was-the-easy-part-evals-observability-and-rollbacks-for-your-diy-claude-code-45f2a9f1c99c (undated). ↩
-
Anthropic, “Handling stop reasons”: https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons (undated; read 1 October 2026). ↩ ↩2
-
AWS, Amazon Bedrock API reference, “Converse”: https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_Converse.html (undated; read 1 October 2026). ↩
-
Google, Gen AI SDK for Python documentation, function calling and automatic function calling: https://googleapis.github.io/python-genai/ (undated; read 1 October 2026). ↩ ↩2
-
AWS documentation SDK examples,
tool_use_demo.py, which setsMAX_RECURSIONS = 5: https://github.com/awsdocs/aws-doc-sdk-examples/blob/main/python/example_code/bedrock-runtime/cross-model-scenarios/tool_use_demo/tool_use_demo.py ↩ -
Google Cloud, deprecation notice for the generative AI modules of the Vertex AI SDK: https://docs.cloud.google.com/vertex-ai/generative-ai/docs/deprecations/genai-vertexai-sdk (September 2026). ↩
-
Vivek Trivedy, “The anatomy of an agent harness”, LangChain, March 2026: https://www.langchain.com/blog/the-anatomy-of-an-agent-harness ↩
-
Anthropic, “Building effective agents”, December 2024: https://www.anthropic.com/engineering/building-effective-agents ↩
-
My own run on fictional data, 1 October 2026. The matching step was re-run for this post and its output is byte-identical to a run from August; the statement, cashbook and prior reconciliation were read from that run’s saved files, not pulled live. The skill and data are public: https://github.com/klameer/audited-ai-close ↩
-
My own runs on the fictional Caldergate data, read from my local test database on 1 October 2026: threads from 12, 27 and 29 August 2026. These questions were asked during testing, by me or by an AI assistant working for me. More on the system: https://codelessops.com/posts/introducing-grounded/ ↩ ↩2 ↩3 ↩4
-
arXiv paper 2607.01641 on “infinite agentic loops”, July 2026: https://arxiv.org/abs/2607.01641 ↩
-
opencode issue 45442, August 2026, a single user’s report: https://github.com/anomalyco/opencode/issues/45442 ↩
-
OpenAI Agents SDK, “Running agents” (
max_turns): https://openai.github.io/openai-agents-python/running_agents/ (undated; read 1 October 2026). ↩ -
Anthropic, Claude Agent SDK for Python reference (
max_turns,max_budget_usd, both unset by default): https://code.claude.com/docs/en/agent-sdk/python (undated; read 1 October 2026). ↩
