Last week you built the loop. This week you build the Act arrow. A tool is how a language model reaches something outside itself — a calculator, a database, a payment API — and it is simultaneously the moment your agent stops being a chat window and starts being able to do damage. Tools are capability plus risk surface, and this week is about designing that surface on purpose.
Week 4 called this step “Act” and left it as a black box. Here it is, opened:
Last week “Act” was a label on an arrow. You stepped through episodes where the agent “called Search” and a result appeared, and we deliberately did not ask how.
Five steps, four of which are ordinary software you write and test. Pages 01 and 02 take them one at a time.
Your Berlin supplier invoiced 4,817 units at €293 each, plus 18% VAT. What is the euro total? Before anyone tells you the answer, find out for yourself what a local model does with this question.
Everything here is copy-and-paste into one window. Start Ollama, then in a terminal run ollama run qwen2.5:0.5b once — after that you are just typing at a prompt. No files, no folders, no Python.
Produce: the filled table below, plus one sentence per round: what changed, and why you think it changed.
Paste the identical question five times and write down all five answers. Press ↑ to repeat the last one.
Nothing to interpret now — no wording, no percentage, just one multiplication. Does that fix it?
The model cannot run anything, so for this round you are the code that runs it. Send this, and it will reply with a request instead of an answer:
Write down the expression it asked for — is it the right formula? Then compute it yourself (phone calculator is fine; the correct value of 4817*293*1.18 is 1665429.58) and hand the result back:
Run the whole round three times. You have just done by hand exactly what the Python on page 03 does automatically.
| Run | Round 1 asked in words | Round 2 given the formula | Round 3 the expression it asked YOU to compute |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 3 | |||
| 4 | |||
| 5 |
Round 1 — asked in words. Five runs, five different answers, spanning five orders of magnitude, none correct (the answer is 1,665,429.58):
Run 4 is the interesting one: it wrote the right formula, 4817 * 293 * (1 + 0.18), and still produced a wrong number.
Round 2 — handed the formula. Wrong every time, and now the wrong answers cluster near the right size, which is worse because they look believable: 999,999.8 · 1,286,483.4 · 1,240,908.4 · 115,085.4 · 1,194,414.8. So this was never a wording problem. Nothing in the model multiplies anything — it predicts the most plausible-looking continuation, and for long multiplication that is a number of roughly the right magnitude.
Round 3 — you as the tool. In our runs the model produced a clean JSON request every time, and the expressions varied: (4817*293)+(4817*293*0.18) — correct — in one run, 4817*293 + 4817*0.18 — wrong — in others. After the result was handed back, it finished correctly every time.
So the two failures are different. Rounds 1–2 are a computation failure, and handing the arithmetic to something that actually computes removes it completely — in Round 3 that something was you. What remains is a planning failure: deciding which expression to compute. That one stays with the model, and it is what page 03 measures.
4817*293 + 18/100 in one run, 4817*293 + 18*293 in another. Arithmetic moved into the tool; planning stayed with the model, and that is what page 03 measures. Bigger models get this particular product right more often, which is also not the point: a token predictor is non-deterministic and unauditable: you cannot promise a controller that the number will be right next month, and you cannot show the work. A tool call gives you exactness, repeatability, and a trace — three things a finance team will actually ask for.A model without tools is a very expensive autocomplete. Every tool you add is a verb the agent acquires: read, compute, file, send, pay. The business value of an agent is almost entirely the value of its verbs.
Every verb is also a way to be wrong at scale. A read tool leaks; a write tool destroys; a spend tool costs money. The schema you write is where “what is this agent permitted to do?” actually gets answered — in code, not in a policy document.
Because every tool call is a structured record — name, arguments, result, timestamp — a tool-using agent is far easier to audit than one that “just knows.” When compliance asks what happened, you replay the calls.
Background, not a prerequisite — you can do pages 01–03 without it. It answers the question one step behind this week: why give a model tools at all?
A model's weights were frozen on a date. It cannot know your Q3 numbers, today's outage, or a policy you published last Tuesday. Closed by: a retrieval or search tool.
Your CRM, your ticket queue, your inventory table were never in anyone's training data — and should not have been. Closed by: a read tool against your own systems.
Predicting the next token is not arithmetic, and it is certainly not running a SQL query or a pricing model. Closed by: a calculator, a code runner, a query tool.
Text cannot file a ticket, send an email, or issue a refund. Closed by: write tools — and this is exactly where the risk arrives.