Last week you installed Ollama and downloaded a model. This page turns that model into an agent that calls real Python tools — no API key, no cost, nothing leaves your machine. Then you measure how reliable it actually is, and use tool design to make it better.
requirements.txt, and a README. Or view the files one by one in the materials folder.Most of you have qwen2.5:0.5b; some of you picked a different model to fit your laptop and your final project. Either is fine — but first confirm the model supports tool calling.
That model cannot take part in this lab. Pull one that lists tools — for example ollama pull qwen2.5:0.5b (398 MB). If you downloaded a model from Hugging Face, prefer a name containing Instruct: a base model has not been trained to follow the tool-call format.
No. The line means the model's chat template knows the format. It does not mean the model will use tools well. In our tests below, one model listed tools and never called a single tool. You find out by measuring — Step 3.
ollama serve in a terminal).ollama, langgraph, langchain-ollama). Tested on Python 3.11.
1_agent_raw.py and 3_agent_langgraph.py and change the MODEL = ... line near the top to your model's exact name. Scripts 2 and 4 take the model name on the command line.A terminal is always standing inside one folder. Commands act on that folder. When you open it, it starts in your home folder — and your Desktop is inside your home folder. So if you save the file to your Desktop, getting there is one word:
How to tell it worked: look at the prompt itself — it shows the folder you are standing in.
Two errors everyone hits once:
cd: no such file or directory: Desktop — you are already in Desktop and tried to go in again. Look at the prompt; if it says Desktop, you are done. Type cd ~ to go home if you want to start over.ls prints nothing — the folder is empty, so the file is not downloaded yet. Do that first, then run ls again and look for 1_agent_raw.py.You should see 1_agent_raw.py listed. If you do, every command below will work. Three more things worth knowing:
cd .. goes back up one folder. cd ~ (Mac) or cd ~ (Windows) returns you home so you can start over.cd and a space, then drag the folder from Finder / File Explorer into the terminal window — the full path types itself.cd Desktop says it cannot find the path, your Desktop is synced to OneDrive. Open the folder in File Explorer, click the address bar, type powershell, press Enter — that opens a terminal already inside it.One more thing before the commands below. Where the code blocks say python, type python3 on Mac or py -3 on Windows. And if you built a virtual environment in Step 1, either activate it or call it directly — ./venv/bin/python 1_agent_raw.py … on Mac.
Stuck anyway? The starter zip also contains run_mac.command and run_windows.bat — double-click one and it does all of this for you. Use it to keep moving, then come back and learn the two commands above; you will need them for your final project.
No framework. The whole agent is about twenty lines. Click any part of the code — the panel beside it explains what that piece does and why it is there.
Click a line on the left, or pick a step above. Nine pieces make up the entire agent — nothing else is hiding.
Run it:
Real output from qwen2.5:0.5b. Now run it with no question — the default is a harder, multi-step invoice question — and compare.
A single successful run tells you almost nothing (Week 8 makes this formal). This script asks the multi-step invoice question k times and checks each final answer against the exact value, 1,798,663.95 USD (4817 × 293 × 1.18 × 1.08).
What we measured — same question, same starter code, default settings, run on the instructor's laptop in September 2026. Speed depends on your hardware; correctness depends mostly on the model.
| Model | Download | Invoice question | What went wrong |
|---|---|---|---|
qwen2.5:0.5b | 398 MB | 0 / 16 | Calls tools, but plans the math wrong — "plus 18% tax" became + 18/100 or + 1800 — and usually skips the exchange rate. |
llama3.2 (3B) | 2.0 GB | 0 / 5 | Skips the calculator and does the arithmetic in its head (wrongly); sometimes writes the tool call as plain text instead of making it. |
deepseek-r1:14b | 9.0 GB | 0 / 3 | Lists tools in its capabilities — and never called one. It reasoned the whole problem out in text. |
| a 27B Qwen-family model | 16 GB | 3 / 3 | The same clean path every time: calculator → exchange rate → calculator. Too large for most laptops. |
The small model failed at planning, not at calling tools. So what if the tool did the planning? This script runs three setups on your model and reports two numbers for each: did a tool produce the right number, and is the final answer right.
| Setup | Tools the model gets | qwen2.5:0.5b, measured |
|---|---|---|
| A. Single step — "What is 4817 × 293?" | calculator, get_exchange_rate | 16 / 16 |
| B. Multi-step invoice question | same two fine-grained tools | 0 / 16 |
| C. Same invoice question | one coarse tool: invoice_total(units, unit_price, tax_rate_percent, from_currency, to_currency) | tool right 12 / 15 final answer right 6 / 23 |
Pooled over several runs. Your numbers will vary from run to run — that variation is itself a result worth reporting.
Look at setup C closely. The model filled in all five arguments correctly almost every time, and the tool returned 1798663.95. Then, writing its final answer, the model reported 179,866.39 or 17,986,639.50 — it moved the decimal point while copying the number.
ollama show lists "tools" for your model. Is that enough to build your final project on it?deepseek-r1:14b lists it and never called a tool on our task. The evidence you need is a pass rate on your project's own questions, measured over several runs — which is exactly what scripts 2 and 4 give you, once you swap in your tools and your questions.Pool your models. Each person runs python 4_tool_design_ladder.py <your-model> 3 on their own laptop — the group will likely have qwen2.5:0.5b plus a few different models. Combine the results into one table: model, download size, and the A / B / C numbers. Then answer together: (1) Which model would you build your final project on, and what evidence from the table supports it? (2) For the weakest model in your group, name one tool-design change that would raise its score, in the spirit of setup C.
Produce: the pooled table, plus three sentences: your model choice with its evidence, and one design change for the weakest model.
The invoice agent had one job. A real agent carries several tools, and the model has to pick — which is where page 02’s triage becomes concrete. This one is a customer-service assistant for an online store:
lookup_order(order_id)Read-only. It cannot change anything, so waiting for a human would only slow the customer down. The harness just runs it.
issue_refund(order_id, amount_usd, reason)It moves money and cannot be undone. The harness stops, prints exactly what is about to happen, and waits for a person to type y.
Run all three cases and watch which tool the model picks — and where it is made to wait:
This agent did not work on the first try. The numbers below are real runs of the three cases on qwen2.5:0.5b, changing only the instructions — no change to the model, no change to the tool code.
| What was changed | “Where is my order?” should call lookup_order | “I want a refund” should call issue_refund | “What’s your return policy?” should call NOTHING |
|---|---|---|---|
| 1 · Plain description, plain system prompt | 1 / 4 | 4 / 4 | 4 / 4 |
| 2 · Stronger tool docstring (“never answer from memory”) | 2 / 6 | — | 6 / 6 |
| 3 · System prompt: “you MUST call lookup_order” | 5 / 6 | 6 / 6 | 2 / 6 |
| 4 · Same, plus “if there is no order number, do NOT call any tool” | 5 / 6 | 6 / 6 | 6 / 6 |
Attempt 3 is the trap. Pushing the model to call tools fixed the question it was ignoring and broke the question it had been getting right — it started offering refunds to people who only asked about the return policy. That is over-triggering, named on page 02, and the fix was not a louder instruction but an explicit rule for when not to act.
llama3.2, a model five times larger. It scored 12 / 24, against 22 / 24 for the 0.5B model — it called a tool for almost every message, including the general questions. Why might a bigger model do worse here, and what does that mean for choosing one?Break it, then fix it. Open 5_demo_order_agent.py and delete the last line of the system prompt — the one that says not to call a tool when there is no order number. Ask all three questions again, three times each, and record which tool got called. Then put the line back and confirm the behaviour returns. Finally, add a third tool of your own (for example cancel_order, which should also be gated) and test whether the model still picks correctly among three.
Produce: your before/after counts for the three questions, and one sentence on what adding a third tool did to the model's accuracy.
Everything above works without a framework, and that is deliberate — you can read the whole loop. This step is your first look at LangGraph, which gives those same pieces names: a model node, a tools node, a conditional edge, and one edge back that makes it a loop. Compare it to the plain version and decide for yourself what the framework bought. Reference: the Week 4 LangGraph page (written for Gemini — running locally, the only line that changes is the model).
Plain loop (script 1) is best for seeing and debugging exactly what happens. LangGraph (script 3) earns its place as the agent grows: several tools, a human-approval step, memory, and a trace you can replay. Both call the same local model through the same Ollama server.
ollama serve in a separate terminal, then try again.ollama show <model>; if tools is not under Capabilities, pull a model that has it, such as qwen2.5:0.5b.ollama list exactly, including the tag after the colon — qwen2.5:0.5b, not qwen2.5.pip install -r requirements.txt again.