Georgia State University — J. Mack Robinson College of Business PATH — Pathways for AI Training & Hiring CIS 4394 Agentic AI  ·  Fall 2026  ·  Dr. Xinyu Fu
03 · Hands-on lab

Build a tool-calling agent on your own laptop.

Last week you installed Ollama and downloaded a model. This page turns that model into an agent that calls real Python tools — no API key, no cost, nothing leaves your machine. Then you measure how reliable it actually is, and use tool design to make it better.

📦 Download the starter code: CIS4394_Week5_Ollama_Starter.zip ↓ — five short Python scripts, a requirements.txt, and a README. Or view the files one by one in the materials folder.
📄 In-class lab sheet (pairs): CIS4394_Week5_Student_Lab.pdf ↓ — Mac and Windows commands side by side, with the observation tables to fill in.
Step 0

Check that your model can call tools

Most of you have qwen2.5:0.5b; some of you picked a different model to fit your laptop and your final project. Either is fine — but first confirm the model supports tool calling.

$ ollama list $ ollama show qwen2.5:0.5b # use your exact model name from `ollama list` Capabilities completion tools <- this line must be here
No "tools" line?

That model cannot take part in this lab. Pull one that lists tools — for example ollama pull qwen2.5:0.5b (398 MB). If you downloaded a model from Hugging Face, prefer a name containing Instruct: a base model has not been trained to follow the tool-call format.

"tools" is listed — are you done?

No. The line means the model's chat template knows the format. It does not mean the model will use tools well. In our tests below, one model listed tools and never called a single tool. You find out by measuring — Step 3.

Step 1

Set up once

  1. Make sure Ollama is running. Open the Ollama app (or run ollama serve in a terminal).
  2. Unzip the starter code and open a terminal in that folder.
  3. Create a virtual environment and install the three packages (ollama, langgraph, langchain-ollama). Tested on Python 3.11.
    # Mac / Linux python3 -m venv venv source venv/bin/activate pip install -r requirements.txt
    # Windows PowerShell python -m venv venv venv\Scripts\Activate.ps1 pip install -r requirements.txt
    # Windows cmd python -m venv venv venv\Scripts\activate pip install -r requirements.txt
  4. Using a model other than qwen2.5:0.5b? Open 1_agent_raw.py and 3_agent_langgraph.py and change the MODEL = ... line near the top to your model's exact name. Scripts 2 and 4 take the model name on the command line.
First, a 60-second terminal skill you will use all semester.

A terminal is always standing inside one folder. Commands act on that folder. When you open it, it starts in your home folder — and your Desktop is inside your home folder. So if you save the file to your Desktop, getting there is one word:

Mac — Terminal
cd Desktop # move into the folder ls # what's in here?
Windows — PowerShell
cd Desktop # move into the folder dir # what's in here?

How to tell it worked: look at the prompt itself — it shows the folder you are standing in.

xinyufu@MacBook ~ % cd Desktop # before: you are home xinyufu@MacBook Desktop % ls # after: the prompt changed. You are there.

Two errors everyone hits once:

  • cd: no such file or directory: Desktop — you are already in Desktop and tried to go in again. Look at the prompt; if it says Desktop, you are done. Type cd ~ to go home if you want to start over.
  • ls prints nothing — the folder is empty, so the file is not downloaded yet. Do that first, then run ls again and look for 1_agent_raw.py.

You should see 1_agent_raw.py listed. If you do, every command below will work. Three more things worth knowing:

  • cd .. goes back up one folder. cd ~ (Mac) or cd ~ (Windows) returns you home so you can start over.
  • Shortcut: type cd and a space, then drag the folder from Finder / File Explorer into the terminal window — the full path types itself.
  • Windows note: if cd Desktop says it cannot find the path, your Desktop is synced to OneDrive. Open the folder in File Explorer, click the address bar, type powershell, press Enter — that opens a terminal already inside it.

One more thing before the commands below. Where the code blocks say python, type python3 on Mac or py -3 on Windows. And if you built a virtual environment in Step 1, either activate it or call it directly — ./venv/bin/python 1_agent_raw.py … on Mac.

Stuck anyway? The starter zip also contains run_mac.command and run_windows.bat — double-click one and it does all of this for you. Use it to keep moving, then come back and learn the two commands above; you will need them for your final project.

Step 2 · 1_agent_raw.py

The whole agent loop, in plain Python

No framework. The whole agent is about twenty lines. Click any part of the code — the panel beside it explains what that piece does and why it is there.

1_agent_raw.py — click a line
import ollama def calculator(expression: str) -> str: """Evaluate an arithmetic expression and return the exact result. Args: expression: e.g. '4817 * 293' """ ... TOOLS = {"calculator": calculator, "get_exchange_rate": get_exchange_rate} for step in range(1, MAX_STEPS + 1): # step cap resp = ollama.chat(model=MODEL, messages=messages, tools=list(TOOLS.values())) msg = resp.message messages.append(msg) if not msg.tool_calls: # model chose to FINISH return msg.content for call in msg.tool_calls: # model chose to ACT name, args = call.function.name, call.function.arguments try: result = TOOLS[name](**args) # YOUR code runs it except Exception as e: result = f"ERROR calling {name}: {e}" messages.append({"role": "tool", "content": result, "tool_name": name})

Start anywhere

Click a line on the left, or pick a step above. Nine pieces make up the entire agent — nothing else is hiding.

Four sentences that survive this page.
  1. The model requests; your code executes.
  2. Tool results go back into the conversation as observations.
  3. The loop continues until the model stops asking for tools.
  4. Your code — not the model — controls the boundaries: the step cap, error handling, which tools exist, who approves.

Run it:

$ python 1_agent_raw.py "What is 4817 * 293?" [step 1] model requests -> calculator({'expression': '4817 * 293'}) harness ran it <- 1411381 [step 2] FINAL ANSWER: The result of 4817 * 293 is 1411381.

Real output from qwen2.5:0.5b. Now run it with no question — the default is a harder, multi-step invoice question — and compare.

Concept check

Map the code to the round trip

Step 3 · 2_pass_k.py

Don't trust one run. Measure.

A single successful run tells you almost nothing (Week 8 makes this formal). This script asks the multi-step invoice question k times and checks each final answer against the exact value, 1,798,663.95 USD (4817 × 293 × 1.18 × 1.08).

$ python 2_pass_k.py qwen2.5:0.5b 3 # your model name, number of runs

What we measured — same question, same starter code, default settings, run on the instructor's laptop in September 2026. Speed depends on your hardware; correctness depends mostly on the model.

ModelDownloadInvoice questionWhat went wrong
qwen2.5:0.5b398 MB0 / 16Calls tools, but plans the math wrong — "plus 18% tax" became + 18/100 or + 1800 — and usually skips the exchange rate.
llama3.2 (3B)2.0 GB0 / 5Skips the calculator and does the arithmetic in its head (wrongly); sometimes writes the tool call as plain text instead of making it.
deepseek-r1:14b9.0 GB0 / 3Lists tools in its capabilities — and never called one. It reasoned the whole problem out in text.
a 27B Qwen-family model16 GB3 / 3The same clean path every time: calculator → exchange rate → calculator. Too large for most laptops.
Read this table carefully. Every model here "supports tools." Only one completed the task. The mechanics of tool calling working is not the same as the agent being reliable — and the only way to know which one you have is to measure it on your task.
Step 4 · 4_tool_design_ladder.py

Can tool design rescue a small model?

The small model failed at planning, not at calling tools. So what if the tool did the planning? This script runs three setups on your model and reports two numbers for each: did a tool produce the right number, and is the final answer right.

$ python 4_tool_design_ladder.py qwen2.5:0.5b 3
SetupTools the model getsqwen2.5:0.5b, measured
A. Single step — "What is 4817 × 293?"calculator, get_exchange_rate16 / 16
B. Multi-step invoice questionsame two fine-grained tools0 / 16
C. Same invoice questionone coarse tool: invoice_total(units, unit_price, tax_rate_percent, from_currency, to_currency)tool right 12 / 15
final answer right 6 / 23

Pooled over several runs. Your numbers will vary from run to run — that variation is itself a result worth reporting.

Look at setup C closely. The model filled in all five arguments correctly almost every time, and the tool returned 1798663.95. Then, writing its final answer, the model reported 179,866.39 or 17,986,639.50 — it moved the decimal point while copying the number.

Discussion question 1In setup C the tool returned the right number and the model still reported a wrong one. Where should the final number come from?
From the tool, not from the model's retelling of it. Critical values — money, dates, IDs, quantities — should be shown to the user straight from the tool's output (or returned as a structured field your code reads), with the model only writing the sentence around them. A second defense is a check in the harness: compare every number in the final answer against the tool results in the trace and flag any that do not match. Small models are good at choosing and filling in a tool; do not also make them responsible for copying digits.
Discussion question 2Setup C moved the plan — multiply, add tax, convert — out of the model and into your Python function. Is this still an agent, or has it become a workflow?
For this question, it is much closer to a workflow: you, the developer, decided the steps, and the model only recognizes the request and fills in five arguments. That is not a defeat — Week 1's rule is to choose the lowest level of autonomy that solves the problem, and a small local model plus a well-designed tool is cheap and far more reliable than setup B. The agentic part comes back when the requests vary and the model has to decide which of several tools to use, in what order, and when to stop. The design question for your project is how much planning to leave to your model, given what you measured it can do.
Discussion question 3ollama show lists "tools" for your model. Is that enough to build your final project on it?
No. The capability line means the chat template supports the tool-call format. deepseek-r1:14b lists it and never called a tool on our task. The evidence you need is a pass rate on your project's own questions, measured over several runs — which is exactly what scripts 2 and 4 give you, once you swap in your tools and your questions.
Group task
👥 Group task · Groups of 3–4 · ~25 minutes

Pool your models. Each person runs python 4_tool_design_ladder.py <your-model> 3 on their own laptop — the group will likely have qwen2.5:0.5b plus a few different models. Combine the results into one table: model, download size, and the A / B / C numbers. Then answer together: (1) Which model would you build your final project on, and what evidence from the table supports it? (2) For the weakest model in your group, name one tool-design change that would raise its score, in the spirit of setup C.

Produce: the pooled table, plus three sentences: your model choice with its evidence, and one design change for the weakest model.

A strong group chooses on evidence, not on the model's name or size: "we pick X because it scored 3/3 on setup B on two different laptops," or "we pick the 0.5B model because our project only needs single-step tool calls, where it scored 16/16." It names the trade-off: a bigger model plans better but may be too slow or too large for a teammate's laptop. And its design change is concrete — merge two tools that always run together into one; replace a free-text argument with a fixed set of allowed values; or have the harness print the tool's number instead of the model's copy of it.
Step 5 · 5_demo_order_agent.py

Two tools, two risk levels, one human gate

The invoice agent had one job. A real agent carries several tools, and the model has to pick — which is where page 02’s triage becomes concrete. This one is a customer-service assistant for an online store:

Auto-run

lookup_order(order_id)

Read-only. It cannot change anything, so waiting for a human would only slow the customer down. The harness just runs it.

Human gate

issue_refund(order_id, amount_usd, reason)

It moves money and cannot be undone. The harness stops, prints exactly what is about to happen, and waits for a person to type y.

Run all three cases and watch which tool the model picks — and where it is made to wait:

$ python 5_demo_order_agent.py "Where is my order ORD-004411?" [step 1] model requests -> lookup_order({'order_id': 'ORD-004411'}) result <- status=shipped, items=2, total=$84.20 [step 2] FINAL ANSWER: Your order ORD-004411 is shipped, 2 items, total $84.20.
$ python 5_demo_order_agent.py "I want a refund for order ORD-004412, it arrived damaged. It was $19.99." [step 1] model requests -> issue_refund({'order_id': 'ORD-004412', 'amount_usd': 19.99, 'reason': 'damaged'}) ┌─ APPROVAL NEEDED ───────────────────────────── │ action : issue_refund │ order_id : ORD-004412 │ amount_usd : 19.99 │ reason : damaged └─────────────────────────────────────────────── Approve? [y/N] y result <- REFUNDED $19.99 on ORD-004412 (damaged)
$ # same question, but this time answer n Approve? [y/N] n result <- DENIED by the human reviewer. Do not retry; tell the customer a person will follow up. [step 2] FINAL ANSWER: I can’t process that refund — a person will follow up with you.
Look at what the denial actually is: the refusal goes back into the conversation as an observation, exactly like a tool result. The model is not overruled in secret — it is told what happened and has to carry on from there. That is the whole shape of a human gate: the model proposes, a person disposes, and the agent keeps its trace either way.

Getting a small model to choose well: four attempts

This agent did not work on the first try. The numbers below are real runs of the three cases on qwen2.5:0.5b, changing only the instructions — no change to the model, no change to the tool code.

What was changed“Where is my order?”
should call lookup_order
“I want a refund”
should call issue_refund
“What’s your return policy?”
should call NOTHING
1 · Plain description, plain system prompt1 / 44 / 44 / 4
2 · Stronger tool docstring (“never answer from memory”)2 / 6—6 / 6
3 · System prompt: “you MUST call lookup_order”5 / 66 / 62 / 6
4 · Same, plus “if there is no order number, do NOT call any tool”5 / 66 / 66 / 6

Attempt 3 is the trap. Pushing the model to call tools fixed the question it was ignoring and broke the question it had been getting right — it started offering refunds to people who only asked about the return policy. That is over-triggering, named on page 02, and the fix was not a louder instruction but an explicit rule for when not to act.

Discussion questionThe same final instructions were run on llama3.2, a model five times larger. It scored 12 / 24, against 22 / 24 for the 0.5B model — it called a tool for almost every message, including the general questions. Why might a bigger model do worse here, and what does that mean for choosing one?
Nothing about “bigger” guarantees better tool discipline. Models differ in how eagerly they reach for a tool — a tendency that comes from how each was trained and tuned, not from parameter count — and a prompt tuned to push one model over the line can push another well past it. The practical lesson is the same one Week 8 makes formal: a model is a choice you test, not a choice you rank by size. Run your own three or four representative messages, including at least one that should call no tool, and count. The message that should trigger nothing is the one most teams forget to test, and it is where over-triggering shows up.
👥 Group task · Pairs · ~15 minutes

Break it, then fix it. Open 5_demo_order_agent.py and delete the last line of the system prompt — the one that says not to call a tool when there is no order number. Ask all three questions again, three times each, and record which tool got called. Then put the line back and confirm the behaviour returns. Finally, add a third tool of your own (for example cancel_order, which should also be gated) and test whether the model still picks correctly among three.

Produce: your before/after counts for the three questions, and one sentence on what adding a third tool did to the model's accuracy.

Step 6 · 3_agent_langgraph.py

The same agent in LangGraph

Everything above works without a framework, and that is deliberate — you can read the whole loop. This step is your first look at LangGraph, which gives those same pieces names: a model node, a tools node, a conditional edge, and one edge back that makes it a loop. Compare it to the plain version and decide for yourself what the framework bought. Reference: the Week 4 LangGraph page (written for Gemini — running locally, the only line that changes is the model).

from langchain_ollama import ChatOllama # Week 4 (cloud): # llm = ChatGoogleGenerativeAI(model="gemini-2.0-flash") # Week 5 (local): llm = ChatOllama(model=MODEL, temperature=0) llm = llm.bind_tools(tools) g.add_node("model", call_model) g.add_node("tools", ToolNode(tools)) g.add_edge(START, "model") g.add_conditional_edges("model", guarded) # tool call, finish, or step cap g.add_edge("tools", "model") # <- the loop
$ python 3_agent_langgraph.py
When to use which

Plain loop (script 1) is best for seeing and debugging exactly what happens. LangGraph (script 3) earns its place as the agent grows: several tools, a human-approval step, memory, and a trace you can replay. Both call the same local model through the same Ollama server.

🎯 Take it to your final project

Choosing your model, on evidence

Pick the model you can run comfortably on the laptops your team will demo on — as a rule of thumb, the download size should fit in your computer's memory with plenty to spare — and then measure it on your own task: swap your tools and questions into script 2 and report the pass rate. If a small model fails at multi-step planning, you have two honest options: design coarser tools so it has less to plan (setup C), or use a larger model on a machine that can run it. Either choice is fine if your numbers support it.
Troubleshooting

If something goes wrong

The Ollama server is not running. Open the Ollama app, or run ollama serve in a separate terminal, then try again.
Your model's template has no tool-calling format. Check with ollama show <model>; if tools is not under Capabilities, pull a model that has it, such as qwen2.5:0.5b.
The name must match ollama list exactly, including the tag after the colon — qwen2.5:0.5b, not qwen2.5.
Your virtual environment is not active, or the packages are not installed in it. Activate it (Step 1) and run pip install -r requirements.txt again.
The model wrote the tool call as ordinary text instead of making a structured call, so the harness had nothing to execute. This is a common small-model failure, not a bug in your code — count that run as a failure in your measurements.
The model is too large for your computer's memory. Try a smaller model, close other applications, and use a smaller k (for example 2) while you are testing.
← Previous02 · Schemas, errors & gates