Skip to content
Huseyin Babal
Go back

How AI agents actually work: what happens on a tool call

This post is the short version of my video. Watch it here: How AI Agents Actually Work: What Happens on a Tool Call

I asked a language model: do I need an umbrella in Istanbul today? I also gave it one tool, get_weather. This is everything it sent back:

{
  "role": "assistant",
  "content": "",
  "tool_calls": [
    {
      "id": "call_ea8moqeo",
      "function": {
        "index": 0,
        "name": "get_weather",
        "arguments": {
          "city": "Istanbul"
        }
      }
    }
  ]
}output

No weather. Just the name of a function and one argument. The model didn’t call anything, and it can’t. All it can do is write text.

So who runs the tool? Everything below runs on my laptop, with qwen3:8b on Ollama, and every output is real.

The request

One message from the user and a list of tools. A tool is a name, a description and a JSON schema for its parameters:

cat weather-request.json
{
  "model": "qwen3:8b",
  "think": false,
  "stream": false,
  "messages": [{"role": "user", "content": "Do I need an umbrella in Istanbul today?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Current weather for a city",
      "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"]
      }
    }
  }]
}
curl -s localhost:11434/api/chat -d @weather-request.json | jq .message

The content comes back empty, and instead there’s a tool_calls entry: the output at the top of this post. To see where that comes from, look at what a model really does.

A language model does one thing

You give it some text, and it gives you a probability for every possible next token. A token is a word or a piece of a word; this model has 151,936 of them. Text in, probabilities out. No memory, no internet, no tools.

A small script asks the model to write exactly one token and to return the five most likely candidates:

sed -n 17,24p next_token.py
for _ in range(steps):
    body = {"model": "qwen3:8b", "prompt": text, "raw": True, "stream": False,
            "options": {"num_predict": 1, "temperature": 0},   # write just one token
            "logprobs": True, "top_logprobs": 5}               # with the 5 most likely ones
    reply = json.load(urllib.request.urlopen(URL, json.dumps(body).encode()))
    top = reply["logprobs"][0]["top_logprobs"]
    show(text, top)
    text += top[0]["token"]                                    # add the top one, ask again
python3 next_token.py "The capital of France is"
 67.4%  ' Paris'
  8.9%  ' a'
  1.9%  ' in'
  1.4%  ' __'
  1.4%  ' ______'output

Now loop: take the top token, add it to the text, ask again.

python3 next_token.py "The capital of France is" 6
'The capital of France is'
    → ' Paris' 67%    (' a' 9%  ' in' 2%)
'The capital of France is Paris'
    → '.' 67%    (',' 15%  '.\n' 12%)
'The capital of France is Paris.'
    → ' The' 49%    (' What' 19%  ' Which' 6%)
'The capital of France is Paris. The'
    → ' capital' 93%    (' E' 1%  ' population' 1%)
'The capital of France is Paris. The capital'
    → ' of' 100%    (' is' 0%  ' city' 0%)
'The capital of France is Paris. The capital of'
    → ' Italy' 38%    (' Germany' 35%  ' Spain' 6%)output

Look at the last step: Italy 38%, Germany 35%. That’s almost a coin flip. Normally the next token is sampled from these probabilities, and that’s why the same question can give you different answers.

A chat is one long text

There is no chat inside the model. Before it sees a conversation, the conversation is turned into one piece of text, with special tokens marking where each message starts and ends. And the model remembers nothing between requests: the whole conversation is sent again every time.

Ollama can show that text. This is what the model really received for the umbrella question:

python3 raw.py weather-request.json
<|im_start|>system


# Tools

You may call one or more functions to assist with the user query.

You are provided with function signatures within <tools></tools> XML tags:
<tools>
{"type": "function", "function": {get_weather Current weather for a city {object <nil> <nil> [city] {"city":{"type":"string"}}}}}
</tools>

For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:
<tool_call>
{"name": <function-name>, "arguments": <args-json-object>}
</tool_call><|im_end|>
<|im_start|>user
Do I need an umbrella in Istanbul today? /no_think<|im_end|>
<|im_start|>assistant
<think>

</think>output

A system message with the tool pasted in as text, instructions to write a JSON object inside <tool_call> tags, the user message, and an open assistant turn. The tool line looks odd because that’s how this Ollama version renders the tool definition into the prompt; the token count matches the chat request exactly, so it is what the model gets.

Now send exactly this text as a plain prompt. No chat API, no tools:

python3 raw.py weather-request.json --run
<tool_call>
{"name": "get_weather", "arguments": {"city": "Istanbul"}}
</tool_call>output

A tool call is just text the model was told to write. Ollama finds the tags and turns them into the JSON from the top of this post.

So who runs the tool?

Your code. The model writes the call, your program reads it and runs the real function. Then it adds two messages to the conversation, the model’s tool call and the result, and sends everything back.

By hand, with me as the program:

curl -s 'wttr.in/Istanbul?format=%C,+%t'
Cloudy, +20°Coutput
jq -c '.messages[]' weather-followup.json
{"role":"user","content":"Do I need an umbrella in Istanbul today?"}
{"role":"assistant","content":"","tool_calls":[{"function":{"name":"get_weather","arguments":{"city":"Istanbul"}}}]}
{"role":"tool","tool_name":"get_weather","content":"Cloudy, +20°C"}output
curl -s localhost:11434/api/chat -d @weather-followup.json | jq -r .message.content
The weather in Istanbul today is cloudy with a temperature of +20°C. You might not need an umbrella, but it's a good idea to carry one just in case.output

In the raw text, the result is one more block, inside <tool_response> tags:

python3 raw.py weather-followup.json | tail -n 12
<|im_start|>assistant
<tool_call>
{"name": "get_weather", "arguments": {"city":"Istanbul"}}
</tool_call><|im_end|>
<|im_start|>user
<tool_response>
Cloudy, +20°C
</tool_response><|im_end|>
<|im_start|>assistant
<think>

</think>output

To the model, it’s all just text.

An agent is that, in a loop

Call the model. If it writes a tool call, run it, add the result, and call the model again. If it writes plain text, you’re done. The model decides what to do next, and your code does it.

The tools often come from an MCP server. MCP, the Model Context Protocol, is a standard way for a program to say: these are my tools, and this is how you call them. This one has two:

"""An MCP server with two tools. Files are limited to ./project."""
import pathlib
from mcp.server.mcpserver import MCPServer

ROOT = pathlib.Path(__file__).parent / "project"
server = MCPServer("workspace")

@server.tool()
def list_files(path: str = ".") -> str:
    """List the files in a project folder."""
    return "\n".join(sorted(p.name for p in (ROOT / path).iterdir()))

@server.tool()
def read_file(path: str) -> str:
    """Read a text file from the project."""
    return (ROOT / path).read_text()

if __name__ == "__main__":
    server.run()tools_server.py

And the agent loop. It asks the MCP server for its tools, then calls the model until there’s no tool call left:

tools = [as_tool(t) for t in (await mcp.list_tools()).tools]
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": task}]
for step in range(1, 11):                       # the agent loop
    msg, tokens = llm(messages, tools)
    print(f"step {step}: {len(messages)} messages, {tokens} tokens sent")
    messages.append(msg)
    if not msg.get("tool_calls"):               # no tool call: the model is done
        return print(f"\n{msg['content']}")
    for call in msg["tool_calls"]:
        name, args = call["function"]["name"], call["function"]["arguments"]
        result = (await mcp.call_tool(name, args)).content[0].text
        print(f"  → {name}({json.dumps(args)})  ← {len(result)} chars")
        messages.append({"role": "tool", "tool_name": name, "content": result})agent.py
python agent.py "What does this project do, and what do I need to install to run it?"
step 1: 2 messages, 216 tokens sent
  → list_files({"path": "."})  ← 37 chars
step 2: 4 messages, 256 tokens sent
  → read_file({"path": "README.md"})  ← 200 chars
step 3: 6 messages, 334 tokens sent

The project is a command-line tool called `weather-cli` that displays today's weather for a given city. To run it, you need to:

1. **Install dependencies**: Run `pip install -r requirements.txt` to install the required packages.
2. **Set up an API key**: Set the `WEATHER_API_KEY` environment variable with your weather API key (e.g., OpenWeatherMap).
3. **Run the tool**: Use `python weather.py <city-name>` to get the weather for a specific city.output

Three calls to the model. Look at the tokens: 216, 256, 334. Every step sends the whole conversation again, plus everything the tools returned. That’s why long agent sessions get slow and expensive.

One note on that run: the answer text varies between runs and I picked a short one for the screen. The steps and the token counts were the same every time.

The whole thing in one paragraph

A language model predicts the next token. A chat is one long text. A tool is a description inside that text, and a tool call is text the model writes. Your code runs the tool and feeds the result back. Do that in a loop, and you have an agent.

Sources


Share this post:

Previous Post
What really happens when you kubectl apply
Next Post
Trying Git 3.0 before it exists: SHA-256, reftable and what broke