Skip to content
TokIQ
Agents

Tool Calling and AI Agents: How Function Calling Really Works

How LLM tool calling works, how to write tool descriptions models use correctly, how agent loops run and stop, and the failures to plan for in agents.

TokIQ Editorial5 min read
In this article
  1. What does a tool definition look like?
  2. How do you write tool descriptions the model uses correctly?
  3. How does an agent loop work?
  4. How should an agent decide when to stop?
  5. What are the most common agent failures?
  6. What about security?
  7. Do you even need an agent?
  8. Where does MCP fit?

Tool calling (or function calling) lets a language model ask your application to run a function: you describe the available tools with names, descriptions and parameter schemas, the model replies with a structured call such as get_order_status(order_id="A-1042"), your code runs it, and the result goes back to the model. An agent is that same mechanism in a loop, where the model keeps choosing tools and reading results until it decides the task is finished.

The model never executes anything itself. It produces text that happens to be a well-formed request. Keep that in mind and most design decisions about agents become clearer, especially the security ones.

What does a tool definition look like?

Across OpenAI, Anthropic and Google the request formats differ slightly, but every tool definition has the same three parts: a name, a natural-language description and a JSON Schema for the parameters.

A weak definition:

{
  "name": "search",
  "description": "Searches for stuff.",
  "parameters": {
    "type": "object",
    "properties": { "q": { "type": "string" } }
  }
}

A better one:

{
  "name": "search_help_center",
  "description": "Full-text search over the public help center articles for the Ledgerly app. Use this when the user asks how a feature works or how to fix a problem. Does not search the user's own account data; use get_account_summary for that. Returns up to 5 articles with title, URL and a short excerpt.",
  "parameters": {
    "type": "object",
    "properties": {
      "query": {
        "type": "string",
        "description": "A short search phrase in English, e.g. 'export invoices to CSV'. Rephrase the user's question into keywords rather than passing it verbatim."
      }
    },
    "required": ["query"]
  }
}

The description is a prompt. The model reads it to decide when to use the tool, what to pass, and what to expect back. Anthropic's tool use documentation stresses detailed descriptions as the most important factor in tool performance, and that matches what you see in practice.

How do you write tool descriptions the model uses correctly?

Answer the questions a new developer would ask before calling an unfamiliar API:

  • What does it do, and what does it not do? "Does not search the user's account data" prevents a whole class of wrong calls.
  • When should it be used instead of a similar tool? Overlapping tools are the leading cause of wrong tool selection. If two tools could plausibly handle a request, say which one wins.
  • What do parameters mean, with examples? Formats (ISO dates, cents vs dollars, IDs with or without prefix) and an example value for anything ambiguous.
  • What does it return? Including what an empty result means.
  • Does it have side effects? "Sends the email immediately; cannot be undone" changes how a careful model behaves, and should change how your app gates it.

Name tools so they read clearly next to each other: search_help_center, get_account_summary, create_support_ticket. Avoid generic names like search or run that invite misuse.

Return useful errors, too. If a call fails, a tool result like "error": "order_id must look like A-1234; got 1234" lets the model fix its own call on the next step. A bare 500 does not.

How does an agent loop work?

The basic loop, which most agent frameworks implement with variations:

  1. Send the model the conversation, the system prompt and the tool definitions.
  2. If the model replies with text only, the turn is done.
  3. If it replies with one or more tool calls, your code validates and runs them.
  4. Append the results to the conversation as tool results.
  5. Go back to step 1.

This pattern, interleaving reasoning with actions and observations, was described in the ReAct paper (Yao et al., 2022), and it is the backbone of most agents built since. Modern APIs support it natively, including multiple tool calls in a single turn.

The loop is simple. Getting it to behave is where the work is.

How should an agent decide when to stop?

Two kinds of stopping conditions, and you need both.

Task-level, in the prompt. Tell the agent what "done" means:

You are done when the user's question is answered or you have
determined it cannot be answered with the available tools.
If after two searches you have not found relevant articles, stop
searching and tell the user what you tried.
Ask the user a clarifying question instead of guessing when the
request could refer to more than one account or invoice.

Without guidance like this, agents tend to either stop too early (answering from their own knowledge without searching) or keep going (re-searching with slight variations, hoping for a better result).

Hard limits, in code. A maximum number of iterations, a token or cost budget, and a wall-clock timeout. When a limit is hit, end gracefully: return what was found so far and say the task was not completed. Never rely on the model alone to stop a loop that spends money.

What are the most common agent failures?

Wrong tool selection. Usually caused by overlapping tools or vague descriptions. Fix the descriptions first; if that is not enough, reduce the number of tools available for this task.

Hallucinated arguments. The model invents an order ID because the user said "my last order". Instruct it to look up IDs with a tool or ask the user, and validate arguments in code before executing.

Looping. The agent repeats a failing call with tiny variations. Detect repeated identical or near-identical calls in your loop and break out, and make tool errors informative enough to fix.

Premature success. The agent reports "Done, I've updated the record" when the update call actually failed. Make tool results explicit about success or failure, and for important tasks, verify the outcome with a separate read call.

Context bloat. Tool results can be huge (a full web page, a 2,000-row query). Long loops fill the context with stale output and the model loses track of the original goal. Trim or summarize tool results before returning them, and return only fields the model needs.

Too many tools. Accuracy tends to drop as you add tools, especially similar ones. Give each task the smallest useful set, or load tools dynamically based on the request.

What about security?

Every tool is a capability that a prompt injection can borrow. If your agent reads web pages or emails and also has a tool that sends messages, an attacker who writes the right text into a page may get it to send messages.

Practical rules:

  • Authorize in code, not in the prompt. The tool should run with the end user's permissions, so a misled model cannot reach data the user could not.
  • Confirm consequential actions. Payments, deletions, outbound messages, permission changes: show the user exactly what will happen and require approval.
  • Validate arguments. Schema validation plus business rules (amount limits, allowed recipients) before execution.
  • Separate read and write tools. A research step that only has read tools cannot damage anything, however confused it gets.

Do you even need an agent?

Often not. If the steps are known in advance (retrieve, then summarize, then format), a fixed pipeline of individual model calls is cheaper, faster, easier to test and easier to debug. Agents earn their complexity when the path genuinely depends on intermediate results: the model needs to look something up to know what to look up next.

A reasonable progression: start with a single prompt; add retrieval if facts are needed (what RAG is); add a single tool call for one action; only then build a loop.

Where does MCP fit?

The Model Context Protocol, introduced by Anthropic in late 2024 and since adopted by other vendors and tools, standardizes how applications expose tools and data to models. It solves the plumbing: one server definition usable from many clients. It does not change any of the advice above. A badly described tool is just as confusing over MCP, and connecting many third-party servers adds both capabilities and attack surface.

For the output side of tool calls (schemas, strict modes, validation), see getting reliable JSON from LLMs. The tools and agents topic breaks these design decisions into practice questions.

Frequently asked questions

What is tool calling in LLMs?

Tool calling, also called function calling, lets a model respond with a structured request to run a function you defined, such as searching a database or sending an email. Your code runs the function and returns the result to the model, which then continues.

Does the model execute the tool itself?

No. The model only outputs the tool name and arguments. Your application decides whether to run it, runs it, and sends back the result, which is why permission checks belong in your code.

What is the difference between tool calling and an AI agent?

Tool calling is a single capability: the model asks for one function to be run. An agent is a loop in which the model repeatedly decides which tool to call next, observes results, and continues until the task is done or a limit is hit.

How many tools can you give an agent?

There is no fixed limit, but accuracy tends to drop as tool lists grow and tools overlap. Keep the set small and distinct for each task, or load relevant tools dynamically.

  • #tool calling
  • #function calling
  • #AI agents
  • #agent loop
  • #MCP

Now practice it

TokIQ turns prompt engineering into short quizzes, with an explanation for every answer.

Coming soon onApp StoreComing soon onGoogle Play

Write better prompts, a few questions a day.

Short quizzes on real prompting decisions, with an explanation for every answer. Free to start on iPhone and Android.

Coming soon onApp StoreComing soon onGoogle Play