How to Build an AI Agent with Tool Calling from Scratch

6 minUpdated:
How to Build an AI Agent with Tool Calling from Scratch

An agent is a loop: send the conversation and tool definitions to the model, execute any tool calls it returns, append the results, and repeat until it answers without calling a tool or hits a step limit. Most of the real work is tool design, error messages, limits and logging.

What is tool calling?

Tool calling lets a language model ask your code to run a function. You describe each tool with a name, a description and a JSON Schema for its arguments. The model replies either with text or with a structured request to call one or more tools.

The model never runs anything itself. Your program decides whether to execute the call, runs it, and sends the result back. That boundary is where you enforce permissions, validation and limits.

All major hosted model APIs and most open-weight serving stacks support this pattern, with small differences in message format. If you want portability, keep your own internal representation of tool calls and convert to each provider’s format at the edge.

How does the agent loop work?

That is the entire core. Written by hand it is usually well under a couple hundred lines, and writing it once is the best way to understand what frameworks do for you.

  • Build the request: system prompt, conversation history and the list of available tools.
  • Call the model and inspect the response.
  • If it contains tool calls, validate the arguments against the schema, execute each allowed tool and capture the result or error as text.
  • Append the model’s tool call and your tool result to the history, keeping the call identifiers the API expects.
  • Repeat until the model returns a final answer, a step limit is reached, or a budget is exhausted.
  • Return the answer along with a trace of every step for debugging.

How do you design good tools?

The model chooses tools from their names and descriptions, so write them like documentation for a new colleague. Say what the tool does, when to use it, when not to, and what the output looks like.

Prefer a few task-shaped tools over many thin API wrappers. A search_orders tool that accepts a customer and a date range beats three endpoints the model must chain correctly.

Return compact, relevant output. Dumping a whole JSON document wastes context and buries the fact the model needs.

Use enums and required fields wherever you can. Constrained arguments leave less room for creative mistakes, and schema validation catches the rest before anything touches a real system.

Design choiceWeak versionBetter version
Nameapi_call_2get_invoice_status
ArgumentsFree-form query stringTyped fields with enums and required keys
ErrorsStack trace or empty stringPlain message saying what went wrong and how to fix the call
OutputFull raw API responseThe few fields the task needs, with IDs for follow-up calls
Side effectsExecutes immediatelyRead tools run freely; write tools need confirmation or policy checks

How do you handle errors and stop conditions?

Tool errors are part of the conversation. Return them to the model as readable messages, and it will often correct its own arguments on the next step. Crashing the loop on the first bad argument makes the agent fragile.

Always set hard limits: maximum steps, maximum tool calls per step, wall-clock timeout and token or cost budget. Without them a confused agent can loop, repeating the same failing call.

Distinguish retryable failures from permanent ones. A timeout from a flaky API can be retried by your code with backoff before the model ever sees it; a permission error should go straight back to the model, or to the user, with a clear explanation.

Detect repetition explicitly. If the same tool is called with identical arguments several times, stop and report rather than trying forever.

How do you keep a tool-using agent safe?

Treat every tool result as untrusted input. A web page, email or file the agent reads can contain instructions aimed at the model, which is the core of prompt injection.

Scope credentials per tool, not per agent. Require human confirmation for irreversible actions such as sending email, payments or deleting data, and log every call with its arguments.

Run code-executing tools in a sandbox such as a container with no network access and a read-only filesystem by default. Give it only the directories and hosts the task needs, and throw it away after each run.

Check authorization inside the tool, using the identity of the real user, rather than trusting that the model picked a permitted record. The model is a planner, not a security boundary.

When should you use a framework or MCP instead?

Stay with a hand-written loop while you have one agent and a handful of tools. Move to a framework when you need durable state, retries across restarts, multi-agent handoffs, streaming UIs or human-in-the-loop checkpoints.

The Model Context Protocol solves a different problem: packaging tools so many clients can use them. If you want the same tools available in Claude Code, an IDE and your own agent, expose them as an MCP server and keep your loop as an MCP client.

Whatever you pick, keep tools as plain functions with clear inputs and outputs. That keeps them testable and lets you move between a hand-written loop, a framework and MCP without rewriting business logic.

ApproachExamplesBest forTrade-off
Hand-written loopProvider SDK plus your codeLearning, small agents, full controlYou build state, retries and tracing yourself
Graph or workflow frameworkLangGraph, LlamaIndex workflowsStateful, branching, resumable agentsMore concepts and dependencies
Multi-agent frameworkCrewAI, AutoGenRole-based agent teamsHarder to debug and to bound cost
Tool protocolMCP serversReusing tools across clientsAnother process and interface to secure

What does a minimal implementation look like?

Start with a tool registry: a map from tool name to a function, its description and its argument schema. Generate the tool list for the API from that registry so definitions and code can never drift apart.

Write one function that runs a single step: call the model, collect tool calls, validate and execute them, and return the new messages. Wrap it in a loop with a step counter and a budget check. Keep the history as plain data so you can save it, replay it or inspect it in tests.

Add a trace object from the start. For each step record the model input size, the tool calls, their arguments, results, errors and duration. You will use this more than any other piece of code when behavior surprises you.

How do you test and observe a tool-using agent?

Unit-test tools like any other function, then test the agent with mocked tools that return fixed results. That isolates model behavior from flaky external APIs and makes runs repeatable enough to compare prompts.

Build a task set with expected outcomes: which tools should be called, which must not be called and what the final answer must contain. Include ambiguous requests where the right move is to ask a clarifying question.

In production, send traces to an observability tool such as Langfuse or an OpenTelemetry backend. Review failed and expensive runs weekly and turn them into new test cases.

Where does it break? Common mistakes

  • Too many tools in one prompt, which increases wrong tool choices and eats context.
  • Descriptions that overlap, so the model cannot tell two tools apart.
  • Trusting model-generated arguments without schema validation and authorization checks.
  • No trace logging, which makes it impossible to see why the agent did something.
  • Evaluating on a few happy-path prompts instead of a test set with ambiguous and adversarial tasks.

Frequently asked questions

Do I need a framework to build an AI agent?
No. The provider SDK plus a loop, tool registry and limits is enough for many production agents. Frameworks help once you need persisted state, resumable runs, complex branching or several cooperating agents. Writing the loop yourself first makes it much easier to judge what a framework adds.
How many tools can one agent handle?
There is no fixed number, but accuracy tends to drop as tools multiply and descriptions overlap. Keep the active set small and task-specific. If you need many tools, group them, load only the relevant subset for each task, or route to specialized sub-agents with their own short lists.
Can the model call several tools at once?
Many current APIs support parallel tool calls in one response. Execute independent calls concurrently and return every result with its matching call identifier. Be careful with calls that have side effects or depend on each other, and serialize those in your own code when order matters.
How is tool calling different from MCP?
Tool calling is the model-level mechanism for requesting a function call. MCP is a protocol for exposing tools, resources and prompts from a separate server so different clients can discover and use them. An MCP client still uses tool calling underneath when it passes those tools to a model.
Free for builders

Get a hand-picked shortlist of repos for your project

Tell us what you are building. A person — not a bot — reviews it and replies within 48 hours with the catalog projects that fit, including licence and difficulty notes.

We use your email only for this request. Privacy policy

Related guides