How to Build an AI Agent: A Step-by-Step Guide With a Real Example
Learn how to build an AI agent step by step: scoping, tools, the agent loop, guardrails, evaluation and deployment, with a worked lead-qualification example.
Building a demo AI agent takes an afternoon. Building one you can trust with real work takes a process. This guide shows how to build an AI agent the way we do it in client projects, using one running example: an agent that qualifies inbound leads for a B2B startup.
If you want the bigger picture first (architecture, tool choices, and when agents make sense), read our guide to AI agent development.
Step 1: Write the job description
Treat the agent like a new hire. Write down the goal, the inputs, the allowed actions, and what "done" looks like.
Step 2: List the tools
Each tool is a function the model can call, with a clear name, a description, and typed inputs. Keep the list short; every extra tool is another way to go wrong.
| Tool | Purpose | Permission |
|---|---|---|
| lookup_company | Fetch company size, industry, and website | Read-only |
| search_web | Find recent news or funding | Read-only |
| get_icp_rules | Load the scoring rubric | Read-only |
| update_crm | Write score and note to the lead record | Write, limited to this lead |
| send_email | Send a templated reply | Write, approval required at first |
| escalate_to_human | Hand off with a summary | Always allowed |
Step 3: Build an evaluation set
Before writing prompts, pull 50 past demo requests and record the right outcome for each: the score band and the correct email. This is your test suite. Without it you are tuning by feel.
Step 4: Write the agent loop
At its core every agent is the same loop: send the goal and tool list to the model, run whatever tool it asks for, give it the result, repeat until it says it's done. Frameworks like LangGraph, the OpenAI Agents SDK, and the Claude Agent SDK give you this loop plus retries, state, and tracing. Add three limits from day one:
- A step limit (say, 15 tool calls) so it can't loop forever.
- A cost limit per task, enforced at the model gateway.
- A timeout, after which it escalates to a person.
Step 5: Write the system prompt
The system prompt is the job description from Step 1, made precise: the goal, the rubric, the rules ("never email anyone not in the request"), the tone of the emails, and when to escalate. Include two or three worked examples. Keep business rules in data the agent loads through a tool (like get_icp_rules) so non-engineers can update them.
Step 6: Run the evals and read the traces
Run all 50 examples. For each failure, read the trace (every model call and tool result) and classify why it failed: missing information, wrong tool choice, misread rubric, or bad output format. Fix the most common category first, then re-run. Repeat until results stop improving.
Step 7: Add guardrails
- Approval on irreversible actions. At launch,
send_emailcreates a draft a person approves in one click. - Output validation. Check the score is a number from 0 to 100 and the email uses an approved template before anything is saved or sent.
- Narrow permissions. The CRM token can only update the lead being processed.
- Audit log. Every action is recorded with the reasoning that led to it.
Step 8: Deploy in shadow mode, then widen
Trigger the agent from your form or CRM webhook. For the first two weeks it proposes actions and a person approves them. Track approval rate, time saved, and cost per lead. When approvals are consistently a formality, remove the approval step for high-confidence cases and keep it for the rest.
Step 9: Monitor and improve
Keep tracing on, alert on failures and cost spikes, and add every new failure to the evaluation set. When a new model is released, run the evals before switching. This is ongoing AI agent development, and it's what keeps the agent useful.
Tools we use to build AI agents
- Models: Claude, GPT, and Gemini, chosen per task, routed through a gateway such as AiKey.
- Orchestration: LangGraph, OpenAI Agents SDK, Claude Agent SDK, or plain code for simple loops.
- Knowledge: Postgres with pgvector, or a managed vector store, plus keyword search.
- Observability: tracing and evaluation tooling, plus cost dashboards per agent.
- Deployment: serverless functions or containers on AWS, GCP, Azure, Cloudflare, or Vercel.
Build it yourself or get help?
If you have engineering time, the steps above will get you a working agent. If you'd rather have it live in a few weeks with evaluation and guardrails done properly, our AI agent development services follow exactly this process. Curious about budget? See our AI agent development cost guide.
AI agent development services
We scope, build, and monitor production AI agents for startups, with guardrails and evaluation built in.
Explore AI agent development servicesFrequently asked questions
Can I build an AI agent without coding?
+
No-code tools can build simple agents that connect a few apps. For agents that touch production systems, need guardrails, or must be reliable at volume, some code is almost always needed.
What programming language is best for building AI agents?
+
Python and TypeScript have the best agent frameworks and SDKs. Pick the one your team already uses so the agent is easy to maintain.
How many tools should an AI agent have?
+
As few as the job needs. Most effective single agents use between three and eight well-described tools. Past that, consider splitting the work across specialised agents.
How do I test an AI agent?
+
Build an evaluation set of real examples with known correct outcomes, run the agent against it after every change, and read the traces of failures. Measure task success rate, cost per task, and escalation rate.