Deploying AI Agents to Production: What Changes After the Demo
What production AI agents need that demos don't: sessions, durable execution, scoped tool access, tracing, cost limits, staged rollout, and a kill switch.
The demo worked. The agent read the request, called three tools, and produced exactly the right result in front of the whole team. Then someone asked: "Great, when can customers use it?"
That's when the real work starts. A demo runs once, on one machine, with one user, and someone watching. A production agent runs thousands of times, for many users at once, overnight, against systems that time out, while nobody's looking. This guide covers what has to change, based on how we deploy agents at ThinkDeck.
1. Sessions: keep each conversation separate
Each user's task needs its own state: history, intermediate results, pending approvals. Store it in a session store (Postgres or Redis are common), keyed per user and task, never in server memory. Otherwise a restart loses everything, and two users can end up sharing context.
2. Durable execution: survive failures mid-task
Agent tasks can run for minutes and call many tools. Any of those calls can fail. Production agents need:
- A task queue so long jobs don't hang a web request.
- Checkpoints after each step, so a crash resumes from step 7 rather than starting over.
- Retries with backoff for flaky tools, and a limit on how many.
- Idempotent actions. If a retry fires
send_invoicetwice, the customer should still get one invoice.
3. Memory: decide what persists
Short-term state belongs to the session. Long-term memory (preferences, past decisions, corrections) needs its own store, a clear policy on what gets written, and a way for users to see or delete what's kept about them. See context engineering for how memory feeds back into the agent.
4. Tool access: least privilege, per user
- Act as the user, not as an admin. Where possible, tools run with the requesting user's permissions, not a shared super-account.
- Scope every token to the narrowest access that works.
- Keep secrets out of prompts. Credentials live in a secrets manager and are injected into tool calls, never shown to the model.
- Gate irreversible actions behind approval until the agent has a track record.
5. Observability: trace every decision
When an agent does something odd at 3 a.m., you need to reconstruct exactly what it saw and why it acted. Log a trace for every run: each model call with its full input and output, each tool call with arguments and results, timings, and cost. On our projects, model calls go through AiKey, our AI gateway, which records usage, latency, and cost per agent and per project, and the orchestration layer adds the tool-call side.
6. Cost controls: set limits before the bill does
- Per-task limits on steps and tokens, so a loop can't run forever.
- Per-agent budgets enforced at the gateway, with alerts before the limit.
- Model routing: smaller, cheaper models for simple steps, larger ones only where reasoning needs it.
Budgets and routing are built into AiKey for exactly this reason. Typical numbers are in our AI agent development cost guide.
7. Staged rollout
Don't flip the switch for everyone at once. Move from sandbox (test data) to shadow mode (a person approves each action) to a canary (a small share of real traffic, acting on its own) to full production. Each stage checks something the previous one couldn't. The full approach is in how to evaluate AI agents.
8. A kill switch and a rollback plan
You need a way to stop an agent instantly, without a deploy: a feature flag that routes work back to people. Version your prompts and tool definitions alongside your code, so "roll back to yesterday's agent" is one command.
Where to host it
Run the agent close to the systems it uses. If your data lives in AWS, deploy there; if you're on Vercel or Cloudflare, their serverless and durable-workflow offerings suit many agents. We deploy across AWS, GCP, Azure, Cloudflare, and Vercel depending on where a client already runs, and keep the agent code portable so moving later isn't a rewrite.
Production readiness checklist
- Session state stored outside the server, per user and task
- Task queue, checkpoints, retries, and idempotent actions
- Scoped, per-user tool credentials kept in a secrets manager
- Approval gates on irreversible actions
- Full tracing of model and tool calls, with cost per run
- Step, token, and budget limits with alerts
- Evaluation set that runs on every change
- Staged rollout plan and an instant kill switch
If that list looks like more than you want to build in-house, it's exactly what our AI agent development service covers, from first scope to monitored production.
AI agent development services
We scope, build, and monitor production AI agents for startups, with guardrails and evaluation built in.
Explore AI agent development servicesFrequently asked questions
What infrastructure does an AI agent need in production?
+
A session store, a task queue with checkpoints and retries, a memory store, scoped tool credentials in a secrets manager, tracing and cost monitoring, budget limits, and a way to switch the agent off instantly.
Should AI agents run on serverless platforms?
+
Many can, especially short tasks. Long-running agents need durable execution (queues or workflow engines) so a timeout or restart doesn't lose progress. The best choice is usually wherever your data and systems already live.
How do you control AI agent costs in production?
+
Limit steps and tokens per task, enforce per-agent budgets at an AI gateway, route simple steps to cheaper models, cache repeated work, and alert before limits are reached.
How do you roll back a misbehaving AI agent?
+
Keep a feature flag that routes work back to people instantly, and version prompts and tool definitions with your code so you can redeploy a previous agent version in one step.