Context Engineering: Why Your AI Agent Forgets, Repeats Itself, or Misses the Point
Context engineering decides what an AI agent sees on every step: instructions, tools, retrieved data, memory and history. How to get it right, with examples.
When an AI agent misbehaves, the model usually gets the blame. More often, the problem is what the model was shown. It asks a customer for their order number twice because the first answer scrolled out of the history. It picks the wrong tool because there were thirty to choose from. It quotes last year's pricing because that page ranked higher in retrieval.
The fix in each case is the same: change what goes into the model's context. That work has a name now: context engineering. It's the part of AI agent development we spend the most time on at ThinkDeck.
What is context engineering?
If the model is the agent's brain, context is everything that brain is looking at right now. A model can only reason about what's in front of it.
Prompt engineering vs context engineering
Prompt engineering is about wording: how you phrase instructions and examples. It still matters. But in an agent, the prompt you wrote is a small part of what the model sees. The rest is assembled at run time: tool definitions, search results, memories, tool outputs, and a growing history of steps. Context engineering is designing that assembly.
What competes for space in the context
| Ingredient | What it's for | Common failure |
|---|---|---|
| System instructions | Role, rules, tone, when to stop | Too long; rules contradict each other |
| Tool definitions | What the agent can do | Too many tools, vague descriptions |
| Retrieved knowledge | Facts from your docs and data | Wrong or stale chunks retrieved |
| Long-term memory | What it learned in past sessions | Everything stored, nothing useful recalled |
| Conversation and step history | What's happened in this task | Grows until key details get lost |
| Task state | Plan, progress, open questions | Missing, so the agent loses its place |
Context windows are large now, but bigger isn't free. Long contexts cost more, respond more slowly, and models tend to pay less attention to details buried in the middle. More context often makes answers worse, not better.
Practical context engineering techniques
1. Offer fewer tools per step
If an agent has many capabilities, don't show them all at once. Group tools by stage of the task, or let a routing step choose the relevant set. Five well-described tools beat thirty overlapping ones.
2. Retrieve just in time
Instead of stuffing every possibly relevant document in up front, give the agent a search tool and let it fetch what it needs when it needs it. Attach metadata (product, plan, date) so it can filter out stale content.
3. Compact the history
On long tasks, replace old turns with a short summary of what matters: decisions made, facts confirmed, questions still open. Keep raw tool outputs only while they're needed; a 2,000-line API response rarely needs to stay in context after it's been read.
4. Give the agent a notepad
Let the agent write structured notes (its plan, findings so far, what it's waiting on) to a scratchpad that's always included. This is the simplest fix for agents that lose their place halfway through.
5. Decide what's worth remembering
Long-term memory needs a write policy: store confirmed preferences and decisions, not every passing remark. Then retrieve memories by relevance to the current task, not by recency alone.
6. Put the important things where they'll be seen
Keep stable instructions at the top and the current task and latest observations at the end. Restating the goal near the end of a long context is a cheap, effective habit.
A worked example
Picture a support agent that keeps asking returning customers for details they've already given. The model isn't forgetful; the context simply doesn't contain those details. Three changes fix it:
- At the start of each session, retrieve a short customer profile (plan, open tickets, last issue) instead of the full ticket history.
- After each resolved issue, write a one-line summary to long-term memory.
- Compact the conversation every ten turns, keeping confirmed facts verbatim.
Same model, same prompt wording. The agent now has the facts it needs on every turn, and each request is smaller and cheaper than sending the whole ticket history.
How to tell if you have a context problem
Look at the exact context the model received on a failed step, not the prompt template, but the full assembled input. Ask: was the information it needed actually there? Was it buried? Was there something misleading next to it? Most of the time, the answer is obvious once you look. That's why tracing every model call is non-negotiable; see how to evaluate AI agents.
Watch context size like a cost line
Bloated context shows up on the bill before it shows up in quality reviews. On ThinkDeck projects, every model call goes through AiKey, our AI gateway, so we can see token usage per agent and per step. A step whose input keeps growing run after run is usually a context engineering problem waiting to happen.
Keep going
Context engineering applies to chatbots too; it's most of what makes a RAG chatbot accurate. For the full agent picture, read what an AI agent is, or talk to our AI agent development team about an agent that keeps losing the thread.
AI agent development services
We scope, build, and monitor production AI agents for startups, with guardrails and evaluation built in.
Explore AI agent development servicesFrequently asked questions
What is context engineering in AI?
+
The practice of deciding what information a model sees at each step (instructions, tools, retrieved data, memory, and history) so it has what it needs and little that distracts it.
How is context engineering different from prompt engineering?
+
Prompt engineering focuses on how instructions are worded. Context engineering covers everything assembled into the model's input at run time, including tool definitions, retrieved documents, memories, and task history.
Does a bigger context window remove the need for context engineering?
+
No. Larger contexts cost more, respond more slowly, and models can overlook details buried in long inputs. Selecting and ordering the right information still matters.
What is context compaction?
+
Replacing older parts of a long conversation or task history with a concise summary of the important facts and decisions, so the agent keeps what matters without the context growing indefinitely.