Your AI Agent May Not Need More Information
If you are using an AI assistant to answer customer questions, prepare quotations, check documents or update your business systems, you may assume that giving it more past examples will make it smarter. That assumption can create a familiar problem: the agent receives so much guidance that it becomes slower, less focused and more likely to miss the instruction that matters.
The practical question is not simply whether your AI agent has memory. It is how much memory you should give it for each type of work. Recent research from IBM Research tested this idea across eight AI models using multi-step tasks involving calendars, messaging, payments and other simulated business applications. The findings are useful even if you run a small Malaysian company rather than an AI laboratory.
TL;DR: More memory does not automatically improve an AI agent. Stronger models may benefit from a complete set of guidelines, while other models perform better with a small core plus task-specific retrieval. Start with relevant instructions, measure results, and remove guidance that does not help.
What This Means
In this research, “memory” does not mean replaying every previous conversation. Instead, the system extracts reusable lessons from earlier work. These lessons can include strategies that succeeded, mistakes to avoid and special cases that commonly cause failures.
For example, an order-processing agent might learn guidelines such as:
- Confirm the delivery address before preparing the final order summary.
- Do not mark an order as completed until payment status is verified.
- Ask for clarification when two customer records have similar names.
The agent does not change its underlying model. Its model weights remain the same. The system simply provides better guidance when the agent handles a new task. This makes the approach portable: the same style of memory can be tested with different AI models without retraining them.
The research evaluated 585 multi-step tasks across nine simulated applications in the AppWorld benchmark. The tasks were measured using Task Goal Completion, or TGC, and Scenario Goal Completion, or SGC. TGC checks whether a task was completed, while SGC applies a stricter test: every variation of a scenario must pass. The benchmark contained 168 normal test tasks and 417 challenge tasks, according to the IBM Research article.
| Memory approach | How it works | Best use |
|---|---|---|
| No memory | The agent works with its original instructions only | Simple, predictable tasks |
| Full guideline set | Every learned guideline is included at each step | Strong models with enough context capacity |
| Curated retrieval | A fixed core plus a few relevant guidelines is selected for each task | Focused work where too much context can distract the agent |
The results showed three broad patterns. DeepSeek-V3.2 improved its TGC score by 9.5 percentage points when given the full guideline set. Claude Opus 4.6 improved by 4.1 percentage points, while GPT-5.5 improved by 2.9 percentage points. By contrast, gpt-oss-120b improved by 16.1 percentage points with curated retrieval rather than the full set. GLM-5 showed no measured improvement in this particular evaluation. These figures come from the published IBM Research results.
The right amount of AI memory depends on the model, the task and the quality of the guidance—not on the belief that more context is always better.
How This Applies to Malaysian SMEs
For a Malaysian SME, this matters when you connect an AI agent to daily operations. Imagine a property agency using an assistant to reply to enquiries from WhatsApp, email and website forms. The agent may need to remember tone, viewing arrangements, document requirements and follow-up rules. If you provide every historical rule on every enquiry, the assistant may spend attention on irrelevant information. A better design gives it the common rules plus the property-specific instructions for that enquiry.
The same principle applies to a trading or distribution business. Your sales agent may handle different product categories, customer groups and delivery zones. A request from a Klang Valley retailer should not necessarily receive the same operational guidance as a request involving Sabah or Sarawak delivery. Instead of loading every exception into one giant prompt, you can maintain a small set of core rules and retrieve the relevant product, delivery and approval guidelines when needed.
Service businesses can benefit as well. A clinic, workshop, tuition centre or professional practice may use an agent to manage bookings and follow-ups. The core memory might cover customer privacy, appointment confirmation and escalation to a human staff member. Task-specific memory could then cover the chosen service, practitioner availability or required documents. This keeps the agent focused while still giving it enough operational context.
You should also pay attention to the difference between a successful average result and reliable completion across variations. In a Malaysian business, an agent that usually prepares a correct quotation may still fail when a customer requests a revised delivery date, mixed products or a special approval. The stricter question is: does it complete every important step when the situation changes? Memory is most useful when it addresses those repeated failure points.
Practical Takeaways
1. Separate permanent rules from task-specific guidance
Permanent rules apply to almost every interaction. Examples include your escalation policy, required approval steps and customer communication standards. Task-specific guidance applies only to one product, service, customer type or workflow. Keep these categories separate so that you can retrieve only what is relevant.
2. Build memory from real mistakes
Do not create a long instruction document based only on assumptions. Review failed replies, missed follow-ups, incorrect stock explanations and incomplete form submissions. Turn repeated problems into short, clear guidelines. A useful guideline should explain what the agent must do, when it applies and what error it prevents.
3. Start with a small test group
Choose one workflow, such as lead qualification or invoice checking. Compare the agent with no learned guidance, a full guideline set and a curated selection. Track completed tasks, human corrections and cases escalated to staff. The IBM evaluation compared three configurations rather than assuming one approach would suit every model, as explained in the source article.
4. Watch the context burden
The research reported that full memory increased average tokens per task by 78% for DeepSeek-V3.2 and 51% for gpt-oss-120b. Curated retrieval increased gpt-oss-120b usage by only 5% while producing the strongest improvement for that model, based on the IBM Research measurements. For your business, the operational lesson is simple: measure how much information the agent receives at each step, especially in long workflows.
5. Keep a human escalation path
Memory should help an agent follow your process, not give it unlimited authority. Set clear boundaries for refunds, sensitive customer records, unusual contractual requests and decisions that require management approval. The agent should know when to stop and ask a person.
- List your five most common AI workflow failures.
- Convert repeated corrections into short guidelines.
- Classify each guideline as core or task-specific.
- Test full guidance against curated retrieval.
- Measure completion, correction rate and escalation rate.
- Remove outdated or contradictory instructions.
- Review the memory set monthly or after a major process change.
The Bigger Picture
This trend points towards a more practical way to improve business automation. You may not need to replace your entire AI system whenever it makes a recurring mistake. In many cases, you can improve the surrounding memory and retrieval process instead. That means your operations team can contribute by documenting lessons from real work, while technical tools handle storage and selection.
However, memory is not a substitute for clear processes. If your staff use different approval rules, product names or customer status labels, an AI agent will receive conflicting signals. Before building a large memory library, standardise the workflow the agent is expected to follow. Good memory preserves a clear process; it does not repair a chaotic one.
The research also shows why model selection should be based on your actual workflow rather than general reputation. A stronger model may benefit from broader guidance, while another model may produce better results with a compact, carefully selected set. The useful question for your company is not “Which AI model is smartest?” It is “Which setup completes our important tasks reliably with the least unnecessary information?”
For a Malaysian SME, the safest next step is a controlled pilot. Pick one repeatable process, document the correct outcome, create a small guideline set and test it against real examples with personal information removed. If the agent improves, expand gradually. If it does not, investigate whether the problem is the model, the instructions, the data or the workflow itself.
More AI memory is not automatically better memory. Give your agent the rules it needs, retrieve the exceptions that matter and keep measuring whether the work is actually completed correctly.
Ready to Streamline Your Operations?
Your business should run itself. AutoRunBiz deploys AI agents to automate your daily operations — WhatsApp orders, invoicing, customer follow-ups, and accounting. Book a free 15-min ops audit to see where automation fits your business →
