source: Hugging Face Blog: How Much Memory Does Your Agent Actually Need?
level: technical
IBM Research evaluated ALTK-Evolve, a method that distills an agent's past trajectories into reusable guidelines and injects them at inference time without weight updates. They tested eight models on AppWorld, a benchmark of 585 multi-step tasks. The key finding: the optimal amount of memory differs by model capability. Strong models with headroom benefit from the full guideline set, while weaker models perform better with a compact core plus task-relevant retrieval. Saturated models show no measurable gain.
For gpt-oss-120b, curated retrieval improved task goal completion by 16.1 percentage points while adding only 5% more tokens. DeepSeek-V3.2 gained 9.5 points with the full guideline set, but token use rose 78%. GLM-5 showed zero improvement, suggesting it was already near its ceiling. The stricter scenario goal completion metric often showed larger gains, such as DeepSeek's 16.1-point jump. Prompt caching can reduce the cost of injecting full guidelines, since the static portion is identical across steps.
The study distinguishes three patterns: strong with headroom, weak or selective, and saturated. Parameter count alone does not determine the pattern; context window size, benchmark headroom, and guideline quality also matter. The learning loop changes only the guidance available to the agent, not the underlying model, making it cheap and portable. Future work includes a learned selector for retrieval, teacher-distilled memory for very weak models, and controlled experiments isolating context window effects.
why it matters: Choosing the right memory dosage can improve agent accuracy while controlling token costs, which is critical for deploying AI agents in production.
source: Hugging Face Blog: How Much Memory Does Your Agent Actually Need?