In our previous article, we compared ALTK-Evolve with ACE, and demonstrated how delivering an agent’s self-distilled guidance—whether retrieving a few principles per task or injecting the full collection—affects both accuracy and cost. This article explores the more fundamental question we left open: How much guidance should we provide?
Equipping an agent with agent memory sounds straightforward: distill experience from past work, place it back into the context, and expect more experience to produce better performance. But that is not always the case. When we expanded our evaluation to 8 models—from a 30B dense model to frontier closed-source systems—one finding stood out:
Agent memory is not a feature you can simply switch on; it is a capability whose dosage must be calibrated to the model.
TL;DR
-
ALTK-Evolve enables agents to learn from their own past trajectories: distilling reusable guidance principles and injecting them at inference time, without updating model weights or requiring human annotations.
-
The right dosage varies by model tier: capable models with room to improve benefit from the full set of guidance principles; weaker models perform best with a compact core set combined with task-specific retrieval; and models that are already near saturation show no measurable gains.
-
Curated retrieval can be either the most accurate or the most cost-effective approach: gpt-oss-120b improved task completion by 16.1 percentage points while increasing token usage by only 5%; at the same time, prompt caching makes the full guidance set affordable in production.
Core Insight: Dosage Depends on Capability
Not all models benefit from the same amount of memory. Across 8 models spanning different capability levels, we observed three recurring patterns:
-
Capable models with room to improve benefit from the full set of guidance principles—including every principle and lessons for rare edge cases. They are capable of absorbing and applying all the information. With the complete self-mined guidance set, DeepSeek-V3.2 (671B MoE) improved its task completion rate by 9.5 percentage points.
-
Smaller or weaker models can be overwhelmed by large guidance sets. For these models, the most effective approach is a compact, high-confidence core set combined with a small number of relevant guidance principles retrieved per task. gpt-oss-120b (117B MoE) improved by 16.1 percentage points with this selective approach; by comparison, the full guidance set delivered a smaller improvement at approximately 50% higher token cost.
-
Models that are already saturated show no measurable gains. We refer to this as the “saturated mode”—a label describing our observation, not a proven explanation. The model may already be near the performance ceiling for these tasks, the guidance may not cover its remaining failure cases, or the model may not be applying the guidance effectively. In our experiments, GLM-5 (745B MoE) fell into this category.
Parameter count alone does not determine which mode a model falls into. Headroom on the benchmark, context-window size, architecture, guidance quality, and task distribution all appear to influence the model’s eventual mode, and disentangling these factors remains ongoing work. Regardless of the specific cause, the practical conclusion holds: the right memory dosage depends on the model, and it can be calibrated.
Learning Happens Around the Model, Not Inside It
The “memory” here is not a replay of past conversations, but a set of guidance principles—effective strategies, mistakes to avoid, and edge cases distilled from the agent’s past trajectories. The overall process is straightforward:
-
The agent attempts to complete a task and generates a trajectory.
-
ALTK-Evolve extracts behavioral guidance from successful and failed runs.
-
The system consolidates these principles into a reusable set.
-
At inference time, the agent receives either the complete guidance set or the subset relevant to the task.
The model weights are not updated. This learning loop changes the guidance available to the agent, not the underlying model itself—which is why it is easy to adopt and portable across the 8 models we tested.
Results Across Different Capability Levels
We evaluated on AppWorld, which contains 585 multi-step tasks (168 test_normal tasks and 417 test_challenge tasks) spanning 9 simulated applications, including calendars, messaging, payments, and more. Tasks are scored in two ways: whether the agent completes each task in full (TGC—Task Goal Completion), and whether all variants of a scenario pass (SGC—Scenario Goal Completion, a stricter all-or-nothing metric). See the appendix for complete definitions.
The Three Configurations We Compared
Because the most confusing aspect of any memory study is often exactly what is included in the context window, we begin by clearly defining the three configurations.
Both memory configurations use the same set of guidance principles. This set was mined once from the AppWorld training set using the process described above. The only difference is how the set is delivered—the full guidance set injects everything at every ReAct step, while curated retrieval delivers a selected subset. The guidance-generation process is unchanged, and no test-set data was used to construct the guidance set.
| Configuration | Contents of the agent context |
|---|---|
| Baseline | No memory—the model’s out-of-the-box state. |
| Full guidance set | Every mined guidance principle is injected at every ReAct step. |
| Curated retrieval | A fixed, high-confidence core set is selected from the same guidance set, and a small number of relevant principles are retrieved per task (fixed portion + variable portion). |
The number of guidance principles a model can mine depends on its own capabilities. We therefore report configurations by strategy—“full guidance set” and “curated retrieval”—rather than by raw count, since raw counts are not comparable across models.
Three Modes at a Glance
Below are representative models from the 8-model evaluation, with task completion rate (TGC) on test_normal:
Figure 1. Representative models from the three observed modes. The bar chart shows TGC for the baseline and best-memory configurations on AppWorld test_normal; the x-axis starts at 40% to make the differences easier to see. Looking at TGC alone understates the larger gains in SGC—see the SGC column in the table below.
The figure plots TGC for readability; the table below adds the stricter SGC metric, where gains are generally larger:
| Model | Mode | Baseline TGC / SGC | Best-memory configuration TGC / SGC | Best configuration | Δ TGC | Δ SGC |
|---|---|---|---|---|---|---|
| gpt-oss-120b (117B MoE) | Weak / selective | 39.9 / 21.4 | 56.0 / 37.5 | Curated retrieval | … | … |
