Six Context Garbage Collectors: The Systems Deciding What an Agent Is Allowed to Forget

Six Context Garbage Collectors: The Systems Deciding What an Agent Is Allowed to Forget

A long-horizon agent doing real work accumulates a lot of debris. Browser snapshots from six turns ago. Reasoning steps that led to a dead end. Tool outputs that were useful once and never again. Every token of that debris costs money and degrades attention on the tokens that still matter. The standard response has been chronological pruning (drop the oldest) or end-of-context summarization (collapse everything into prose when you hit the limit). Both approaches treat history as a token buffer rather than as a collection of objects with different future values.

Six research systems from mid-2025 through late September 2026 challenge that frame directly. Each proposes a different answer to a surprisingly hard question: how does a system decide what an agent is allowed to forget? The answers range from guideline optimization run against failure trajectories, to object-indexed lifecycle management with sidecar recovery, to probing the model’s own hidden states in the moment before each action.

One important caveat before the list: four of the six systems covered here (StateComp, Forget Reasoning, PaMER, and DRSR) come from substantially overlapping author groups at the TierFlow Team, Renmin University of China, and Tsinghua University, and all four use the same WorkBuddyBench evaluation. That is not a reason to dismiss them, but it is a reason not to treat their results as four independent confirmations. The remaining two (ACON and Self-GC) come from independent teams and use different benchmarks, giving a useful cross-check.

What follows is a system-by-system account of each approach, with exact numbers from the papers, a comparison table, and a note on what each design choice costs you. If your team is building production AI agents that run for dozens of steps or longer, these are the papers to read before you decide how to handle context.


1. ACON: Compression Guidelines Optimized from Failure

Paper: “Acon: Optimizing Context Compression for Long-horizon LLM Agents” (arXiv:2510.00615, October 2025). Authors: Minki Kang (KAIST), Wei-Ning Chen, Dongge Han, Huseyin A. Inan, Lukas Wutschitz, Yanzhi Chen (Microsoft / University of Cambridge), Robert Sim, Saravan Rajmohan (Microsoft).

ACON approaches context garbage collection as a compression guideline problem. The insight is that a compressor’s failures are informative: when a compressed trajectory fails where the full trajectory succeeds, the failure encodes exactly what information the compression lost. ACON collects paired trajectories of this type and asks a capable LLM to analyze what went wrong. The resulting analysis updates a natural-language compression guideline, which is then used to constrain how the compressor condenses future histories and observations. Because optimization happens in natural language space rather than parameter space, the entire approach is gradient-free and works with any API-accessible model.

The results span three benchmarks, each requiring at least 15 interaction steps: AppWorld, OfficeBench, and Multi-objective QA. ACON reduces peak token usage by 26 to 54 percent while largely maintaining task performance for large models (GPT-4.1). For smaller models (GPT-4.1-mini and Qwen-14B), ACON actually improves performance by removing the distracting long context that those models struggle to reason over: +32 percent on AppWorld, +20 percent on OfficeBench, +46 percent on Multi-objective QA. The optimized compressor can also be distilled into a smaller model, preserving over 95 percent of the teacher’s accuracy while reducing module overhead.

ACON framework: unbounded context growth motivates agent context optimization via guideline learning
ACON motivates task-specific compression by showing unbounded context growth in long-horizon agents. Source: arXiv:2510.00615.

The limitation worth naming: guideline optimization is offline and environment-specific. You need to run trajectories, collect failures, and refine the guideline before deployment. That loop takes time and compute. Teams that need a single compressor to work across heterogeneous environments without environment-specific tuning will find ACON more expensive to operate than its gradient-free framing suggests.


2. Self-GC: Runtime Objects with Recoverable Lifecycles

Paper: “Self-GC: Self-Governing Context for Long-Horizon LLM Agents” (arXiv:2607.00692, July 2026). Authors: Xubin Hao, Hongjin Meng, Xin Yin, Jiawei Zhu, Chenpeng Cao (Xiaohongshu). This system is in production deployment.

Self-GC makes an architectural claim that the others do not: agent history is not a token buffer at all. It is a collection of runtime objects, and managing it correctly requires lifecycle control over those objects, not position-based cleanup. The name deliberately echoes garbage collection in operating systems: the system does not merely reclaim unused tokens; it governs whether each context object is still needed by future computation.

In practice, Self-GC assigns stable identifiers to every user turn and tool span. A side-channel planner (a separate model call that does not pollute the main agent context) reads the indexed objects and proposes one of three actions for each: fold (move the exact payload to a recoverable sidecar and leave a pointer), mask (preserve structural boundaries while eliding low-signal middle content), or prune (remove obsolete content with no recovery guarantee). The harness then rehearses the proposed plan locally, checks that no cut-turn edits violate protocol, and commits the plan only at a safe turn boundary. Critically, the harness also estimates whether committing the plan now would break enough of the provider prefix cache to offset the savings; if it would, the plan stays pending until cache expiry.

Self-GC runtime framework: context-engine hook exposes indexed objects to side-channel planner; harness rehearses and commits at safe boundaries
Self-GC runtime framework: the context-engine hook exposes indexed objects to a side-channel planner; the harness rehearses proposed edits and commits only at safe turn boundaries. Source: arXiv:2607.00692.

On a 33-session Hard Set of long agent traces with heavy browser, shell, and web-fetch pressure, Self-GC reaches 84.85 percent no-impact rate at 43.95 percent pruning. Heuristic baselines achieve only 54.55 to 69.70 percent no-impact despite comparable pruning rates, meaning they delete tokens that future continuations actually needed. On a 332-session Production Suite, three planner backbones (Qwen3.6-Plus, Qwen3.7-Max, GLM-5.1) reach no-impact rates of 91.27 to 94.58 percent, while baselines stay at 77.71 to 87.46 percent. In the production account-level deployment at Xiaohongshu, daytime average input tokens drop 10 to 15 percent, with peak reductions near 20 percent.

The limitation worth naming: Self-GC requires the harness to support object indexing, stable identifiers, sidecar storage, and cache-aware commit logic. That is a meaningful engineering investment. Teams running agents on a provider’s hosted API without harness control cannot use Self-GC without first building the infrastructure layer it assumes.


A note on the next four systems

StateComp, Forget Reasoning, PaMER, and DRSR are all authored by substantially the same group at the TierFlow Team, Renmin University of China, and Tsinghua University, all posted to arXiv within two days of each other in late September 2026, and all evaluated on the same WorkBuddyBench benchmark (260 tasks, plus a fixed 40-task Eval40 subset). The numbers are not directly comparable across papers because different baselines and model backbones are used in each, but the pattern of results is consistent: all four show token savings alongside reward preservation or improvement. The research on agent memory as a learned policy gives useful independent context for reading this cluster of work.


3. StateComp: The Compression Decision as a State Transition

Paper: “StateComp: Learning When to Compress History in State-Conditioned Long Horizon Agents” (arXiv:2609.27298, September 23, 2026). Authors: Mingxuan Wang et al. (TierFlow Team), Yanbiao Ma (Renmin University of China), Jungong Han (Tsinghua University).

StateComp attacks a specific failure mode of threshold-based compression: the same number of tokens can mean very different things depending on what the agent has done. A trajectory that is long because the agent has just assembled critical evidence should not be treated the same as one that is long because the agent repeated the same search six times. StateComp trains a router on the hidden-state representations of a frozen language model to classify past interactions as KEEP or READY. READY does not mean “delete immediately.” It means the interaction has passed the point where it adds state-dependent value and can be safely summarized.

The training data is constructed via a two-stage annotation pipeline over 244,526 interaction-state pairs, identifying 15,633 positive (READY) transitions and 1,832 historical interactions that receive a compression point. The router predicts compression timing from the current agent state; adjacent READY spans are grouped and replaced with summaries. Original content is not kept for retrieval once removed.

On WorkBuddyBench (260 tasks), StateComp reduces total agent and summarization tokens from 698.17 million to 333.24 million, a 52.27 percent reduction, while mean reward moves from 0.6987 to 0.7026. Representation extraction runs 12.67 times faster than the baseline.

The limitation worth naming: because StateComp discards removed content without a retrieval path, information lost at compression time is permanently gone. If the agent reaches a later subtask that needed the original evidence, there is no mechanism to recover it. PaMER (below) addresses exactly this gap.


4. Forget Reasoning: Entropy as the Deletion Signal

Paper: “When Can Agents Forget Their Reasoning? ICLR for Reasoning Block Compression in Long-Horizon LLM Agents” (arXiv:2609.29875, September 24, 2026). Authors: Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen (TierFlow Team), Yanbiao Ma (RUC), Jungong Han (Tsinghua).

Where StateComp targets interaction history at the block level, Forget Reasoning targets a specific content type: reasoning steps. The central observation is that tool calls, actions, and observations carry factual record that future steps depend on, but reasoning blocks are often transitional. Once the agent has acted on a conclusion, the reasoning that produced it may no longer need to be in context.

Forget Reasoning uses a frozen proxy model to estimate the predictive entropy of each reasoning block: low-entropy blocks (the agent’s future behavior is unlikely to change based on what these blocks contain) are candidates for deletion. The method is training-free. When blocks are removed, the compressed history is written back to context so that future requests see the shorter version. Actions, tool calls, and observations are always preserved.

The paper measures a striking phenomenon it calls “trajectory amplification”: a small local deletion in reasoning produces nonlinear changes in total computation across subsequent steps, because downstream reasoning depends on what context was present. This means token savings compound. On WorkBuddyBench (260 tasks), reward improves from 0.699 to 0.718 while input tokens fall 25.5 percent, output tokens fall 14.4 percent, and cache read tokens fall 33.3 percent. A frozen proxy’s hidden state at a context boundary predicts whether a reasoning block will be reused in future steps with AUROC 0.844.

The paper also reports a behavioral risk finding that has implications beyond compression. When the agent’s derived state exists only inside reasoning steps rather than having been externalized to code, files, or tool outputs, future behavioral risk is 67.3 percent. After externalization, it drops to 36.2 percent. This suggests that agents that externalize intermediate conclusions before compressing reasoning are meaningfully safer to compress. Activation patching experiments show that KL divergence repair (the degree to which removed reasoning can be reconstructed from later layers) climbs from 0.047 at layer 4 to 0.978 at layer 31, confirming that the information is not lost from the model’s processing; it gets reconstituted deeper in the network.

The limitation worth naming: the proxy entropy signal can only assess whether a block’s content is uncertain; it cannot directly assess whether later steps will need that specific text as a reference anchor. A reasoning block with low entropy might still contain the only instance of a specific calculation result that a future step needs to quote. The 67.3 percent behavioral risk figure suggests this is not a hypothetical failure.


5. PaMER: Memory Needs Are Already Encoded Before the Action

Paper: “Memory Control Signals Emerge Before Action in Long Horizon Agents” (arXiv:2609.27286, September 23, 2026). Authors: Mingxuan Wang et al. (TierFlow Team), Yanbiao Ma (RUC), Jungong Han (Tsinghua).

PaMER (Preaction Memory with Evidence Retrieval) asks a question the prior systems do not: does the model itself already know whether compression or retrieval is needed, before it is told? The paper studies hidden states at the token immediately before each agent action, across 72,912 decision points from 2,521 long agent trajectories. A linear probe trained on these preaction hidden states predicts compression need with AUROC 0.831 and recall need (whether earlier evidence needs to come back into context) with AUROC 0.765. Both figures substantially outperform observable controls based on context length, turn count, and tool type.

Layerwise analysis reveals asymmetry: compression signals strengthen monotonically with model depth, while recall signals emerge earlier and peak at an intermediate layer. This suggests the two memory operations draw on different representational resources inside the network. The last input token’s representation outperforms mean pooling across the entire context, indicating that the model concentrates memory decision information at the boundary position.

The practical finding is that approximately 26 percent of the full context is sufficient to preserve most of the predictive signal for compression and recall decisions. The remaining history can be offloaded to external storage and retrieved selectively. PaMER implements this by using preaction representations to trigger compression and trigger step-level evidence retrieval from historical context stored outside the active window. PaMER+ selects only the specific raw historical steps that the current task requires, reducing retrieval noise. The design directly addresses StateComp’s limitation: removed content does not disappear; it is externalized and returned when the hidden state signals need for it.

The limitation worth naming: PaMER uses a shared frozen Qwen3.5-9B model for representation extraction. The AUROC numbers measure predictive accuracy of a linear probe, not causal confirmation that the model actually uses these signals to govern its own behavior. The paper is careful to frame the claims as predictive rather than causal.


6. DRSR: Set-Level Deletion Risk with an Abstention Gate

Paper: “DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents” (arXiv:2609.27276, September 23, 2026). Authors: Mingxuan Wang et al. (TierFlow Team), Yanbiao Ma (RUC), Jungong Han (Tsinghua).

DRSR identifies a flaw shared by most compression systems: they score history blocks independently. Two blocks may each appear safe to delete when scored alone, because either one preserves the same fact. Deleting both removes the only remaining copy. This is not a hypothetical edge case; it is a structurally guaranteed failure mode of any method that assigns deletion scores block by block without modeling what remains after the deletion set is applied.

DRSR reformulates compression as risk-constrained selection over deletion sets. The training signal is exact and counterfactual: given a completed trajectory, jointly remove a candidate set of history blocks and measure how the likelihood of the recorded next output changes under the edited context. This converts an online-unavailable quantity (true deletion harm) into an offline training label. A lightweight scorer then learns to predict set-level harm from information available before the next action: the current decision state, the proposed deletion set, the information that would remain, and pairwise interactions within the deletion set.

At deployment, the scorer evaluates candidate deletion sets in a batch without additional language model calls. Pruning proceeds only when the candidate satisfies protocol constraints, recency protection, a token budget limit, and the learned risk gate. If no candidate passes, DRSR abstains rather than forcing a deletion it is not confident about. The abstention mechanism is the mechanism with the clearest practical benefit: other systems delete regardless; DRSR declines when the evidence for safety is insufficient.

DRSR pipeline: offline counterfactual supervision, relational set-risk prediction, and online risk-gated selection with abstention
DRSR pipeline: offline counterfactual supervision labels deletion sets by their measured effect on recorded next output; a relational set-risk predictor gates online deletion with an abstention option. Source: arXiv:2609.27276.

On WorkBuddyBench Full260, DRSR increases mean reward from 0.699 to 0.802 while reducing total model tokens by 20.82 percent. On the fixed Eval40 comparison subset, DRSR achieves 0.794 reward at 1.211 million tokens per task, using 35.85 percent fewer tokens than the uncompressed agent. These are the strongest results of any single system in this group on the shared benchmark, though they are produced by the same team under the same evaluation conditions as the other TierFlow papers.

The paper’s mechanistic analysis also offers an important nuance for teams building compression systems: the agent decision plane changes as the task progresses, and the information from retained context matters as much as the information removed. Both the deleted content and what remains after deletion jointly determine risk, which is why singleton scores systematically underestimate harm.

The limitation worth naming: DRSR’s offline training requires completed trajectories with known next outputs. That means you cannot train a DRSR-style scorer before you have a population of successful trajectories to supervise on, which is harder to obtain for new task domains than it sounds.


Comparison Table

System Deletion unit Decision signal Training required Token reduction Task gain Benchmark
ACON (Oct 2025) Observations and history Failure-driven compression guidelines optimized in language space No gradient updates; guideline optimization loop on failure trajectories 26 to 54% peak tokens +20 to +46% small-model task performance AppWorld, OfficeBench, Multi-objective QA
Self-GC (Jul 2026) Indexed context objects (turns, tool spans) Side-channel planner proposing fold/mask/prune actions; harness enforces recoverability No; planner uses existing model 10 to 20% online input tokens (production deployment) 84.85% no-impact at 43.95% pruning (Hard Set) Production traces (332-session suite), online at Xiaohongshu
StateComp (Sep 2026) Interaction-history blocks Router trained on hidden states; KEEP or READY classification from current agent state Yes; imbalance-aware router trained on 244K state pairs 52.27% total tokens Reward 0.6987 to 0.7026 WorkBuddyBench (260 tasks)
Forget Reasoning (Sep 2026) Reasoning blocks only; actions and observations preserved Frozen proxy entropy; low-entropy reasoning blocks are deletion candidates No; training-free at deployment 25.5% input tokens; 33.3% cache read tokens Reward 0.699 to 0.718 WorkBuddyBench (260 tasks)
PaMER (Sep 2026) History blocks (compress) and external evidence (retrieve) Preaction hidden state of frozen Qwen3.5-9B; AUROC 0.831 compression, 0.765 recall Yes; linear probe trained on labeled decision points Substantially reduces context (26% context preserves most signal) Competitive reward on WorkBuddyBench; adds retrieval recovery WorkBuddyBench (260 tasks)
DRSR (Sep 2026) Deletion sets of history blocks (joint, not singleton) Set-level risk scorer trained on counterfactual likelihood change; abstains if no safe set Yes; offline supervision from completed trajectories with known next outputs 20.82% total tokens (Full260); 35.85% (Eval40) Reward 0.699 to 0.802 (Full260) WorkBuddyBench (260 tasks and fixed Eval40)

What the Design Choices Tell You

The six systems split roughly into two camps. ACON and Self-GC approach the problem from the outside: they define a policy for what to compress without looking at the model’s internal state. ACON optimizes that policy from failures. Self-GC enforces it through object-level lifecycle rules. StateComp, Forget Reasoning, PaMER, and DRSR all look inside the model in some way, either by probing hidden states (PaMER), training on hidden state representations (StateComp), using proxy entropy (Forget Reasoning), or measuring output likelihood change under deletions (DRSR).

The inside-out approaches produce the strongest numbers on token reduction and the most interpretable mechanistic findings. PaMER’s layerwise analysis and Forget Reasoning’s behavioral risk measurements give you something you can reason about when deciding whether to compress reasoning in your own system. But the outside-in approaches have different practical advantages: Self-GC is actually deployed and produces the only online production evidence in the group. ACON works with any API model without model-internal access.

The harness distillation pattern connects to ACON’s distillation finding: if you can train a smaller, cheaper model to replicate what a larger compressor does, you dramatically reduce the per-step overhead of running compression at all. ACON preserves 95 percent of teacher accuracy after distillation; that is a strong enough result that distilling your compressor should be a standard step in any production deployment.

Three practical questions fall out of the literature:

Do you need retrieval after compression? StateComp and Forget Reasoning do not retrieve; what they remove is gone. PaMER adds step-level evidence retrieval. DRSR’s abstention gate takes a different path: when evidence seems risky to remove, the system simply does not remove it. If your tasks involve long gaps between gathering evidence and using it, retrieval or abstention is worth the engineering cost.

Can you afford to look at the model’s internal state? DRSR and PaMER both need hidden-state access or likelihood re-scoring over completed trajectories. That means you need inference infrastructure, not just API access. Teams running agents against a hosted API without model weights cannot use these methods without a local copy of the model used for scoring.

Does your agent externalize its intermediate conclusions? Forget Reasoning’s finding that behavioral risk drops from 67.3 percent to 36.2 percent when derived state is externalized from reasoning to code, files, or tool outputs has a design implication that has nothing to do with which compression system you pick. If your agent is designed to write its conclusions into working memory before it compresses its reasoning, compression is safer regardless of which method you use.


Frequently Asked Questions

What is context garbage collection in LLM agents?
Context garbage collection refers to methods that decide which parts of an agent’s accumulated interaction history can be safely removed from the active context window without harming future task completion. Unlike simple token truncation, these methods reason about future dependencies, task state, and the risk of losing specific evidence before deciding what to delete, summarize, or externalize.

How is DRSR different from other compression systems?
DRSR scores deletion sets rather than individual history blocks. Two blocks that each appear safe to delete individually may together hold the only copy of critical information. DRSR uses counterfactual supervision from completed trajectories to train a lightweight scorer that predicts set-level harm, and includes an abstention gate that prevents deletion when no candidate set meets the safety threshold.

Does Self-GC work with any agent framework?
Self-GC is harness-portable but not framework-agnostic. It requires the harness to expose turn and tool-span boundaries, assign stable object identifiers, support a side-channel planner call, persist folded payloads to sidecar storage, and commit plans only at safe turn boundaries. Teams using a provider’s hosted API without harness control need to build this infrastructure before Self-GC can run.

What does “trajectory amplification” mean in reasoning compression?
Trajectory amplification is the phenomenon, described in the Forget Reasoning paper, where a small local deletion in reasoning produces nonlinear changes in total computation across subsequent steps. Each step’s reasoning depends on what context was present, so removing one block can cascade, compounding both the savings (when the block was genuinely dispensable) and the errors (when it was not).

Why does ACON improve smaller model performance while preserving large model performance?
ACON removes context that smaller models find distracting. Large models can attend effectively over long, redundant histories; smaller models lose track of relevant evidence when it is buried in accumulated tool outputs. ACON’s failure-driven guidelines remove low-value content that smaller models would otherwise be distracted by, producing gains of 32 percent on AppWorld, 20 percent on OfficeBench, and 46 percent on Multi-objective QA.

Should agents externalize intermediate conclusions before compressing their reasoning?
Yes. The Forget Reasoning paper finds that behavioral risk when deleting reasoning blocks is 67.3 percent when derived state exists only inside the reasoning text, but drops to 36.2 percent when the agent has externalized that state to code, files, or tool outputs. Designing agents to write key conclusions into working memory before compressing reasoning is a meaningful safety measure, independent of which compression method is used.