Six Systems Making the Agent Harness a Trainable Artifact

Six Systems Making the Agent Harness a Trainable Artifact

On September 22, 2026, two independent teams published papers on the same day with a shared premise: the executable scaffold surrounding a language model is not a fixed engineering artifact, it is a second learnable object sitting alongside model weights. The Growing Harness paper from Shenzhen Institutes of Advanced Technology reported reducing LLM inference calls by 76.0–91.8% and deployed-agent inference cost by 74.4–98.6% by moving recurring control decisions out of model context and into compiled code. The JAZ framework from MIT CSAIL showed that long-horizon memory and continual self-improvement, workflows that typically require specialized external systems, can emerge from prompting a minimal loop rather than building dedicated harnesses. Neither team was working in isolation. At least four other significant systems, published between June and September 2026, converge on the same structural insight from different directions.

Why This Comparison Is Timely

Three papers landed in four days between September 22 and 25, 2026. Growing Harness (arXiv:2609.26760) and JAZ (arXiv:2609.26891) both appeared on September 22. Pistis Technical Report (arXiv:2609.28554) from ByteDance arrived on September 23, introducing Pistis-Auto-Harnessing as a production-grade harness optimization system bundled with a 27B multimodal model family. These join Harness-R1 (arXiv:2608.02276, August 3), HELIX (arXiv:2608.13951, August 14), and Self-Harness (arXiv:2606.09498, June 8), which introduced the idea of a model improving its own harness without human engineers or stronger external agents.

What makes this cluster worth examining together is not just timing. Each system carves out a different answer to the same question: what exactly can be learned outside the model, and who does the learning? The answers range from compiling recurring decisions into executable code (Growing Harness), to training a dedicated 9B engineer using online RL (Harness-R1), to treating harness evolution as a data pipeline for the next model update (HELIX). Taken together they define a new optimization surface that conventional fine-tuning and prompt engineering both miss.

The Shift: Static Runtime Around a Fixed Model → Two Independently Optimizable Layers

Standard agent deployment keeps two things fixed: the model weights and the surrounding runtime. The runtime, which the harness literature calls the agent harness, includes system prompts, tool schemas, context construction logic, action validation, retry and recovery policies, memory management, and stopping rules. It is hand-engineered once, then shipped. Every deployment decision that the harness does not encode falls back to the model, which must reconstruct it from context on every call.

The systems below reject that asymmetry. They treat the harness as a mutable object with its own optimization loop, while the deployed model remains frozen. The architectural shift is:

model weights (learned) + harness (hand-written, static) → model weights (learned) + harness (separately learned or evolved)

That second layer has very different properties from the first. Harness edits are cheap to execute, do not require gradient computation, can be regression-tested before deployment, and can target failure modes that are invisible to a reward signal computed at the end of a long trajectory. The entries below are ordered from systems already bundled with production models (Pistis-Auto-Harnessing) to the most theoretically minimal (JAZ), with the rationale for that ordering explained in the synthesis section.

Pistis-Auto-Harnessing: Automated Harness Optimization as a Production Component

The Pistis Technical Report (ByteDance, September 23, 2026) introduces a 27B and 9B multimodal model family built on Qwen3.6 and Qwen3.5, together with a post-training framework called Interleaved Distillation and Reinforcement Learning (IDRL). Bundled with that model family is Pistis-Auto-Harnessing (PAH), a system-level method that improves the inference workflow around a frozen Pistis-Agentic model without updating model parameters or changing the interaction budget.

PAH’s mechanism is a closed outer loop. An Optimization Agent reads development trajectories produced by the frozen policy and proposes bounded revisions to prompts, workflow structure, skills, evidence representation, and routing logic. Each candidate is validated: a proposed harness revision is retained only when it improves a pre-specified development metric. The selected harness is then frozen for evaluation. The paper describes this as a “trace-guided propose–validate–update procedure” that applies whenever an executable harness and measurable development feedback are available, making it domain-agnostic despite being instantiated on multimodal search.

On VDR-testmini, PAH raises accuracy from 26.8% to 28.6% under a matched environment-interaction limit. That 1.8-point gain arrives without touching model parameters. The broader Pistis-27B-Agentic results, against Qwen3.6-27B as baseline, show improvements of 8.8 points on BrowseComp-VL, 3.3 on MMSearch, 3.2 on VDR-testmini, and 9.7 on LiveVQA, with some of those gains attributable to the IDRL training and some to PAH. What distinguishes PAH from the other systems is context: it is not a research prototype but a component shipping alongside a production model family, making it the most deployment-proximate approach in this comparison.

HELIX: Model-Harness Co-evolution as a First-Class Principle

HELIX (Tianyu Fan, Chao Huang, University of Hong Kong, August 14, 2026) argues that the standard framing of “make the model better” misses a recursive dependency: the harness governs how the model executes today, and simultaneously structures the trajectories from which the model learns tomorrow. Treating model and harness as independent artifacts severs that feedback loop. HELIX closes it with a three-stage build–update–rebuild cycle.

In the build stage, HELIX constructs and evaluates explicit, source-traceable harness candidates for a fixed model. It decomposes four agent codebases (OpenCode, Pi Mono, Nanobot, Hermes Agent) into typed ports, atoms, recipes, product shells, and runtime policies. The system exposes 96 ports per product contract and enumerates 45=1,024 coupled recipes (or 4,096 when acceptance is independently selected). Pre-execution checks ensure each declared composition is valid before running. The evidence plane retains traces, test results, policy decisions, and provenance, preserving what the paper calls “outcome richness rather than collapsing to a single success bit.” In the update stage, the matched rollouts across harness candidates become “verified sibling trajectories”: paired outcomes capturing successes, regressions, near misses, and alternative solutions that serve as structured training signal. In the rebuild stage, the improved model re-enters harness evolution, since a stronger model may favor a different runtime design.

HELIX’s evaluation runs 65 harness candidates in one evolution round and finds a fixed harness that improves task coverage by 4.0% over Pi Mono. The complete portfolio exposes up to 58.0% more post-hoc coverage. The learning side is equally concrete: a 200-slot sibling slice from the deeper validation yields 438 verified SFT, critic, filter, and preference records for the next model update. HELIX is the only system in this comparison that makes harness evolution feed back into model training as a design goal. The open-source release at https://github.com/HKUDS/HELIX provides the substrate for reproducing this loop.

Growing Harness: Turning Recurring Decisions into Compiled Code

Growing Harness (Laizhen Li, Jiarui Li et al., Shenzhen Institutes of Advanced Technology, September 22, 2026) starts from a deliberate weakness: a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller. No ReAct loop, no Self-Ask structure, no predefined stopping policy. The system then asks whether task feedback alone can turn that empty scaffold into a capable agent, by converting recurring control decisions into persistent executable code rather than re-generating them through model calls on every task.

The mechanism has three parts. First, function-level execution tracing: every failed task produces a directed acyclic graph recording which harness functions executed, which LLM calls were made, which tool calls occurred, and where errors arose. This graph localizes the failure to a bounded set of functions rather than requiring the optimizer to reason about the whole program. Second, joint repair: the optimizer receives a window of K failures together and synthesizes a complete executable candidate that addresses the active failures as a group. Third, a held-out gate: each candidate must improve aggregate success on a separate gate set without reducing any prior capability before its edits are accepted. Accepted repairs accumulate into the shared harness across rounds.

Results across BrowseComp-Plus and WebArena-Verified with three deployment models (4B to 120B parameters): Growing Harness achieves the highest mean success rate in five of six benchmark-model settings, trailing the best mean by 0.7 percentage points in the sixth. Relative to a Tool-Calling agent, it reduces LLM calls by 76.0–91.8% and deployed-agent inference cost by 74.4–98.6%. The scale-sensitivity result is notable: on WebArena-Verified, Growing Harness success stays at 44.7–45.3% as model scale changes from 4B to 120B, while Tool-Calling falls to 6.7% with the 4B model. The learned harness compensates for model weakness by encoding recurring decisions as cheap code rather than expensive inference.

Overview of the Growing Harness training loop: from a strategy-free scaffold, each round traces failed executions, localizes the faulty functions, and repairs function-level code; candidates are accepted only if they fix enough failures and preserve held-out success
Source: Li et al., arXiv:2609.26760, 2026

Harness-R1: A Dedicated Harness Engineer Trained by Online Reinforcement Learning

Harness-R1 (Shuai Shao, Kangning Zhang et al., Shanghai Jiao Tong University and Xiaohongshu Inc., August 3, 2026) takes the most mechanistically distinct approach: it makes harness editing itself a learned capability. All other systems in this comparison use a fixed editing policy, whether that is a prompted LLM, a rule-based proposer, or a search procedure. Harness-R1 instead post-trains a dedicated 9B harness engineer using online reinforcement learning, so that the quality of the edits it produces improves with experience.

The setup separates two models: a frozen target agent that executes tasks, and a trainable harness engineer that reads batches of the target’s failure trajectories and produces executable runtime patches. The patch is not a prompt rewrite or a workflow description. It is a concrete code overlay that wraps the target agent’s execution loop at four lifecycle hooks: episode initialization (sets starting context and episode state), pre-decision (augments context with retrieved guidance and interface constraints before the agent acts), pre-action (a runtime guardrail that may canonicalize, rewrite, or veto the proposed action before it reaches the environment), and post-feedback (inspects returned observations and triggers recovery when the trajectory stalls). After a patch passes validation, the patched target reruns the same tasks from the batch. The realized performance change, not the patch’s text quality, rewards the engineer via group-relative policy optimization (GRPO). Cold-start supervised fine-tuning initializes the editing policy before online training begins.

Across WebShop, ALFWorld, and DBBench, Harness-R1 raises vanilla Qwen3.5-9B success from 44.3% to 53.6%, a gain of 9.3 percentage points. After direct target-agent fine-tuning, a target-specific engineer raises the average from 59.2% to 64.2% (a further 5.0 points), suggesting that the harness engineer and the target can co-evolve in sequence. The paper notes that gains hold both before and after fine-tuning the target, which means the engineer’s learned editing policy transfers across capability levels of the agent it wraps.

Harness-R1 training loop: mined target failures become an evidence bundle; the harness engineer writes executable patches at four lifecycle hooks in the frozen target's runtime; the patched target reruns the same tasks and the resulting reward trains only the engineer via GRPO
Source: Shao et al., arXiv:2608.02276, 2026

Self-Harness: The Agent Improves Its Own Runtime Without External Guidance

Self-Harness (Hangfan Zhang, Shao Zhang et al., Shanghai AI Laboratory, June 8, 2026) addresses a scaling problem: as new models are released rapidly, manually redesigning a model-specific harness for each one is expensive and unreliable. A harness that works well for one model may be suboptimal for another because different models exhibit distinct tool-use habits, error modes, and sensitivities to prompting. Self-Harness proposes that the target model itself should close this gap, without relying on human engineers or a stronger external agent to do the optimization.

The three-stage loop is iterative. In Weakness Mining, the current harness runs the target model on a set of tasks, producing execution traces with verifiable outcomes. The model then clusters failed traces, reasoning about recurring failure patterns rather than isolated mistakes. In Harness Proposal, the same model (now acting as proposer) generates a set of diverse but minimal harness modifications, each targeting a specific failure mechanism. The constraint that edits be minimal is explicit: Self-Harness is not supposed to add generic instructions or inflate the prompt. In Proposal Validation, candidate modifications are tested through regression, and an edit is promoted only if it improves performance without degradation on held-out tasks. Accepted candidates are merged into the next harness version.

Results on Terminal-Bench-2.0 across three model families: MiniMax M2.5 held-out pass rate improves from 40.5% to 61.9% (an absolute gain of 21.4 percentage points), Qwen3.5-35B-A3B from 23.8% to 38.1%, and GLM-5 from 42.9% to 57.1%. Qualitative analysis shows that the changes are model-specific rather than generic. For MiniMax M2.5, the evolved harness encourages the agent to create required output files earlier and stop unproductive tool-use loops. For Qwen3.5-35B-A3B, it focuses on checking dependencies in advance and avoiding repeated failed commands. For GLM-5, it helps the agent preserve environment settings across shell commands. The harness learns the pathologies of each model and encodes targeted corrections, which is exactly what manual harness engineering does, but without human intervention.

Three paradigms of harness improvement contrasted: human harness engineering (left), Meta-Harness using a stronger external agent (center), and Self-Harness where the agent improves its own operating harness (right)
Source: Zhang et al., arXiv:2606.09498, 2026

JAZ: The Harness as a Programmable Language Primitive

JAZ (Zhening Li, Joshua Liu et al., MIT CSAIL, September 22, 2026) approaches the problem from the opposite direction. Rather than optimizing an existing harness, it asks whether a single language-level primitive can eliminate the need for specialized external harnesses altogether. The framework introduces invoke, defined as a function whose implementation is generated by the LLM at call time. Every call to invoke prompts the model to write the function body, execute it, and return the result. This makes the LLM a first-class participant in the language rather than a downstream endpoint.

Two properties distinguish invoke from earlier code-mode agent loops like CodeAct or RLMs. First, recursive sub-invokes: because the language the model writes already contains the invoke primitive, the model can create subagents through ordinary code rather than through a separate orchestration layer. Second, full variable access: everything the model sees, including the user prompt and the REPL history, is also a variable the model can reference and manipulate in its code. The REPL history is not just a log shown to the agent; it is a data structure the agent can read, write, and restructure. This collapses the boundary between the agent’s observational record and its computational state. JAZ implements invoke with a hook system for constraints, monitoring, and budget control, making it extensible without adding more harness structure.

The evaluation targets two capabilities that normally require dedicated external systems. For long-horizon memory (requiring recall beyond the model’s context window), JAZ with GPT-5.4 nano achieves 70% on the recall-heavy subset of StuLife, compared to Letta (MemGPT) at 62% and CodeAct+subagents at 32%, at half the cost of Letta. For continual self-improvement, JAZ on AppWorld (using GPT-5.4 at the top level and GPT-5.4 nano for sub-invokes) achieves 74%, compared to ACE (a specialized self-improvement harness) at 70%, CodeAct+subagents at 71%, and plain CodeAct at 68%. JAZ achieves both by prompting the minimal loop rather than engineering specialized systems. The interesting part is the implication: if a well-designed primitive can subsume these specialized harnesses, the useful work in harness engineering shifts toward defining the primitive, not building the structures on top of it.

How They Compare

System What gets optimized Who does the optimizing Model weights frozen? Regression gate Feeds model training? Key measured gain
Pistis-Auto-Harnessing Prompts, workflow, skills, evidence handling, routing Optimization Agent (LLM) reading dev trajectories Yes Development metric improvement required No (separate from IDRL) VDR-testmini 26.8% → 28.6%
HELIX Typed ports, atoms, recipes, runtime policies across 4 codebases Structured 65-candidate evolution with evidence plane Yes (in execution; model updated in rebuild stage) Pre-execution checks + verified sibling classification Yes, explicitly: sibling trajectories become SFT, critic, and preference records +4.0% task coverage (fixed harness); +58.0% portfolio coverage
Growing Harness Executable control code in the harness (function-level) LLM optimizer scoped to active failure traces Yes Success-first held-out gate with rollback Not reported LLM calls −76.0–91.8%; inference cost −74.4–98.6%
Harness-R1 Executable patches at 4 lifecycle hooks (init, pre-decision, pre-action, post-feedback) Dedicated 9B engineer trained by GRPO on realized task success Yes (target frozen; engineer is trained) Patch validation before installation Not reported Vanilla target 44.3% → 53.6% (+9.3 pp) across 3 benchmarks
Self-Harness System prompts, tools, memory, verification rules, orchestration logic Same fixed model (no stronger external agent) Yes Regression testing on held-in and held-out splits Not reported MiniMax M2.5 held-out pass rate 40.5% → 61.9% (+21.4 pp)
JAZ Entire harness is supplanted by a single invoke primitive The LLM itself at call time (no separate optimizer) N/A (no harness to freeze; harness is generated dynamically at call time) None (no persistent harness state) Not reported StuLife recall: 70% vs Letta 62%; AppWorld: 74% vs ACE 70%

What This Category Reveals

The entry ordering above moves from most production-mature to most theoretically minimal, which happens to match a spectrum of opinion about what the harness is for. At the production end, Pistis-Auto-Harnessing and HELIX treat the harness as a first-class engineering artifact to be systematically evolved. In the middle, Growing Harness, Harness-R1, and Self-Harness treat it as a failure-correction surface to be patched incrementally. At the minimal end, JAZ treats it as a design problem to be dissolved: if the primitive is expressive enough, the harness structures built on top of it become unnecessary.

The critical ingredient separating the systems that generalize from those that risk overfitting is the regression gate. Every system that reports held-out gains, Growing Harness, Self-Harness, and Harness-R1, uses some form of explicit regression control that blocks updates harmful to prior capability. Systems without it, or with a weak one, run the risk of compiling environment-specific tricks rather than transferable control patterns.

The open question the category has not yet answered is transfer across task families. All six systems demonstrate gains within the distribution on which the harness was evolved. None demonstrates that a harness learned on one task domain transfers usefully to a substantially different one. This is not a minor caveat. It determines whether learned harnesses are general infrastructure or per-deployment configuration, which has large consequences for how they fit into an AI engineering stack.

HELIX surfaces a second question: if harness evolution is also a data pipeline for model training, does the choice of harness design space introduce systematic biases into the resulting training data? The same harness interventions that improve execution today shape which trajectories are available for tomorrow’s model update. That feedback loop needs deliberate management.

Limitations and Open Questions

The benchmark coverage is narrow. Growing Harness evaluates on BrowseComp-Plus and WebArena-Verified. Harness-R1 uses WebShop, ALFWorld, and DBBench. Self-Harness uses Terminal-Bench-2.0. These are useful but not representative of the full range of production agent deployments, which typically involve multi-step workflows across proprietary APIs, stateful backend systems, and unpredictable user inputs. Gains on research benchmarks should not be projected uncritically to production performance.

The September 2026 burst (Growing Harness, JAZ, Pistis) arrived in a four-day window, and several papers from this period share overlapping author networks, particularly in the Chinese academic AI community. That does not diminish the individual contributions, but it cautions against interpreting the burst as five independent confirmations of the same idea. It may partly reflect coordinated publishing within a research cluster.

Self-Harness and Growing Harness both acknowledge the overfitting risk directly. A harness evolved on a finite training set can encode rules that appear effective on that set while reducing robustness on distribution shift. The held-out gate in Growing Harness and the regression split in Self-Harness mitigate this risk but do not eliminate it. Harness-R1 addresses it differently by training the editing policy itself with RL, making the engineer generalize across failure types rather than memorizing specific patches, though it notes gains on cross-target settings remain smaller than same-target gains.

None of the systems address concurrent evolution: what happens when a harness is being optimized while the model it wraps is simultaneously being fine-tuned or updated? HELIX names this as the build–update–rebuild cycle, but the empirical results do not yet show multiple full cycles completing on the same codebase. That sequential dependency may be a practical bottleneck.

What This Means for Engineering Teams

If any of these systems stabilize toward production readiness, the practical consequence is a new layer in the AI engineering stack. The current pattern for most teams is: pick a model, write a harness, tune prompts, ship. The pattern these systems propose adds a continuous optimization loop on the harness side, separate from model fine-tuning, running from execution traces rather than from labeled datasets. That requires different instrumentation than current agent development practices provide. Teams need trace collection infrastructure, a regression testing framework that can run on live or near-live data, and a clear definition of which harness components are editable versus which are treated as fixed architecture.

The cost economics are compelling enough to take seriously. A 74–98% reduction in inference cost from harness learning, as Growing Harness reports, dwarfs what is achievable from quantization or KV-cache optimization on a per-token basis. The mechanism is different: rather than making each inference cheaper, the harness removes inference steps entirely by replacing LLM decisions with compiled code. For teams running large-scale agent deployments, the relevant question is not whether the academic benchmarks translate exactly, but whether the mechanism, converting recurring control into cheap code, is real. The mechanism is real. Earlier work on regularized agent harness self-improvement pointed in the same direction; the 2026 systems now show it at scale across multiple benchmarks and model sizes.

For teams evaluating which approach to adopt or experiment with, the comparison table above suggests three useful questions. First, what components of the harness are actually causing failures? Growing Harness and Harness-R1 both work best when failures are traceable to specific control decisions in executable code. If failures arise from model knowledge gaps rather than control structure, harness optimization will not help much. Second, is the failure distribution stable enough to learn from? All six systems assume that the failures observed during training are representative of future failures. If the task distribution shifts frequently, a learned harness may need frequent re-evolution. Third, does the team have the regression testing infrastructure to validate harness edits safely? Without it, any of these systems risk introducing silent regressions on task types not covered by the training set.

Teams building production AI infrastructure should track harness distillation approaches alongside these learning-based systems, since distillation and evolution are complementary: one compresses a strong agent into a cheap harness, the other grows a harness from failure. Both serve the same underlying goal of reducing the marginal cost of each deployment decision. If you are building systems where AI agents must handle complex multi-step workflows at scale, the cost and reliability gains from harness optimization are worth a structured evaluation now, before the category matures further and the best patterns become settled.

Key Takeaways

  • Growing Harness (September 22, 2026) reduces LLM inference calls by 76.0–91.8% and deployed-agent inference cost by 74.4–98.6% by compiling recurring control decisions into executable code rather than regenerating them through model context.
  • Self-Harness achieves absolute held-out pass-rate gains of up to 21.4 percentage points (MiniMax M2.5, Terminal-Bench-2.0) by running the target model as its own harness proposer, without a stronger external agent.
  • Harness-R1 is the first system to make harness editing itself a learned capability via online RL (GRPO), training a dedicated 9B engineer from the realized task success of its patches on a frozen target.
  • HELIX is the only system that explicitly treats harness evolution as a data pipeline: a 200-slot sibling slice yields 438 verified SFT, critic, filter, and preference records for the next model update.
  • JAZ argues the opposite direction: a single expressive invoke primitive can subsume long-horizon memory and self-improvement harnesses through prompting alone, without specialized external systems.
  • Regression gates (held-out splits, validation metrics, pre-execution checks) are the critical safety mechanism separating transferable harness improvements from overfitted environment-specific rules.
  • None of the six systems has demonstrated harness transfer across substantially different task families, which remains the open question with the largest practical consequence for production deployments.

Work With Origins AI

Origins AI builds production AI infrastructure for engineering teams. If your agents are spending significant inference budget reconstructing decisions that could live in executable harness code, talk to our team.