RACaP Moves Robot Code Generation Out of the Real-Time Loop

RACaP Moves Robot Code Generation Out of the Real-Time Loop

A September 2026 paper from researchers at the Chinese University of Hong Kong, HKUST (Guangzhou), and Knowin AI introduces RACaP, a robot learning framework that separates the slow, expensive work of writing and testing physical control code from the fast, scene-responsive work of executing tasks. By moving code generation entirely to an offline evolution stage and exposing frozen typed Policy APIs at deployment, RACaP achieves 2.5 times the success rate of the strongest Code-as-Policies baseline on LIBERO-PRO, a 1.9x speedup in median policy execution time, and a 13.2x per-decision speedup after distillation to an 8B-parameter model.

Why Code-as-Policies Has a Runtime Problem

Code-as-Policies (CaP) methods generate executable robot programs directly from language instructions. This is attractive because it does not require training a new neural policy for each new task; a large language model writes the program at the moment the task arrives. CaP-X added visual feedback and repair loops, allowing the system to revise programs after observing physical execution failures.

But putting code generation on the physical execution path creates two problems. First, it is slow. One coding iteration in CaP-X takes 6.8 to 23.8 seconds; a multi-turn trial takes 60.8 to 113.6 seconds. This is incompatible with the responsiveness that physical manipulation tasks require. Second, the generated program often mixes reusable physical mechanisms with task-specific constants: a particular object name, a coordinate value, a grasp height derived from the observed failure. The reusable part cannot be extracted without rewriting the program from scratch for the next task.

More recent self-improving CaP methods attempt to accumulate experience from past runs. RATS collects programs through curiosity-driven play. ASPIRE validates and retrieves program repairs. But both retain source generation or repair during task execution, and play-time objectives do not guarantee that learned programs transfer to downstream tasks. The authors show this directly in a controlled experiment: in LIBERO-90, the play objective in RATS reduces in-domain success from 21.1% to 11.1% and zero-shot success on LIBERO-PRO from 17.8% to 8.3%. An ambiguous play objective can actively hurt task performance.

The authors frame this as the “evolution-to-execution problem”: current CaP methods lack an effective evolution stage, so previous experience cannot be made generalizable for runtime execution to benefit from. Their answer is to separate the two phases entirely.

How RACaP Works

RACaP, introduced by Zexi Li, Yehang Zhang, Haojian Huang, Bohan Zhou, Wenqian Li, Chenxu Wang, Yifan Chang, Yangkai Wei, Tianyi Zhang, Ying-Cong Chen, Kaiwen Zhou, Yinchuan Li, and James Cheng (arXiv:2609.29394, September 24, 2026), defines a deployable policy with three frozen components: a Policy API library (Aθ), a ReAct agent harness (πφ), and experience-based long-term memory (M).

The Policy APIs encode reusable geometry and contact mechanisms. They expose typed arguments that the runtime agent can steer: for example, an argument for object identity or approach direction, but not the low-level grasping trajectory itself. This separation is intentional: “Code should store physical mechanisms that improve success across tasks. ReAct should retain scene-dependent choices, including action order, function arguments, recovery, and stopping.” The interface between code and runtime decisions must be strong enough to execute reliable control and flexible enough to expose the choices that change scene to scene.

The ReAct harness maintains two memory systems: task-objective-driven working memory for the current episode and long-term experience memory from which it retrieves general lessons about physical manipulation. Using visual observations, perceptual measurements, and tool reports, the ReAct agent can change its action sequence, API arguments, recovery strategy, and stopping decision without generating any new source code.

Evolution happens before deployment in two phases. Phase 1 uses capability curriculum learning to build basic competence, establishing reliable execution, recovery behaviors, skill breadth, and orchestration. The curriculum gradually increases task difficulty so the system acquires foundational capabilities before facing the full task distribution. Phase 2 starts from Phase 1’s foundation and runs autonomous self-evolution: the system selects failures, proposes API changes, tests candidate changes, and promotes a candidate only when its paired performance gains on a development set exceed its paired regressions. This conservative acceptance criterion prevents the system from replacing a working mechanism with one that helps on some tasks but hurts on others.

RACaP architecture diagram: evolution stage on the left with curriculum learning and visual critic refining Policy APIs, deployment stage on the right showing ReAct agent calling frozen typed Policy APIs with visual observations
Source: Zexi Li, James Cheng et al., arXiv:2609.29394, 2026

Results: What They Measured

The authors evaluated RACaP across four settings: in-domain learning on LIBERO-90, zero-shot transfer to LIBERO-PRO, long-horizon tasks on LIBERO-Long, and cross-embodiment evolution.

On LIBERO-90 in-domain: RACaP achieves 54.4% success. On zero-shot transfer to LIBERO-PRO (a harder benchmark with novel scenes the system was not evolved on): 45.0%, compared with at most 17.8% for CaP baselines. On LIBERO-Long, which requires managing sequences of causally dependent subtasks across a long horizon: 46.0%, compared with at most 4.0% for CaP baselines. The LIBERO-Long gap is the most striking: CaP baselines essentially fail on long-horizon tasks while RACaP reaches nearly half the task set.

On LIBERO-PRO specifically, RACaP achieves 2.5 times the success rate of the strongest CaP baseline with a 1.9 times speedup in median policy time. The speedup comes directly from removing runtime code generation from the execution path: the ReAct agent calls typed functions with typed arguments instead of generating programs.

For edge deployment, the authors distilled the ReAct decision traces from GPT-5.6 into Qwen3-VL-8B-Instruct using rejection-sampled fine-tuning. The distilled 8B model achieves a 13.2 times per-decision speedup compared to GPT-5.6 and reduces repeated physical calls from 16 to 4 (a 75% reduction). This is the path toward on-robot deployment where inference must run at hardware speeds: train a large capable model to discover good decisions, distill the decision-making function into a small model that fits on the robot’s compute, and leave the physical mechanisms in the frozen code layer.

Limitations and Open Questions

The evaluations are on LIBERO benchmarks, which are simulation-based manipulation tasks. The paper reports results on a structured benchmark rather than on a physical robot in an uncontrolled environment. Physical deployment would add perception uncertainty, calibration issues, and physical variability that simulation does not fully capture. The authors acknowledge this and describe LIBERO as the evaluation setting.

The fixed Policy API library can become a bottleneck when tasks require genuinely novel physical manipulation that existing APIs cannot express. The two-phase evolution system can add new APIs during evolution, but it cannot discover capabilities that the robot hardware cannot execute or that were not encountered during the evolution phase. Zero-shot transfer works when the target tasks can be composed from the evolved API set; it may fail when they cannot.

The cross-embodiment evolution result is reported but the abstract does not give numbers. This would be an important result to understand in detail: whether APIs evolved on one robot hardware configuration transfer to another is a significant question for practical deployment where hardware configurations vary.

The distillation result (13.2x speedup) assumes that the teacher model (GPT-5.6) makes acceptable decisions on the training set and that rejection sampling produces a training set that generalizes. The reduction in repeated physical calls from 16 to 4 is valuable but still represents multiple physical contact events per task episode. Whether this is acceptable depends on the task and hardware.

What This Means for Engineering Teams

The evolution-to-execution separation in RACaP reflects a design principle with broad applicability beyond robotics: expensive, slow reasoning belongs in an offline phase; fast, cheap function calls belong in the runtime phase. Teams building AI agents for any domain where latency and reliability matter should examine where they currently do slow reasoning on the critical path and whether that reasoning could be moved offline.

The typed Policy API interface is essentially a contract. Code on one side of the contract evolves to become more capable and general; runtime decisions on the other side become cheaper because they only need to choose which function to call and what arguments to pass. This is structurally similar to how well-designed software systems separate stable interfaces from volatile implementations. The fact that it produces a 13.2x inference speedup and removes the code-generation latency from physical execution is a strong validation of the principle.

The distillation pathway, from a large capable teacher that discovers good decisions to a compact on-device model that executes them, is increasingly the practical path for deploying AI in constrained environments. The 8B-parameter model running on robot hardware is not discovering new physical manipulation strategies; it is executing the decision function learned from the larger model. The intelligence is in the evolved code layer and the distilled decision model combined, not in either alone.

Teams working on production AI agent systems should also pay attention to the conservative evolution acceptance criterion: promote a candidate only when gains exceed regressions. This is a simple principle that prevents unstable evolution and could be applied to any system that iteratively improves a capability library, a tool set, or a prompt collection. The risk of runtime failure from an untested capability change is significantly higher than the risk from a change that passed paired regression evaluation.

Key Takeaways

  • RACaP separates code generation (offline evolution) from task execution (runtime ReAct), eliminating the 6.8-to-23.8-second coding iteration that CaP-X places on the physical execution path.
  • On zero-shot transfer to LIBERO-PRO, RACaP reaches 45.0% compared to 17.8% for the best CaP baseline; on LIBERO-Long, 46.0% versus at most 4.0% for CaP baselines.
  • LIBERO-PRO results: 2.5x the success rate of the strongest CaP baseline with a 1.9x speedup in median policy time.
  • Rejection-sampled fine-tuning of GPT-5.6 decision traces into Qwen3-VL-8B-Instruct produces a 13.2x per-decision speedup and reduces repeated physical calls from 16 to 4.
  • Two-phase evolution, capability curriculum then conservative autonomous self-evolution, prevents the play-objective problem that reduced RATS success by 10 percentage points in controlled comparison.
  • The typed Policy API interface cleanly separates reusable physical mechanisms (in code) from scene-dependent decisions (in the ReAct agent), enabling transfer without rewriting programs.

Work With Origins AI

Origins AI builds production AI systems for engineering teams. If your team is deploying AI agents in latency-sensitive environments and needs help structuring offline capability development versus runtime decision-making, talk to our team.