Accept 7.4/10 cs.RO trendtoknow-paper-summaries codex-pro/gpt-5.6-luna

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry · October 01, 2026 · cs.RO, cs.AI, eess.SY

📌 Highlights

RPG turns offline manipulation demonstrations into a simulation practice curriculum, then improves a persistent robot execution system by editing reusable symbolic skills and the system prompt rather than model weights.

The strongest reported result is 209/220 successes, or 95.0%, on held-out initializations of 22 simulated tasks after 15 rounds, compared with 166/220 for ASPIRE and 132/220 for the strongest CaP-Agent0 configuration.

The physical transfer result is 30/30 complete-task successes across Store Ball in Drawer, Fold Towel, and Transfer Bowl after common calibration and hardware adaptation.

The main caveat is that the final simulation system still fails 11 episodes, with seven failures in towel folding, and the physical benchmark covers only three quantitatively evaluated tasks.

  • Weight-frozen improvement through shared skills and prompt revision
  • Privileged execution plus video diagnosis
  • Cross-task regression gates for candidate and merged revisions
  • 95.0% held-out simulation success and 30/30 physical trials
  • Deformable manipulation remains the dominant failure mode

🎯 Introduction

The paper frames embodied-agent development as a tradeoff between Code-as-Policy and Agent-as-Policy. Executable programs can run efficiently and incorporate explicit feedback loops, but robust symbolic perception checks are difficult to specify; multimodal agents provide flexible visual and semantic reasoning, but frequent model calls can make low-level control slow. RPG targets a hybrid execution system that delegates repeatable behavior to reusable code while retaining multimodal reasoning for decisions that benefit from visual or semantic judgment.

The paper-level objective is autonomous improvement without model-weight updates. Given an offline dataset, RPG identifies manipulation capabilities, constructs related simulation tasks, diagnoses failures using execution feedback and privileged state, revises a shared skill library and system prompt, and validates changes across tasks before retaining them. The approach is designed so that a change to a shared skill can benefit multiple tasks that reuse it.

The method is positioned against prior robot-program generation, execution-feedback, skill-library, and simulation-task-construction systems. Its distinctive combination is dataset-guided practice, Runtime-versus-Privileged execution comparison, source-video analysis, and cross-task validation of both individual candidates and merged revisions.

🔬 Methodology

RPG coordinates six specialized agents: Constructor, Runtime Agent, Privileged Agent, Video Analyzer, Implementor, and Merger. The initial library contains nine low-level CaP-X tools and six higher-level manipulation skills for pick, place, pick-and-place, table-supported transfer between arms, move-above, and push. The resulting initial practice library contains 15 entries. The Constructor reviews 3–8 trajectories per selected ABC source task under a 7.5-minute budget, creates a MuJoCo task with the YAM robot model, and writes a task-specific check_success evaluator that humans verify before practice.

The practice suite contains seven CaP-X tasks and fifteen ABC-derived or analogous tasks spanning transport, sorting, stacking, extraction, insertion, articulated closing, bimanual manipulation, and cloth folding. Task specifications and evaluators are frozen before Practice. Examples in the construction table show how organizing sunglasses becomes closing a loaded case, organizing makeup becomes closing a preloaded drawer, and rolling towels becomes a single fold.

At each round, the Runtime Agent attempts every task five times using only deployed-robot observations and execution history. A separate Privileged Agent runs from the same initial states with simulator state added to its inputs. The Video Analyzer receives sampled frames, skill calls, return values, perception results, commanded and measured states, simulator ground truth, evaluator metrics, and available source videos. The Implementor proposes at most one candidate revision per underperforming development task, and candidates must preserve valid syntax, interfaces, and deployment-observable inputs.

For n=5 development trials, the empirical task and suite objectives are defined by equation 1. Candidate revisions are eligible only under equation 2: mean success must increase and no task may lose more than one success out of five. Eligible revisions are merged and retested against the original system before retention. The final Go Real stage applies common calibration and hardware adaptation, then removes the Privileged Agent and Video Analyzer from deployment.

tab:construction

Caption: Representative correspondences between source tasks in the offline dataset and constructed practice tasks.

Content:
Source taskConstructed practice task
Organize sunglassesClose the case; opening and loading are encoded in the initial state.
Organize makeupClose the drawer; preceding manipulation is already satisfied.
Roll towelsFold once; a cloth manipulation analogue with a different final shape.
No corresponding source taskTransfer an object between arms using a task specification.

Why It Matters: The construction examples show how practice tasks isolate capabilities without reproducing source scenes exactly.

equation 1

Formula:
\[\widehat q_i(S)=\frac1n\sum_{k=1}^{n}e_i(\tau_{ik}(S)),\qquad \widehat J(S)=\frac1K\sum_{i=1}^{K}\widehat q_i(S). \label{eq:empirical}\]

Meaning: These estimates define per-task success and the mean suite score used for candidate comparison.

equation 2

Formula:
\[\mathsf{Gate}(\widetilde S;S_r)=\mathbf1\!\left[ \Delta\widehat J>0\ \land\ \min_i\Delta\widehat q_i\geq-0.20 \right]. \label{eq:gate}\]

Meaning: This gate controls whether a candidate or merged system can be retained without a severe regression on any task.

📊 Experiments

The simulated YAM robot has two 6-DoF arms, two grippers, RGB-D observations, and 14-dimensional absolute joint-position commands. The evaluation contains seven CaP-X tasks and fifteen ABC-derived or analogous tasks, for 22 tasks total. Perturbations randomize rigid-body displacement and yaw within ±3 cm and ±0.5 rad, cloth within ±4 cm and ±0.3 rad, table height within ±8 mm, camera translation within ±4 mm, camera rotation within ±0.008 rad, and depth noise and bias. Practice uses five fixed development seeds per task; frozen round-end systems are evaluated retrospectively on ten held-out seeds per task, yielding 220 episodes per point.

For the main comparison, CaP-Agent0 receives up to 50 model turns, while RATs receives up to 50 turns and a 1,200-second wall-clock limit. RPG and ASPIRE receive up to 30 model turns, 12 model-based observation calls, and 1,200 seconds. RPG's self-improvement uses Fable 5.1 for offline revision and Gemini 3.8 Flash for online task-level programs. RPG reaches 209/220 successes, or 95.0%, versus 166/220 for ASPIRE, 132/220 for GPT-6 Astra Pro CaP-Agent0, 107/220 for Gemini CaP-Agent0, and 92/220 for RATs.

The practice trajectory rises from 63/220 successes, or 28.6%, after round 1 to 209/220, or 95.0%, after round 15. The library grows from 15 to 38 entries, comprising 23 new skills and 66 modifications. The largest single jump is from 43.2% to 73.6% between rounds 3 and 4, when the system prompt changes without a skill edit; after the prompt is fixed, success increases by another 21.4 percentage points by round 15. The curve is not strictly monotonic, with one-success decreases at rounds 6 and 14.

The five-round diagnostic ablation gives 78.2% with both Video Analyzer and Privileged Agent, 51.8% without the Video Analyzer, 50.9% without the Privileged Agent, and 50.9% without both. In the library-only comparison, fixing the runtime raises success from 26/40 to 37/40, or from 65.0% to 92.5%. On hardware, RPG completes 10/10 trials for Store Ball in Drawer, Fold Towel, and Transfer Bowl, while CaP-Agent0 with GPT-6 Astra Pro completes 4/10, 9/10, and 1/10 respectively. Five additional physical tasks are qualitative only and are not baseline comparisons.

tab:per-task

Caption: Simulation task completion on ten held-out seeds per task. CaP-Agent0 and ASPIRE are reproduced from their released code. The RATs column reports any-time success. Bold marks the highest observed count in each row.

Content:
TaskCaP-Agent0 Gemini 3.8 FlashCaP-Agent0 GPT-6 Astra ProCaP-Agent0 Opus 5CaP-Agent0 Fable 5.1RATs Gemini 3.8 FlashASPIRE Gemini 3.8 FlashRPG Gemini 3.8 Flash
Lift cube9/1010/1010/1010/1010/1010/1010/10
Nut assembly4/106/104/105/100/103/109/10
Restack cubes6/109/108/107/105/1010/1010/10
Spill wipe2/1010/106/109/105/108/1010/10
Stack cubes8/1010/109/107/1010/1010/1010/10
Two-arm handover5/103/103/106/102/106/1010/10
Two-arm lift2/103/103/102/106/108/1010/10
Group mean (%)51.472.961.465.754.378.698.6
Extract from fixture8/1010/102/102/102/107/1010/10
Insert into fixture2/101/100/100/104/102/1010/10
Place dish on rack1/109/102/105/104/107/1010/10
Serve onto plate9/1010/109/108/102/1010/109/10
Fold towel1/100/100/100/100/100/103/10
Transfer object0/100/100/105/101/109/109/10
Lift rod8/1010/109/1010/1010/1010/1010/10
Pick up cup10/1010/1010/1010/1010/1010/1010/10
Place bottles in bin0/101/100/100/100/103/1010/10
Close sunglasses case9/106/106/106/109/1010/1010/10
Close drawer6/105/106/103/109/1010/1010/10
Sort cubes (standard)3/102/104/105/101/1010/109/10
Sort cubes (extended)0/100/100/100/100/108/1010/10
Stack blocks4/107/100/101/102/105/1010/10
Transport cup10/1010/109/107/100/1010/1010/10
Group mean (%)47.354.038.041.336.074.093.3
Mean success (%)48.660.045.549.141.875.595.0
Total successes107/220132/220100/220108/22092/220166/220209/220

Why It Matters: The full task-level table makes the aggregate advantage auditable and exposes remaining weak cases such as towel folding.

tab:guidance-ablation

Caption: Five-round diagnostic ablations from the same 15-skill initialization. Each variant is evaluated on 220 held-out episodes.

Content:
Video AnalyzerPrivileged AgentSuccess
YesYes78.2%
NoYes51.8%
YesNo50.9%
NoNo50.9%

Why It Matters: The ablation indicates that both diagnostic sources matter and that their information is complementary.

tab:intervention

Caption: Library-only comparison before and after the library revisions, with ten paired initializations per task. The runtime remains fixed, and only the library changes.

Content:
TaskBefore repairAfter repair
Two-arm lift6/1010/10
Place dish on rack5/109/10
Nut assembly5/108/10
Place bottles in bin10/1010/10
Total26/4037/40
Success rate65.0%92.5%

Why It Matters: The controlled comparison isolates the contribution of executable library changes.

fig:architecture

fig:architecture

Caption: RPG uses privileged execution and video analysis to guide skill development. The top row shows source scenes alongside representative simulation practice tasks. The Runtime Agent and Privileged Agent share the same model and skill library, but only the latter receives simulator state $S_t$ in addition to robot observations $O_t$. For development tasks, the Video Analyzer compares the two agents' execution records with available videos from the offline dataset to diagnose failures. The Implementor develops new skills, refines existing skills, and revises the system prompt. Candidate changes are evaluated across tasks. The Merger combines selected revisions, and the resulting system is retest...

Why It Matters: This is the core system diagram: it shows how offline demonstrations, privileged simulation feedback, video analysis, implementation, and cross-task retention form a closed improvement loop.

sec:transfer-case

sec:transfer-case

Caption: Case study of self-improvement on bottle placement. (1) The Runtime Agent attempts the task by calling the shared skill pick\_and\_place\_tall\_object(), but the rollout fails when the bottle collides with the bin wall. (2) Failure diagnosis attributes the error to insufficient placement-path validation and proposes a more robust procedure: validate the full trajectory, release only after the bottle is safely inside the bin, and verify the final placement. (3) The revised skill improves task success from 1/5 to 5/5 on the development trials. The candidate is eligible for merging only if mean task success increases while no individual task drops by more than 20 percentage points. Selected el...

Why It Matters: The case study makes the repair mechanism concrete: execution evidence is converted into a reusable skill change and then protected by a cross-task gate.

fig:improvement-curve

fig:improvement-curve

Caption: Task completion and revision activity. (A) Round-end systems are evaluated retrospectively on the same 22 tasks $ $ 10 held-out seeds after development is frozen. Dashed lines show ASPIRE, RATs, and the matched-model and strongest CaP-Agent0 configurations (abbreviated CaP). Table~tab:per-task gives the full model comparison. (B) Stacked bars count accepted skill additions (green) and modifications (light blue) in each round. Purple squares mark rounds in which the Runtime Agent's system prompt was revised (rounds 1--4); their height does not encode a count. The curve represents one improvement run. Rounds do not represent matched computational budgets.

Why It Matters: The figure connects performance gains to system changes, including the large improvement associated with the round-4 prompt revision and later skill modifications.

fig:real-sequences

fig:real-sequences

Caption: Physical task execution. Representative execution sequences on the physical YAM robot. The first three rows show the quantitatively evaluated tasks: Store Ball in Drawer, Fold Towel, and Transfer Bowl. Bowl transfer uses table-supported staging. The remaining rows show additional qualitative deployments on Insert Cylinder, Serve Fruit onto Plate, and Sort Utensils. Table~tab:real reports outcomes for the three quantitative tasks.

Why It Matters: It visually links the simulation-developed system to the three physical tasks used for quantitative evaluation and additional qualitative deployments.

fig:zero_shot_real

fig:zero_shot_real

Caption: Qualitative zero-shot real-world deployment. Representative execution sequences for five additional physical tasks: Open Scissors, Uncap Marker, Pull out Tissue, Erase Whiteboard, and Unscrew Bottle Cap. The frozen RPG system receives the task instruction and executes each task without task-specific real-world tuning, demonstrations, or retry-based adaptation. These examples are qualitative and are not included in the quantitative baseline comparison.

Why It Matters: These examples probe whether the frozen shared skill system can compose behaviors for new physical instructions without task-specific real-world adaptation.

🔮 Conclusion

The paper's evidence supports RPG as a practical recipe for persistent robot-system improvement: use offline data to create capability-focused practice tasks, diagnose failures with privileged execution and video evidence, revise shared executable skills and prompts, and retain changes only after cross-task validation. The reported system improves from 28.6% to 95.0% on held-out simulation episodes and completes all 30 quantitative physical trials after common adaptation.

The strongest interpretation is not that RPG eliminates the need for stronger models, but that shared skill code and prompt structure can be high-leverage targets for improving a fixed model. The controlled library-only comparison and physical results support that interpretation, while the remaining towel-folding failures and non-monotonic held-out curve show that retention gates do not make improvement risk-free.

🛠️ Future Research Improvements

The paper explicitly identifies rollout efficiency and revision integration as bottlenecks. Each practice round collects Runtime Agent and Privileged Agent trajectories across all 22 tasks, while API rate limits constrain parallel execution. Quota-aware scheduling could prioritize uncertain tasks, recent regressions, or skills with broad dependency impact rather than evaluating every task uniformly at every stage.

Dependency-aware proposal integration is also needed because overlapping revisions to shared skills can conflict. Future evaluations should add repeated independent improvement runs, compute-aware comparisons, broader physical tests, and stronger deformable-object coverage. The paper's seven towel-folding failures and local ABC-sorting regression make these particularly important research targets.

🏭 Potential Industry Use Scenarios

A warehouse or light-manufacturing robot could maintain a shared library of grasp, placement, insertion, sorting, and bimanual-transfer skills. New demonstrations or failure traces could be converted into simulation practice tasks, allowing repairs to be validated across existing workflows before deployment.

Service and household robots could use the same pattern for long-tail manipulation instructions such as opening containers, handling utensils, moving objects between hands, or interacting with drawers and fixtures. The qualitative zero-shot deployments suggest potential for composing skills on new task instructions, but the evidence remains preliminary because those five tasks were not quantitatively compared against baselines.

For production use, the critical operational requirements would be reliable task evaluators, safe simulation-to-real calibration, audit logs for every candidate revision, and conservative cross-task gates that prevent a local improvement from damaging established behaviors.

💬 Critical Analysis

The strongest aspect of the work is the separation of improvement roles. The Runtime Agent remains constrained to deployable observations, while the Privileged Agent and Video Analyzer are used only as development-time diagnostic tools. This makes the information asymmetry useful for debugging without silently giving the deployed policy access to unavailable state.

The experimental design includes meaningful controls: fixed development seeds for candidate checks, held-out seeds for round-end evaluation, a library-only comparison with fixed runtime components, and matched physical calibration for baselines. These controls support the claim that skill-library and prompt revisions contribute materially. However, the paper reports one improvement curve, and the practice rounds are not matched for computational budget, so the exact efficiency and variance of the process remain unclear.

The result is also platform- and evaluator-dependent. RPG relies on human-checked task evaluators, MuJoCo task construction, and a YAM-specific adaptation procedure. The large physical gain over CaP-Agent0 is compelling on the three selected tasks, especially drawer closure and bowl transfer, but broader claims about general-purpose robot autonomy require more robots, objects, environments, repeated runs, and statistically designed physical evaluations.

Original Abstract

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/