Accept 7.2/10 cs.CL trendtoknow-paper-summaries codex-pro/gpt-5.5

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Shuai Bai, Jiayong Deng, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang, Yizhen Jiang, Yuheng Jing ยท September 18, 2026 ยท cs.CL, cs.SE

๐Ÿ“Œ Highlights

RecreationWorld studies hybrid computer-use agents that must operate GUIs, write code, build applications, and visually verify their own outputs in one long-horizon loop.

The paper introduces RecreationBench with 250 tasks across Ubuntu, macOS, Windows, Android, and Web, using hidden programmatic and visual assertions derived from running references.

GPT-6 Astra leads the reported benchmark with 58.06% average score, but only reaches 100% programmatic coverage on 2.80% of applications, so complete behavioral reconstruction remains rare.

  • Core claim: recreation is a scalable, verifiable way to train and evaluate hybrid GUI-code agents.
  • Strongest result: GPT-6 Astra scores 58.19% programmatic, 57.92% VLM, and 58.06% average on the aggregate benchmark.
  • Transfer signal: recreation-trained checkpoints improve across ProgramBench, GameCraft-Bench, Vision2Web, OSWorld 2.0, and WeaveBench.
  • Main limitation: agents reproduce static structure more reliably than interactions, computed outputs, and complete task behavior.
  • Builder takeaway: plausible UI reconstruction is not enough; hidden action-conditioned tests are essential.

fig:overview

fig:overview

Caption: Overview of . Left: five-platform recreation tasks. Center: transfer to five out-of-distribution benchmarks, shown individually (pale) and on average (dark) relative to each model's first checkpoint. Right: five-platform programmatic and visual benchmark scores.

Why It Matters: It summarizes the paper's three central claims: five-platform recreation, transfer beyond recreation, and benchmark scoring across programmatic and visual channels.

tab:main-results

Caption: Main results on . Scores and threshold coverage are averaged equally across platforms; the average score is the mean of Prog and VLM. Appendix~app:platform-results reports the per-platform results.

Content:
MetricGemini 3.7Kimi 3GLM-5.3Qwen3.7-Grok 4.6Qwen3.8--0902Claude 5Claude 4.8GPT-5.6GPT-6
Programmatic score (%)24.9132.0726.309.1839.0235.5345.9932.4040.6358.19
VLM score (%)17.3430.7422.469.1234.4534.0742.3429.8143.4957.92
Average score (%)21.1231.4124.389.1536.7334.8044.1631.1042.0658.06
Prog $>90\%$ (% apps)2.402.002.000.001.602.005.531.605.2017.60
Prog $=100\%$ (% apps)0.000.000.000.000.000.000.800.400.402.80

Why It Matters: This is the benchmark's main model comparison and shows that high average scores do not imply full application-level correctness.

๐ŸŽฏ Introduction

The paper starts from a practical gap in computer-use agents: GUI agents can observe and manipulate running applications, while coding agents can write and run software, but real software work requires these abilities to be interleaved. A terminal-only agent cannot see whether the interface it produced matches intent, and a GUI-only agent cannot construct the software behind an interface.

RecreationWorld turns this hybrid requirement into a benchmarkable task. Given a running reference application, the agent must infer behavior through interaction and deliver source code for a faithful candidate implementation. The objective is not source similarity; it is observable behavioral fidelity under hidden reference-grounded tests.

The broader objective is both evaluation and scalable supervision. Because a running reference can act as an oracle, the framework can derive hidden tests and use verifier-backed rollouts to select training trajectories from open-source applications.

fig:task-overview

fig:task-overview

Caption: Application recreation in . Across five platforms, a hybrid agent autonomously interleaves GUI exploration, implementation, and verification to reconstruct a running reference; hidden programmatic and visual assertions score the delivered candidate.

Why It Matters: It clarifies the core task loop builders would need to reproduce: explore the reference, implement, run, verify, and revise.

๐Ÿ”ฌ Methodology

Each recreation task has three parts: an input consisting of a high-level task prompt, access to a running reference application, and a prepared GUI-plus-development environment; an interaction phase where the agent may explore, implement, build, launch, inspect, and revise; and an output consisting of a complete source submission that satisfies a platform-specific delivery contract.

The task is formalized around a reference application and a submitted source artifact. The agent observes the executable reference rather than receiving a complete specification, then returns a source submission whose built candidate is evaluated only through observable behavior. This keeps the benchmark open to different implementation architectures while preserving an objective scoring surface.

RecreationBench construction turns references into hidden test suites. The pipeline prepares a reproducible reference, inventories features and reachable behavior, authors tests with fixtures, actions, and expected observations, validates them against clean reference instances, sends survivors through human review, and freezes the suite for candidate evaluation.

The framework deliberately preserves platform-native semantics. Desktop tasks use source plus build/launch contracts and accessibility APIs; Android delivers Gradle projects that build APKs and uses UiAutomator; Web uses a pinned React web stack producing a self-contained index.html and is evaluated through DOM/ARIA plus visual checks.

fig:task-construction

fig:task-construction

Caption: Reference-grounded task construction, from reference preparation and behavior inventory through test generation, validation, and suite freezing.

Why It Matters: It shows how executable references are converted into validated, frozen benchmark tasks rather than hand-written static specs.

Table 1

Caption: Platform-specific delivery contracts and programmatic test interfaces. Visual assertions supplement these APIs on every platform.

Content:
PlatformDeliveryTest API
UbuntuSource + build/launchAT-SPI
macOSSource + build/launchAXUIElement
WindowsSource + build/launchUI Automation
AndroidGradle project โ†’ APKUiAutomator
WebPinned React web stack โ†’ self-contained index.htmlDOM / ARIA

Why It Matters: It defines what a submitted candidate must deliver and which native programmatic interface evaluates behavior on each platform.

Task input reference

Formula:
\[A^*\]

Meaning: Denotes the running reference application that the agent can operate but must reconstruct behaviorally.

Task output submission

Formula:
\[S\]

Meaning: Denotes the complete source submission returned by the agent under the platform's build and launch contract.

Web repackaging cap

Formula:
\[0.10\]

Meaning: Caps the Web task aggregate when direct repackaging or replay of reference implementation artifacts is detected.

lstlisting 1

Steps: ['Install, launch, screenshot, and inspect the reference and clone renderings.', 'Load reference and clone screenshots as RGB arrays.', 'Compute absolute pixel differences over regions such as app, toolbar, summary, and chart.', 'Report mean absolute error and within-threshold pixel percentages.', 'Use the measured render similarity to decide whether to add remaining behavior such as persistence and settings.']

Why It Matters: This source-backed trajectory excerpt shows an agent using visual comparison as part of the recreation loop rather than relying only on code generation.

lstlisting 2

Steps: ['Use ImageMagick for fast root-window captures.', 'Script menu traversal with keyboard navigation.', 'Capture a sequence of reference screenshots while moving through menu states.', 'Run the same capture procedure over the candidate and reference windows.', 'Compare menu-state screenshots to guide reconstruction.']

Why It Matters: This excerpt demonstrates procedural GUI exploration and comparison, a key behavior for hybrid agents reconstructing desktop application interactions.

๐Ÿ“Š Experiments

The training setup sources tasks from high-quality open-source GUI applications on GitHub across the same five platform families. Qwen3.8-Max generates recreation trajectories using GUI and coding tools; task-specific behavioral verifiers perform rejection sampling; and the final supervised fine-tuning mixture contains 7,000 selected trajectories from each platform, for 35,000 total trajectories.

Out-of-distribution evaluation uses ProgramBench, GameCraft-Bench, Vision2Web, OSWorld 2.0, and WeaveBench. ProgramBench reports test-pass fractions, GameCraft-Bench reports a build-gated rubric score, Vision2Web uses visual scores for Level 1 and mean visual-functional scores for Levels 2 and 3, and OSWorld 2.0 and WeaveBench report task-specific partial scores and shortcut-audited overall scores. The paper reports that both training sweeps finish above their first measured checkpoints on all five benchmarks, though not monotonically.

RecreationBench itself contains 250 tasks, 50 per platform. Evaluation replays frozen hidden suites in clean environments and reports Prog and VLM separately, macro-averaging applications within platforms and weighting platforms equally. The main aggregate table shows GPT-6 Astra leading with 58.19% Prog, 57.92% VLM, and 58.06% average, followed by Claude 5 at 44.16% average and GPT-5.6 at 42.06%.

The benchmark audit supports the claim that the suite tests more than static screens: macro-averaged across platforms, 53.2% of cases require one hop, 24.1% require two or more hops, 94.2% check outcomes, and 40.7% require exact expected results. The failure analysis reports that static structure is easier than button, computation, interaction, and other action-conditioned categories.

fig:transfer-curves

fig:transfer-curves

Caption: Transfer of recreation training to five out-of-distribution benchmarks. Rows are training arms and columns report benchmark-specific metrics on independent vertical scales. Training progress is normalized within each run.

Why It Matters: It is the main evidence that recreation training can improve broader coding and hybrid computer-use performance.

fig:benchmark-composition

fig:benchmark-composition

Caption: Benchmark composition. Top: functional domains for desktop and Android applications and release-native categories for Web. Bottom left: native-application source size and framework distributions. Bottom right: repository-star distributions for non-Web tasks.

Why It Matters: It documents task diversity across domains, frameworks, source sizes, and project popularity.

fig:testcase

fig:testcase

Caption: Generated Windows test case for Logbert. Loading a fixed log and selecting Statistic leads to exact UI Automation and visual assertions. Pale panels show intermediate states, and the inset enlarges the final chart; evaluation uses the full application-window capture.

Why It Matters: It makes the evaluation contract concrete by showing one fixture-driven interaction that produces both structured and visual assertions.

fig:trajectory-behavior

fig:trajectory-behavior

Caption: Reference investigation and final verification. Bars show platform-balanced means with 95\% bootstrap intervals.

Why It Matters: It supports the paper's analysis of how agents explore references, vary inputs, and whether they relaunch and inspect after final edits.

tab:main-results

Caption: Main results on . Scores and threshold coverage are averaged equally across platforms; the average score is the mean of Prog and VLM. Appendix~app:platform-results reports the per-platform results.

Content:
MetricGemini 3.7Kimi 3GLM-5.3Qwen3.7-Grok 4.6Qwen3.8--0902Claude 5Claude 4.8GPT-5.6GPT-6
Programmatic score (%)24.9132.0726.309.1839.0235.5345.9932.4040.6358.19
VLM score (%)17.3430.7422.469.1234.4534.0742.3429.8143.4957.92
Average score (%)21.1231.4124.389.1536.7334.8044.1631.1042.0658.06
Prog $>90\%$ (% apps)2.402.002.000.001.602.005.531.605.2017.60
Prog $=100\%$ (% apps)0.000.000.000.000.000.000.800.400.402.80

Why It Matters: This is the benchmark's main model comparison and shows that high average scores do not imply full application-level correctness.

tab:testcase-audit

Caption: Navigation depth and outcome specificity audit of the frozen suites.

Content:
PlatformStart1 hop2+ hopsOutcome (Exact)
Ubuntu22.557.819.794.6 (41.6)
macOS34.544.021.690.2 (52.6)
Windows26.954.218.992.9 (51.1)
Android16.542.241.393.4 (29.2)
Web13.467.619.0100.0 (28.8)
Macro avg.22.853.224.194.2 (40.7)

Why It Matters: It verifies that the benchmark is not just testing launch screens: most cases leave the start surface and most check action-conditioned outcomes.

๐Ÿ”ฎ Conclusion

The paper's practical conclusion is that hybrid computer-use agents need benchmarks and training tasks that force them to close the loop between interface observation and implementation. Recreation is a strong fit because the reference application supplies an executable oracle and because hidden tests can score independent implementations without requiring source-level imitation.

The reported results are promising but sobering. Recreation training transfers across five out-of-distribution benchmarks, yet the held-out RecreationBench scores show that frontier agents still fail many behaviors that require action-conditioned state changes, exact outputs, or end-to-end functional reconstruction.

For AI builders, the useful message is not simply that one model leads another. It is that agent systems should be instrumented for breadth of reference exploration, controlled input variation, build/relaunch discipline, and final visual plus programmatic verification.

๐Ÿ› ๏ธ Future Research Improvements

The paper points toward improving agents' ability to test input-output behavior rather than merely reproduce static layouts. Same-field controlled input variation appears in only 13.4% to 17.0% of observed native trajectories, suggesting that agents often under-sample the behavior needed to infer rules, persistence, and computations.

Another immediate research direction is stronger final verification. The strict final edit, relaunch or reload, and GUI observation sequence occurs in only 23.6% to 47.5% of trajectories, so many submissions may end after code changes that were never tested in the final state.

Benchmark methodology could also improve around causal analysis. The paper carefully distinguishes observed workflow correlations from causal claims; future work could isolate runtime design, verification prompts, memory/state handling, or GUI exploration strategies in controlled ablations.

  • Train agents to vary inputs systematically when inferring application behavior.
  • Require or reward final build-relaunch-inspect loops before handoff.
  • Study whether programmable interaction runtimes improve quality, not only overhead.
  • Add stronger controls for Android egress and installed-package recoverability where possible.

fig:trajectory-behavior

fig:trajectory-behavior

Caption: Reference investigation and final verification. Bars show platform-balanced means with 95\% bootstrap intervals.

Why It Matters: It supports the paper's analysis of how agents explore references, vary inputs, and whether they relaunch and inspect after final edits.

๐Ÿญ Potential Industry Use Scenarios

The clearest industry scenario is evaluating and training agents that rebuild, migrate, or prototype software from observed behavior. A recreation task resembles real workflows such as replacing a legacy desktop tool, porting an internal app to a new stack, or generating a first functional clone for modernization.

The benchmark design also maps well to enterprise agent QA. Programmatic assertions through accessibility or DOM APIs plus visual checkpoints could be used to verify that generated software preserves action outcomes, not just first-screen appearance.

Another credible use is agent workflow telemetry. The paper's analysis categories, including GUI observation, GUI interaction, reading, writing/editing, build/execution, controlled input variation, and final verification, provide useful operational metrics for teams building long-horizon software agents.

  • Legacy application recreation and migration from observed behavior.
  • Cross-platform prototype generation for desktop, Android, and Web.
  • Automated UI regression and behavioral fidelity testing for generated apps.
  • Training data generation for agents that need both GUI interaction and code execution.

fig:testcase

fig:testcase

Caption: Generated Windows test case for Logbert. Loading a fixed log and selecting Statistic leads to exact UI Automation and visual assertions. Pale panels show intermediate states, and the inset enlarges the final chart; evaluation uses the full application-window capture.

Why It Matters: It makes the evaluation contract concrete by showing one fixture-driven interaction that produces both structured and visual assertions.

Table 1

Caption: Platform-specific delivery contracts and programmatic test interfaces. Visual assertions supplement these APIs on every platform.

Content:
PlatformDeliveryTest API
UbuntuSource + build/launchAT-SPI
macOSSource + build/launchAXUIElement
WindowsSource + build/launchUI Automation
AndroidGradle project โ†’ APKUiAutomator
WebPinned React web stack โ†’ self-contained index.htmlDOM / ARIA

Why It Matters: It defines what a submitted candidate must deliver and which native programmatic interface evaluates behavior on each platform.

๐Ÿ’ฌ Critical Analysis

The strongest part of the paper is its evaluation framing. By using running references as oracles and scoring both programmatic and visual behavior, RecreationBench avoids a common failure mode in agent demos: applications that look convincing but do not implement the underlying interactions or computed outputs.

The main caveat is that the results mix many factors: model identity, harness details, platform differences, tool interfaces, runtime behavior, and training data. The paper is appropriately cautious that trajectory analyses are descriptive and do not prove that a particular behavior, such as broader GUI exploration or final-loop closure, caused higher scores.

Reproducibility depends heavily on the released environments and test suites. The paper says it releases the benchmark, environments, and test suites, which is important because hidden reference-grounded evaluation is only useful to builders if they can inspect the contracts, reproduce scoring, and understand platform-specific isolation boundaries.

As a benchmark, RecreationBench appears demanding and directionally useful, but the field should treat the headline average as a partial fidelity score, not as a pass rate for usable reconstructed applications. The 2.80% full programmatic pass rate for GPT-6 Astra is the clearest reminder that hybrid agents are still early for reliable software reconstruction.

  • Strength: evaluates behavior rather than source imitation.
  • Strength: spans desktop, mobile, and web with native automation interfaces.
  • Weakness: source-blindness is asymmetric because Web necessarily exposes client implementation.
  • Weakness: descriptive workflow analysis cannot isolate causal drivers of model performance.
  • Open question: whether better final verification policies can materially raise full-suite pass rates.

fig:trajectory-source-growth

fig:trajectory-source-growth

Caption: Normalized source growth over trajectory progress. Curves show model--platform medians; triangles mark the first code-writing turn and the diagonal denotes constant-rate growth.

Why It Matters: It shows that agents often create a large scaffold early and then make smaller edits, with notable platform-specific variation.

tab:main-results

Caption: Main results on . Scores and threshold coverage are averaged equally across platforms; the average score is the mean of Prog and VLM. Appendix~app:platform-results reports the per-platform results.

Content:
MetricGemini 3.7Kimi 3GLM-5.3Qwen3.7-Grok 4.6Qwen3.8--0902Claude 5Claude 4.8GPT-5.6GPT-6
Programmatic score (%)24.9132.0726.309.1839.0235.5345.9932.4040.6358.19
VLM score (%)17.3430.7422.469.1234.4534.0742.3429.8143.4957.92
Average score (%)21.1231.4124.389.1536.7334.8044.1631.1042.0658.06
Prog $>90\%$ (% apps)2.402.002.000.001.602.005.531.605.2017.60
Prog $=100\%$ (% apps)0.000.000.000.000.000.000.800.400.402.80

Why It Matters: This is the benchmark's main model comparison and shows that high average scores do not imply full application-level correctness.

Web repackaging cap

Formula:
\[0.10\]

Meaning: Caps the Web task aggregate when direct repackaging or replay of reference implementation artifacts is detected.

Original Abstract

Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomously decide when to explore an interface, implement software, and run and visually verify their artifacts. We introduce RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web, plus a unified harness with native GUI control and coding tools. The running reference serves as an oracle for hidden behavioral tests, providing execution-grounded rewards. We scale trajectory generation with high-quality open-source applications. Models trained on these trajectories improve across five out-of-distribution coding and hybrid computer-use benchmarks and more frequently verify their rendered outputs, providing evidence of transfer beyond recreation. For held-out evaluation, we introduce RecreationBench, comprising 250 diverse tasks across domains and platforms. Reference-grounded programmatic and visual assertions cover action-conditioned outcomes at multiple interaction depths; each is validated on the reference and by human reviewers before the suite is frozen for automatic scoring. GPT-6 Astra leads at 58.1% overall, but passes all programmatic tests on just 2.8% of tasks. Agents reproduce static interface structure more reliably than interactions and computed outputs, while generated applications remain smaller and more monolithic than their references. We release the benchmark, environments, and test suites.