Accept 7.4/10 cs.LG trendtoknow-paper-summaries codex-pro/gpt-5.5

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh ยท September 30, 2026 ยท cs.LG, cs.AI, cs.CL

๐Ÿ“Œ Highlights

cua-speedrun reframes CUA benchmarking around performance, speed, and cost under standardized infrastructure rather than treating success rate as the only meaningful outcome.

  • The platform standardizes VMs, execution pipeline, action-observation contracts, timing, and agent interfaces across multiple CUA benchmarks.
  • The paper evaluates 56 agent configurations on OSWorld and 21 on OSWorld2 using representative task sets, while also extending GPT-6 Astra evaluations to MyPCBench and CUA-World.
  • No single model family dominates score, time, and cost: on OSWorld, the high-performance time frontier runs from Claude Opus 5 (low) at 87.6% in 86 s to GPT-6 Astra (low) at 90.8% in 90 s and Astra (xhigh) at 91.6% in 127 s.
  • Reasoning effort has non-monotonic effects: Gemini 3 Flash Preview improves from 33.6% to 57.6% moving low to medium effort while mean task time drops from 492 s to 268 s.
  • Faster I/O can backfire: GPT-6 Astra (low) environment processing falls from 12.68 s to 0.34 s per task under FastCUA, but total task time rises from 89.5 s to 99.0 s.

fig:cua-speedrun-overview

fig:cua-speedrun-overview

Caption: overview. is a standardized platform for benchmarking the performance, speed, and cost of computer-use agents. It combines representative task selection with standardized infrastructure on Modal, enabling consistent comparisons across models, agent harnesses, and reasoning settings. Our findings reveal that more reasoning can reduce both task time and model cost for CUAs, while faster environment I/O can make agents slower.

Why It Matters: It gives the paper's core benchmark story in one visual: standardized CUA speedrunning with representative tasks and frontier analysis.

๐ŸŽฏ Introduction

Computer-use agents operate graphical user interfaces through screenshots, keyboard actions, and mouse actions to complete user-specified goals. The paper starts from a capability milestone: recent frontier models report successes from 78.7% to 86.1% on OSWorld-Verified, above the 72.4% human reference score, and frontier models also perform strongly on long-horizon benchmarks.

The paper's central complaint is that capability progress has outpaced speed and cost evaluation. CUA benchmarks require complex infrastructure, and the community uses different local VMs, cloud desktops, benchmark implementations, API interfaces, hyperparameters, and harnesses. Those differences confound wall-clock comparisons because an observed speed result may reflect runtime setup or harness behavior rather than the agent itself.

cua-speedrun is introduced to isolate these factors. It standardizes evaluation through a serverless cloud-based provider, a common virtual machine setup, a common execution pipeline, portable agent interfaces, and representative benchmark subsets. The stated goal is to make speed and deployment practicality measurable, comparable, and improvable.

๐Ÿ”ฌ Methodology

The benchmark design separates four dimensions that are often entangled in CUA evaluation: the agent, the benchmark, the environment infrastructure, and the agent loop. Each task specifies an initial desktop state, an instruction, and a verifier, either programmatic or LLM-based. Agents observe 1920 ร— 1080 RGB screenshots and act through keyboard and mouse actions. The system measures task time only from instruction delivery to termination or task limit, while provisioning, setup, initialization, and verification are outside the timed interval.

The infrastructure adapts Gym-Anything VM runtime to Modal sandboxes, with separate sandboxes for agents and environments. For self-hosted open-weight models, the paper uses vLLM inference servers on fixed L40S GPUs. The platform records total task time, time spent executing environment operations, time spent executing agent operations, and model-call cost. CUA-AutoDebug validates that keyboard, mouse, and screenshot contracts behave consistently across VM runtimes.

Representative task selection is the paper's main methodological addition for faster benchmarking. A task is embedded by the partial scores and exact completions achieved by M agents; subsets are selected by minimizing energy distance between the selected subset and full benchmark over these task embeddings. The procedure is validated leave-one-agent-out, then the final deployed subset is built using all available agents after choosing a sufficient K.

FastCUA is an experimental fast I/O mode that optimizes networking, action execution, and image processing to reduce action-to-observation latency to 2-28 ms, more than an order-of-magnitude improvement over typical 2-3 s runtimes. The authors evaluate it not as a guaranteed speedup, but as a controlled test of whether lower infrastructure latency actually reduces end-to-end agent wall time.

equation 1

Formula:
\[z_i=(C_{1i},\ldots,C_{Mi},B_{1i},\ldots,B_{Mi})\in\mathbb{R}^{2M}.\]

Meaning: This task embedding combines partial scores and exact-completion indicators across agents, making two tasks similar when agents show similar success and failure patterns on them.

equation 2

Formula:
\[\begin{split} \mathcal{E}(S_K)=\frac{1}{\overline{D}}\bigg( &\frac{2}{KN}\sum_{i\in S_K}\sum_{j=1}^{N}D_{ij} -\frac{1}{K^2}\sum_{i,i'\in S_K}D_{ii'} -\frac{1}{N^2}\sum_{j=1}^{N}\sum_{j'=1}^{N}D_{jj'} \bigg). \end{split}\]

Meaning: The first term compares selected tasks to the full benchmark, the second discourages redundant selected tasks, and the third is constant for the full benchmark.

equation 3

Formula:
\[K^*=\min\left\{K:\min_{q}\min_{k\in\{K-1,K,K+1\}} \rho_q(k)\geq 0.95\right\}.\]

Meaning: This rule chooses the smallest subset size that preserves agent rankings robustly, requiring the 0.95 Spearman threshold to hold at K-1, K, and K+1.

๐Ÿ“Š Experiments

The main experiments evaluate 56 configurations on OSWorld and 21 on OSWorld2 using the representative task sets from the subset-selection method. The configurations include frontier proprietary and open-weight models, public reference agent implementations or native computer-use APIs, selected reasoning-effort settings, and fixed tasks and infrastructure. The paper also reports selected comparisons for harness design, standard versus fast I/O, and single versus batched tool calls.

The reported metrics are mean verifier score with partial credit, mean task time, and mean model cost per task. OSWorld and OSWorld2 are the central frontier benchmarks in the main body, while CUA-World and MyPCBench are used to test long-horizon and personalized computer use with GPT-6 Astra through Codex at four reasoning settings. CUA-AutoDebug is used to validate agent designs and interface correctness on representative tasks.

On OSWorld, the high-performance time frontier includes Claude Opus 5 (low) with 87.6% in 86 s, GPT-6 Astra (low) with 90.8% in 90 s, and Astra (xhigh) with 91.6% in 127 s. Gemini 3.8 Flash (low) matches the highest score and takes almost the same time as GPT-6 Astra, but costs $0.11 rather than $0.71 per task. At the low-cost end, GPT-5.6 Luna (low, direct API) costs about $0.01 per task but scores 61.8%. On OSWorld2, the time frontier is formed by GPT-6 Astra configurations: Astra (high) achieves 75.0% in 914 s, while xhigh reaches 76.9% in 1,237 s for an additional 1.9 points at 35% more time.

The ablations are the most useful part for builders. Gemini 3 Flash Preview on OSWorld improves from 33.6% to 57.6% when moving from low to medium reasoning while task time drops from 492 s to 268 s because model calls fall from 56.0 to 32.5 despite tokens rising from 1,598 to 4,125. For Gemini 3.8 Flash, low and medium both score 91.6% on OSWorld, but medium increases task time from 127 s to 250 s and approximately doubles cost; on OSWorld2, moving low to medium improves performance from 35.3% to 59.9%, while moving medium to high adds only 0.2 points and increases time by 49% and cost by 72%. FastCUA reduces environment processing from 12.68 s to 0.34 s per task, but GPT-6 Astra (low) total task time increases from 89.5 s to 99.0 s.

  • Benchmarks named in the main-body evaluation: OSWorld, OSWorld2, CUA-World, MyPCBench, and CUA-AutoDebug for interface validation.
  • Central metrics: mean verifier score, mean task time, mean model cost per task, Pareto frontiers over performance-time and performance-cost, and a joint frontier over all three.
  • Representative task sets selected in the main body: 50 of 295 OSWorld tasks, 52 of 63 OSWorld2 tasks, 26 of 143 CUA-World tasks, and 38 of 184 MyPCBench tasks.
  • Repeated-evaluation stability example: across five seeds, GPT-6 Astra (low) scores 90.8 ยฑ 1.8% on OSWorld with mean task time of 89.5 ยฑ 2.2 s.

fig:results-overview

fig:results-overview

Caption: Different agents define the speed and cost frontiers. Mean score vs. mean task time and model cost on OSWorld (top, 54 configs) and OSWorld2 (bottom, 19 configs), excluding MiniMax M3. Dashed lines connect the observed Pareto-frontier points, which are circled. Colors and marker shapes identify models, and darker shades indicate increased reasoning effort. The plots compare complete agent configs. Costs use recorded charges or token usage at the applicable API prices. Astra (low) uses five-seed means for each I/O setting. Full configs and plots in Appendix~app:full-results.

Why It Matters: This is the main frontier evidence for choosing agents by score, time, and cost rather than only success.

fig:results-reasoning-interactions

fig:results-reasoning-interactions

Caption: Less reasoning can make an agent slower. Gemini 3 Flash Preview at low, medium, and high effort on OSWorld. The panels show mean task time, generated tokens including reasoning, and model calls across all 50 tasks. Scores are 33.6\%, 57.6\%, and 59.6\%, respectively. Medium and high effort generate more tokens than low effort but require fewer calls and complete tasks faster.

Why It Matters: It explains why lower token output can still be slower if it causes more model calls and bad GUI actions.

fig:results-reasoning

fig:results-reasoning

Caption: The useful reasoning setting changes across benchmarks. Gemini 3.8 Flash at low, medium, and high effort; labels also report mean cost per task. Each curve keeps the model, agent harness, and execution settings fixed. Additional reasoning brings no gain on OSWorld, whereas moving from low to medium substantially improves OSWorld2 performance.

Why It Matters: It shows benchmark-dependent reasoning returns, a key deployment tuning lesson.

fig:results-fast-io

fig:results-fast-io

Caption: Faster I/O can be offset by more waiting and agent interaction. GPT-6 Astra (low), with normal and fast I/O. Panel (a) separates agent time, agent-requested waits, and remaining environment processing, averaged over five seeds per setting; error bars show the standard deviation of mean task time across seeds. Panel (b) shows two consecutive screenshots after opening Save As: the dialog is absent in the first and visible in the second, requested 5.8\,s later with no intervening action.

Why It Matters: It is the paper's clearest systems result: reducing environment latency can expose model-policy mismatch and increase total time.

๐Ÿ”ฎ Conclusion

The paper concludes that CUA benchmarking needs to evaluate capability and efficiency together. cua-speedrun provides standardized infrastructure for measuring performance, speed, and cost, identifies performance-time and performance-cost Pareto frontiers on OSWorld and OSWorld2, and demonstrates that those frontiers shift across benchmarks.

The strongest practical conclusion is that CUA speed is an interaction property, not simply a model-latency property. Reasoning effort, step count, harness design, action batching, generated tokens, environment latency, and UI update timing all influence total task time. For builders, the right agent is therefore the configuration that sits on the relevant workload frontier under the product's score, latency, and cost constraints.

fig:results-joint-frontier

fig:results-joint-frontier

Caption: Choosing an agent by performance, time, and cost. Joint Pareto frontier on OSWorld; the OSWorld2 frontier is in Appendix~app:full-results, Fig.~fig:joint-osworld2. Circled points are the configurations on the Pareto frontier. Time and cost axes are logarithmic. =-1

Why It Matters: It captures the conclusion's operational message: deployment choice should be a multi-objective frontier decision.

๐Ÿ› ๏ธ Future Research Improvements

A natural next research direction is agent training and inference policy design for low-latency desktop interaction. The FastCUA result suggests current agents are not calibrated to screenshots that arrive before UI state updates; future CUAs may need explicit wait policies, UI-state prediction, or training data from faster I/O regimes.

The task-subset selection method should be stress-tested as new model families arrive. Because the representative subset is built from existing agent performance patterns, builders should verify that ranking preservation still holds for agents with qualitatively different interaction styles, specialized desktop training, or much lower latency.

The paper also points toward benchmark design that jointly scores success, time, and cost. CUA-World and MyPCBench are included in the main body as long-horizon and personalized benchmarks, but the detailed main-body frontier analysis centers on OSWorld and OSWorld2; broader frontier reporting across personalized and long-horizon workloads would make the deployment signal stronger.

๐Ÿญ Potential Industry Use Scenarios

The most immediate industry use is regression testing for GUI agents. Teams building desktop, browser, or enterprise workflow CUAs can use cua-speedrun-style infrastructure to compare models and harnesses under stable VM, timing, and verifier conditions before changing production agents.

The benchmark is also useful for cost-performance procurement. A product team can plot candidate configurations on workload-specific score, time, and cost frontiers, then choose a model and reasoning setting that satisfies service-level latency and budget targets rather than overpaying for unused accuracy.

Another plausible use is infrastructure and harness optimization. The paper shows that action batching, harness choice, generated-token volume, and I/O latency can materially change wall-clock behavior, so platform teams can evaluate runtime improvements against actual end-to-end task time instead of microbenchmark latency alone.

๐Ÿ’ฌ Critical Analysis

The paper's major strength is measurement discipline. By making task time start and stop at a clearly defined interval, separating agent and environment operations, fixing infrastructure, and tracking model cost, cua-speedrun attacks the reproducibility gap that makes speed claims hard to interpret across CUA papers and product demos.

The experimental insights are valuable because they are counterintuitive but operationally plausible. More reasoning can reduce total time by reducing bad steps; lower I/O latency can slow agents by giving them stale screenshots; and the same model can shift when wrapped in a different harness. These are exactly the kinds of effects that single success-rate leaderboards hide.

The main caveat is that many detailed configurations, full plots, and some validation analyses are outside the main-body evidence boundary, so this summary cannot verify all supporting details. The main body provides enough evidence for the core claims on OSWorld and OSWorld2, but builders should inspect the released code, main PDF, and full configuration records before making procurement or architecture decisions.

The representative-subset method is promising but should be treated as a benchmark engineering tool, not a universal guarantee. It preserves rankings under the reported leave-one-agent-out criterion and selects compact task sets, including a large OSWorld reduction, but new agents that fail or succeed in different ways could alter the task geometry used by the energy objective.

equation 2

Formula:
\[\begin{split} \mathcal{E}(S_K)=\frac{1}{\overline{D}}\bigg( &\frac{2}{KN}\sum_{i\in S_K}\sum_{j=1}^{N}D_{ij} -\frac{1}{K^2}\sum_{i,i'\in S_K}D_{ii'} -\frac{1}{N^2}\sum_{j=1}^{N}\sum_{j'=1}^{N}D_{jj'} \bigg). \end{split}\]

Meaning: This is the critical mechanism behind benchmark reduction, and its validity depends on whether past agent-performance patterns remain predictive for new agents.

Original Abstract

Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces standardized infrastructure and task sets, with a focus on evaluating the speed and efficiency of CUAs. cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks. Across four different CUA benchmarks, we evaluate how reasoning effort, agent harnesses, and environment latency affect performance, speed, and cost. We find no single model family is optimal for all three; none of the open-weight models are on the frontier, and also, unintuitively, for some models increasing the reasoning effort can speed up task completion, while faster environment input-output can slow down overall task completion time. We also demonstrate that we can effectively reduce the evaluation task set of most CUA benchmarks without degrading overall statistical power, allowing for more efficient benchmarking and comparison. We believe cua-speedrun will enable structured progress towards fast, efficient CUAs, unlocking new real-world use cases and applications. All code, infrastructure, and analysis are available at https://cuaspeedrun.com.