Accept 7.5/10 cs.CL trendtoknow-paper-summaries codex-pro/gpt-5.6-luna

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma ยท October 01, 2026 ยท cs.CL, cs.AI, cs.DB

๐Ÿ“Œ Highlights

Argo-Bench is a consequential evaluation framework for data agents. Its core contribution is to place ordinary-looking stakeholder requests inside a large, coherent ERP warehouse, hide the simulator's ground-truth state, and score what the agent files or changes rather than merely checking whether a query matches an answer key.

  • 210 tasks span trust and safety, FP&A, marketplace, accounting, and growth workflows.
  • The environment contains 235 tables and 7.49 billion rows representing a simulated 2024 New York City food-delivery platform.
  • Claude Opus 5.5 is the strongest of 14 evaluated models, but solves only 34.8% of tasks and averages 59.5 points.
  • The dominant failure modes are reading the wrong record, optimizing the wrong objective, measuring the wrong quantity, and filing overconfident forecasts.
  • The benchmark is realistic enough to expose enterprise reasoning gaps, but its single-city, single-year, simulator-defined world limits direct claims about production performance.

๐ŸŽฏ Introduction

The paper targets data-agent workflows that have moved beyond simple text-to-SQL questions toward long-horizon analysis, durable dashboards, fraud enforcement, promotion decisions, and operational planning. It argues that established benchmarks can make enterprise data work appear nearly solved while evaluating a narrower capability: generating SQL over fragmented, relatively documented public schemas. Spider 2.0-Snow's best reported score had risen to 96.7%, and the top BIRD entry reached 82.4% against a human 93.0%, but these figures do not test whether an agent can navigate mutually constraining enterprise modules or act on its findings.

The authors identify three gaps. Existing datasets are patchworks of public or local databases rather than one coherent enterprise system; many valuable tasks require Python-based statistics, machine learning, or optimization and cannot be completed in SQL; and answer-key grading cannot reliably evaluate decisions whose outcomes are not directly observed. Argo-Bench addresses these gaps with a simulated but calibrated food-delivery business, a hidden latent state, an ERP-scale warehouse, and a consequence-based grader.

๐Ÿ”ฌ Methodology

The simulated world models a dominant New York City food-delivery platform in 2024. The authors draw on 34 public datasets and reports, using identity donors for public entities, shape donors for distributions, and published anchors for totals that the simulator must reproduce without copying records. The world contains 81 million orders and about 3.4 million active customers, models courier, customer, and merchant incentives, and incorporates fraud patterns such as order theft, GPS spoofing, account renting, promotion abuse, account takeover, and payout diversion.

The simulator exports its state into an Oracle E-Business Suite 12.2-style warehouse. The schema was designed with three experienced ERP consultants, uses standard EBS table and column conventions, and intentionally omits latent state such as future outcomes and fraud labels. The warehouse is coherent and free of legacy data drift by design, which isolates the ability to understand and explore enterprise data organization. Agents operate in a sandboxed Python environment, query the warehouse, and file explicit actions through mission control.

Tasks are organized into five business areas: trust and safety, FP&A, marketplace, accounting, and growth. Of 210 tasks, 178 file more than a figure, and 99 expose only a partial 2024 snapshot so that forecasts can be judged against unseen months. The reference workflow demonstrates the intended pattern: join a small subset of relevant tables, recover the minimum-pay rule, backtest the forecast, and file an interval. The central business rule is represented below.

Minimum-pay true-up rule

Formula:
\[\max(\text{individual floor},\; \$19.56 \times \text{hours} - \text{pay}) - \$0.054\text{ per hour}\]

Meaning: This is the business rule recovered by the reference forecasting workflow in the main-body example. It controls how the agent constructs minimum-pay adjustments before fitting and filing the May forecast interval.

๐Ÿ“Š Experiments

The evaluation uses 210 Argo-Bench tasks over the private-seed world, with the agent seeing only the warehouse snapshot specified by each task. Agents run in isolated gVisor sandboxes with 25 preinstalled Python libraries, no outbound internet access, and a mission-control interface for filings. Runs end after 500 model turns. The task mix includes action-taking, forecasting, dashboard publication, budget allocation, and reporting, with forecasts evaluated on months after the agent's cutoff. The task distribution and partial-view design are summarized in the attached figure.

Argo-Bench is positioned against BIRD, Spider 2.0, BEAVER, DSBench, DABstep, ฯ„-bench, and CRMArena-Pro. Unlike the comparison benchmarks, it combines an ERP-scale coherent system, Python and ML support, action grading, and latent-state ground truth. The paper evaluates 14 proprietary and open-weight models at their highest tested reasoning effort. Claude Opus 5.5 leads overall at 34.8% solved and 59.5 mean score; GPT-6 Astra leads forecasting, while Claude Sonnet 5.5 leads compliance.

The reported findings show that more reasoning effort helps some models but has diminishing returns. Claude Opus 5.5 gains 17.0 points from low to medium effort, 5.9 from medium to high, and 4.0 from high to extra-high; Claude Sonnet 5.5 gains 14.9 and 11.5 over its last two steps. Gemini 3.8 Flash and Muse Spark 1.3 average 197 and 250 model calls per task without scoring higher than Opus, and 57% of Muse's warehouse spend goes to tasks on which it scores below 5. The paper's qualitative findings are equally important: agents often use a sound method on the wrong record, optimize the wrong objective, or measure a quantity that the warehouse does not actually store.

The forecast results expose a calibration problem. Across 4,553 forecast series from 3,346 runs on 72 tasks, nominal 80% intervals contain the realized value only 44.8% of the time. In an incentive-allocation task, Claude Opus 5.5 saves $3.09 million of an attainable $3.12 million when it recognizes that quests substitute for surge pay, whereas GPT-6 Astra's alternative objective loses $86,281 and scores 0. These results suggest that objective specification and evidence validation can matter as much as raw reasoning depth.

Table 1

Caption: Argo-Bench pairs an ERP-scale warehouse with a Python sandbox for machine learning and optimization, and grades an agentโ€™s actions against a latent state. Tables and rows are per database. Entries marked n/r are not reported by the benchmark. A coherent system is a single enterprise application whose tables must reconcile. Actions are graded when the benchmark scores what the agent changes or files rather than an answer it returns.

Content:
BenchmarkTables per DBRows per DBDataCoherent systemPython and MLActions gradedGround truth
BIRD (Li et al., 2023)7.3549KPublicโœ—โœ—โœ—Gold SQL
Spider 2.0 (Lei et al., 2025)52.6โ€ n/rPublicโœ—โœ—โœ—Gold SQL
BEAVER (Chen et al., 2024)101.5n/rPrivateโœ“โœ—โœ—Logged SQL
DSBench (Jing et al., 2025)n/rn/rPublicโœ—โœ“โœ—Answer keys
DABstep (Egg et al., 2025)n/r138KRealโœ—โœ“โœ—Answer keys
ฯ„-bench (Yao et al., 2025)32.8KLLM-madeโœ—โœ—โœ“Goal state
CRMArena-Pro (Huang et al., 2026)2555KLLM-madeโœ“โœ—โœ—Generator
Argo-Bench2357.49BSimulatedโœ“โœ“โœ“Latent state

โ€ Spider 2.0-Lite, as computed by Chen et al. (2024).

Why It Matters: The comparison isolates Argo-Bench's combined design: enterprise-scale relational structure, Python and ML support, action grading, and latent-state ground truth.

Table 2

Caption: Results on Argo-Bench (210 tasks), each model at the highest reasoning effort we ran, extra-high where offered. Solved is the share of tasks scoring at least 95, and score is the mean task score out of 100, with a modelโ€™s forecast tasks floored at 0 as a block. Domain columns give mean scores. Steps (model calls) and cost (API spend, excluding BigQuery) are per-task averages. Best results are in bold.

Content:
ModelSolved (%)ScoreFcst.FraudFin.Dash.Comp.StepsCost ($)
GPT-6 Astra27.651.839.548.085.871.085.7232.71
GPT-6.1 Sol24.849.538.447.687.163.482.8240.49
GPT-6 Sol17.636.827.935.475.837.375.7380.85
GPT-6 Luna7.615.02.119.156.110.142.9390.06
Claude Opus 5.534.859.535.366.190.084.681.6824.71
Claude Sonnet 5.528.651.834.662.575.654.186.9773.74
Claude Sonnet 59.517.32.523.757.021.523.8552.08
Claude Haiku 4.51.45.50.49.319.20.035.7310.20
Gemini 3.8 Flash16.726.23.833.461.336.585.71973.94
Muse Spark 1.313.820.73.824.046.826.385.72504.89
Kimi K314.328.415.833.673.318.881.0854.52
GLM 5.3 Flash12.421.23.229.857.523.061.9920.35
DeepSeek V4.1 Flash17.625.44.630.856.941.775.01250.42
Qwen 3.8 Max17.124.46.124.566.135.565.5772.72

Why It Matters: The table shows that even the leading model is far from reliable across enterprise workflows, with large domain differences and substantial variation in reasoning steps and cost.

fig:overview

fig:overview

Caption: grades the filings of an agent with an incomplete view of the world. Top: one task from each of the five business areas (as in Figure~fig:tasks). Bottom: the forecasting task fc-12 (Appendix~app:examples). The agent sees the warehouse only up to April 30, and only the grader sees May's actual value. The reference solution reads 6 of the 235 tables, recovers the true-up rule, files a forecast for May with an 80\% interval, and grades 94 out of 100, while repeating April's \$1.44 million misses by \$1.41 million and grades $-1$.

Why It Matters: It makes the benchmark's central distinction concrete: agents must reconstruct hidden business state from an incomplete warehouse and then file an action whose future consequences are graded.

fig:calibration

fig:calibration

Caption: The simulated world matches the city's reported economics, and its warehouse follows ERP conventions. (a) Deviation of per-delivery economics from the 2024 NYC DCWP quarterly anchors. Courier pay runs 6--9\% above the anchor after the April minimum-pay increase, and all other figures stay within 5\%. (b) Shares of the warehouse's 235 tables and 7.49 billion rows by EBS module group. Custom extensions alone hold about a third of both the tables and the rows.

Why It Matters: It shows how the authors connect public economic anchors to an ERP-shaped warehouse, supporting the benchmark's realism claim while exposing its deliberate simulation boundary.

fig:tasks

fig:tasks

Caption: Tasks in each business area vary by expected final action and the snapshot of data given. (a) Tasks per business area, grouped by action taken. Multi-part tasks are counted by their first filing. (b) The last month of 2024 visible to the agent in the 99 tasks with a partial view. Partial views are used mostly in forecasting tasks, which are graded against the months after the cutoff.

Why It Matters: It clarifies the benchmark distribution: tasks differ both in the business area and in whether the agent must act on accounts, forecast, publish a data source, allocate a budget, or report figures.

๐Ÿ”ฎ Conclusion

Argo-Bench demonstrates that enterprise data-agent competence requires more than generating executable SQL. The agent must discover how records relate across ERP modules, infer latent business rules from observable evidence, use statistical or optimization tooling when needed, and file an action whose consequences align with the stakeholder's objective.

The main empirical conclusion is sobering: the best evaluated model solves only 34.8% of tasks and averages 59.5 points, while many failures arise from careful but semantically misaligned analysis. The benchmark's hidden-state design and executable reference solutions make these failures more diagnosable than answer-key disagreement, although the simulator remains an abstraction of production enterprise data.

๐Ÿ› ๏ธ Future Research Improvements

The paper directly motivates broader environment coverage: add foreign-currency representations, shared accounts across multiple businesses, longer histories spanning major market shocks, and additional ERP formats such as SAP S/4HANA. These extensions would test whether agents can transfer schema and business reasoning beyond one city, one year, and one Oracle-style system.

A stronger next generation should also vary simulator assumptions, evaluate realistic tail behavior, include controlled warehouse messiness, and repeat each model-and-effort setting across seeds. Prompt variants should remain explicit because the reported results show that hints about the objective or holdout can materially change performance. Forecast evaluations should emphasize empirical coverage and calibration alongside point accuracy and skill scores.

๐Ÿญ Potential Industry Use Scenarios

The task design maps naturally to enterprise copilots for fraud and abuse operations, FP&A forecasting, courier or workforce incentive planning, accounting reconciliation and back-pay reporting, and growth analytics. In each case, a production system could use the benchmark's pattern of combining warehouse exploration, Python analysis, explicit filing, and consequence-based review.

The results also indicate where human approval remains important. Fraud bans can trade prevented loss against wrongful customer loss, incentive plans depend on the exact operational objective, dashboards can be structurally valid but numerically wrong, and forecasts can be sharply overconfident. A practical deployment should therefore require provenance, objective declarations, counterfactual or holdout checks, calibrated uncertainty, and approval gates before durable actions.

๐Ÿ’ฌ Critical Analysis

The strongest aspect of Argo-Bench is alignment between task difficulty and grading. The warehouse is large enough to require navigation, yet the latent simulator state provides an objective reference for outcomes that a real warehouse cannot reveal. Explicit filings also make partial credit and operational intent clearer than comparing intermediate SQL strings. The reference workflow shows that solvability is not merely asserted: the agent can recover a business rule from a small subset of the warehouse and produce a graded forecast.

The central caveat is ecological validity. A greenfield, internally consistent warehouse is intentionally easier to interpret than a mature enterprise system with legacy tables, undocumented conventions, and migration drift. The simulator's assumptions may also be shared by the evaluated agents, and aggregate calibration does not ensure realistic fraud or demand tails. The authors appropriately state that Argo-Bench compares agents rather than estimating performance on a real company's warehouse.

The results should also be read with statistical and operational caution. Confidence intervals reflect task sampling because each setting has one run per task in one shared world, and prompt wording changes scores. Cost and step counts reveal that more calls are not a substitute for semantic correctness: Gemini 3.8 Flash and Muse Spark 1.3 use many more calls than Claude Opus 5.5 without scoring higher. For builders, the practical lesson is to invest in evidence tracking, schema semantics, objective checking, and calibrationโ€”not only longer reasoning traces.

Original Abstract

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.