Accept 7.4/10 cs.CL trendtoknow-paper-summaries codex-pro/gpt-5.5

AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

Ruifeng Yuan, Yizhi Li, Yaxin Du, Fengyu Cai, Yiqi Liu, Hou Pong Chan, Chenghua Lin, Yun Chen · September 30, 2026 · cs.CL

📌 Highlights

AutoDataBench evaluates data intelligence through controlled data interventions rather than broad, confounded auto-research performance. The benchmark fixes training pipelines, tools, base models, and resource budgets so that improvements are attributable to submitted training data.

The strongest reported overall model is Kimi-3 with 60.68, driven by knowledge injection, while GPT-5.6-Sol leads mean ID tool-use performance at 81.98±0.34 and mean ID retrieval performance at 40.46±1.32.

The main caveat is that target-task optimization does not guarantee generalization. In retrieval, GPT-5.6-Sol has the highest mean ID score but the lowest reported OOD score among LLMs at 26.15.

  • Three benchmark dimensions: data diagnosis and repair, data organization, and data construction.
  • Three tasks: polluted function-calling data, embedding-model training, and knowledge injection.
  • Evaluation combines ID optimization, hidden OOD evaluation, trajectory analysis, and prediction of intervention effects.
  • Auto-research trajectories improve all five reported downstream coding scores when reused during mid-training.

🎯 Introduction

The paper starts from a gap in existing auto-research and capability benchmarks: strong performance on software engineering, GPU kernels, math, scientific reasoning, or terminal tasks does not directly show whether an LLM can understand and improve training data. The authors argue that data quality and composition substantially affect model performance, while evaluating data choices often requires costly training experiments.

AutoDataBench targets this missing capability by evaluating whether LLM agents can diagnose data problems, decide what to select or repair, construct useful new examples, and revise those choices using empirical feedback. The benchmark's objective is not to test whether an agent can produce plausible explanations or data-processing code in isolation, but whether its interventions improve trained models under fixed procedures.

The paper frames data intelligence as a capability that could support future self-improving systems: agents that inspect evidence, design data, train models, evaluate outcomes, and use trajectories as training material for later models.

🔬 Methodology

AutoDataBench defines each task by source data, base model, training procedure, and resource budget. An evaluated LLM iteratively inspects data, implements scalable interventions through code and auxiliary models, submits a training set, and receives target-task feedback. Hidden evaluation data are withheld during optimization and used only to evaluate the checkpoint selected by target-task feedback.

The three task instantiations cover complementary data problems. Tool use uses approximately 40k single-turn examples based on XLAM-FUNCTION-CALLING-60K, with seven probabilistic corruption operators applied to callable training examples and a maximum submitted training set of 16k examples. Retrieval uses the RLHN collection, provides anchors and positives by default, and requires agents to mine hard negatives under a 600k-pair trial limit and a cumulative budget equivalent to ten such trials. Knowledge injection asks agents to construct context-question-answer examples from post-cutoff facts for TALKIE under offline privileged-context distillation.

The knowledge-injection objective is the most explicit mathematical component. The frozen teacher sees context c, question q, and answer prefix, while the student sees only q and the teacher-forced answer prefix. Training then uses a fixed forward KL objective to align the student's answer-token distributions with the context-conditioned teacher, holding the loss function and hyperparameters constant across data interventions.

  • All evaluated LLMs run in a shared framework through external APIs.
  • Each run uses a single NVIDIA A800 GPU and a 24-hour wall-clock limit.
  • Each run permits at most 20 training-and-evaluation rounds.
  • Each training trial restarts from the same initial checkpoint under fixed training settings.

equation 1

Formula:
\[P_t = P_T(\cdot\mid c,q,y_{<t}),\]

Meaning: The teacher distribution over answer tokens at position t, conditioned on supporting context, question, and previous answer tokens.

equation 2

Formula:
\[Q_t = P_S(\cdot\mid q,y_{<t}).\]

Meaning: The student distribution over answer tokens at position t, conditioned on the question and previous answer tokens without access to the supporting context.

📊 Experiments

The experiments evaluate seven frontier LLMs from the GPT, Claude, Kimi, DeepSeek, GLM, and Qwen families. For each LLM-task combination, the authors run three independent standard runs, totaling 63 standard runs. The selected checkpoint in each run is the one with the highest ID score, and Table 1 reports three-run ID means with sample standard deviation, ID best scores, and mean hidden-evaluation OOD scores at the selected checkpoints.

Tool use evaluates macro-averaged AST accuracy on 2,000 in-distribution examples across five categories, with OOD evaluation on held-out BFCL single-turn subsets. Retrieval optimizes mean nDCG@10 across ArguAna, FEVERHardNegatives, FiQA2018, HotpotQAHardNegatives, and SCIDOCS, then evaluates OOD on ClimateFEVERHardNegatives, CQADupstackGamingRetrieval, CQADupstackUnixRetrieval, Touche2020Retrieval.v3, and TRECCOVID. Knowledge injection optimizes Novel accuracy on post-1930 facts and evaluates hidden Retention on pre-1930 knowledge.

The main table shows different winners by task and by generalization. Kimi-3 has the highest Overall score at 60.68 and the highest Knowledge ID mean at 60.60±2.31. GPT-5.6-Sol leads Tool use ID mean at 81.98±0.34 and Retrieval ID mean at 40.46±1.32, but its Retrieval OOD score is 26.15, below the expert reference of 34.76 and below the other LLMs in the table. This is the strongest evidence that target-task feedback can reward strategies that do not transfer.

Trajectory analysis shows that most runs improve with iteration: 61 of 63 standard runs improve on their initial datasets, including all 21 tool-use runs, all 21 retrieval runs, and 19 of 21 knowledge-injection runs. In prediction-enabled runs for GPT-5.6-Sol and Kimi-3, retrieval forecasts track observed scores with mean absolute errors of 0.64 and 1.55 points, but four of six model-task combinations achieve lower best observed scores than the corresponding standard-run mean, suggesting that explicit forecasting may add optimization burden.

tab:main_results

Caption: Performance across tasks (scores $ 100$). Overall averages the three task-wise ID means (Novel for knowledge injection). For each LLM and task, Mean/Best denote the mean/maximum across three standard runs; $ $ indicates sample standard deviation. OOD reports mean hidden-evaluation scores at ID-selected checkpoints: BFCL for tool use, held-out retrieval datasets for retrieval, and Retention for knowledge injection. Bold marks the best LLM result.

Content:
MethodOverallTool use ID meanTool use ID bestTool use OODRetrieval ID meanRetrieval ID bestRetrieval OODKnowledge ID meanKnowledge ID bestKnowledge OOD
Baseline47.9870.45--55.1834.38--31.1339.10--47.84
Expert56.8981.65--60.2340.62--34.7648.40--48.55
DeepSeek-V4-Flash55.2080.97±0.5981.4057.2636.22±1.2637.5929.5348.40±7.7157.3049.38
Qwen-3.7-Max56.8080.82±0.2481.0057.9039.22±2.5141.2530.1850.37±1.3151.4047.33
Claude-4.758.3181.00±0.1581.1558.9538.99±2.3140.8629.6154.93±0.5755.4050.37
Qwen-3.8-Max58.5681.92±0.3382.2559.4439.40±2.6541.6629.6554.37±5.3657.9047.74
GPT-5.6-Sol58.8581.98±0.3482.2557.9240.46±1.3241.5226.1554.10±3.4256.6047.62
GLM-5.259.1081.15±0.3981.6058.2438.55±2.6241.5230.0357.60±0.5258.2047.15
Kimi-360.6881.37±1.5082.4559.3240.07±1.8341.6529.8560.60±2.3162.4048.71

Why It Matters: It contains the paper's decisive benchmark comparisons across baselines, expert references, LLM agents, ID optimization, and OOD generalization.

tab:auto_research_transfer

Caption: Benchmark as Data Engine: Transfer of auto-research trajectories to general coding tasks. Both models use a 5B-token mid-training budget followed by identical SFT.

Content:
MBPP BaseMBPP PlusCRUXEval InputLiveCodeBench v5SWE-bench Multilingual
Control85.4572.7572.0044.8927.67
+ Auto-Research86.2473.2874.1245.8033.33
∆+0.79+0.53+2.12+0.91+5.66

Why It Matters: It reports that adding AutoDataBench trajectories improves MBPP Base, MBPP Plus, CRUXEval Input, LiveCodeBench v5, and SWE-bench Multilingual under a matched training setup.

fig:autodatabench_overview

fig:autodatabench_overview

Caption: Overview of AutoDataBench. (a) Three dimensions of data intelligence: data diagnosis and repair, data organization, and data construction. (b) Agents iteratively propose data interventions, construct training datasets, and use feedback from a fixed training and evaluation pipeline. Final evaluation measures target-task improvement and generalization to held-out evaluation distributions.

Why It Matters: This figure defines the benchmark loop and shows how the three data-intelligence dimensions map to tasks and evaluation.

fig:optimization_curves

fig:optimization_curves

Caption: Best-so-far ID scores in the best-performing standard run of each LLM on each task (scores $ 100$; Knowledge uses Novel). Runs are selected by their maximum ID score, with ties resolved by the earliest run. Each trajectory ends at its own last successful evaluation; failed evaluations are excluded. These selected runs illustrate optimization behavior, not average performance or equal-compute comparisons.

Why It Matters: This result figure shows how agents improve through iterative data interventions and makes clear that final scores can arise from very different optimization paths.

fig:prediction_trajectories

fig:prediction_trajectories

Caption: Predicted and observed scores across training evaluations for GPT-5.6-Sol and Kimi-3. Horizontal lines show the mean best score across standard runs. The first evaluation has no forecast.

Why It Matters: This figure supports the paper's behavioral probe: whether LLM agents can anticipate the empirical effects of their own data changes.

🔮 Conclusion

The paper concludes that AutoDataBench offers a controlled way to evaluate LLM data intelligence through data repair, organization, and construction. Agent-curated data approaches expert references on tool use and retrieval and substantially improves knowledge acquisition relative to the baseline and expert reference in several cases.

The practical takeaway is mixed but useful: frontier LLM agents can improve datasets through iterative experimentation, yet strong target-task scores may fail to transfer to hidden distributions. The benchmark therefore argues for evaluating optimization outcomes, generalization, and prediction of data effects together rather than relying on a single leaderboard number.

The additional mid-training experiment supports a second role for the benchmark: AutoDataBench is not only an evaluator but also a data engine whose trajectories can train coding models on evidence inspection, hypothesis formation, and iterative action.

🛠️ Future Research Improvements

The most direct improvement is broader task coverage. The authors emphasize extensibility, but the main body instantiates only three initial tasks. Future versions should test more modalities, larger model families, more training paradigms, and domains where data quality has different failure modes.

The prediction protocol deserves expansion. Because prediction-enabled runs are limited to GPT-5.6-Sol and Kimi-3 on each task, broader evaluation is needed to separate genuine data-effect reasoning from post-hoc adaptation to feedback.

Future work should better isolate causal mechanisms inside successful trajectories. The paper provides concrete examples such as GLM-5.2 targeting argument accuracy, Qwen-3.8-Max changing retrieval source sampling and negative assignment, and Kimi-3 adjusting QA-pair construction, but these are descriptive rather than controlled ablations of individual interventions.

  • Add more benchmark tasks and training paradigms.
  • Run prediction-enabled evaluations across all candidate LLMs.
  • Introduce intervention-level ablations within agent trajectories.
  • Measure whether trajectory mid-training benefits persist at larger scale.

🏭 Potential Industry Use Scenarios

AutoDataBench is directly relevant to teams building data-curation agents for post-training, retrieval systems, tool-use models, and knowledge-update pipelines. Its controlled setup gives engineering teams a way to test whether an agent's proposed data changes actually improve trained models under fixed infrastructure.

The hidden-evaluation design is valuable for production ML workflows where overfitting to visible validation metrics is a real risk. Builders can adapt the benchmark pattern by keeping deployment-like evaluations hidden until after data-policy selection.

The mid-training result suggests a training-data product opportunity: collect high-quality research-agent trajectories that include inspection, diagnosis, hypothesis formation, data construction, and evaluation feedback, then reuse them to improve coding or repository-level reasoning models. The evidence is promising but limited to the reported Qwen2.5-Coder-14B setup with a 5B-token budget and 80M upsampled auto-research trajectory tokens.

  • Benchmark data-curation agents before deploying them into model-training pipelines.
  • Use hidden OOD checks to detect benchmark over-optimization.
  • Mine agent research trajectories as training data for coding and repository-understanding models.
  • Compare agent-generated data against expert references under identical training recipes.

💬 Critical Analysis

The paper's strongest design choice is its data-centrality. By fixing non-data components, AutoDataBench makes a cleaner measurement target than broad auto-research benchmarks where better results may come from hyperparameters, framework changes, compute choices, or other confounders.

The main empirical story is nuanced. The benchmark shows that LLMs can create useful data interventions, especially in tool-use repair and knowledge construction, but it also shows that visible target-task feedback can mislead agents. Retrieval is the cautionary example: GPT-5.6-Sol nearly matches the expert ID reference but generalizes poorly to held-out retrieval distributions.

Reproducibility depends on released code, task data, fixed training recipes, and access to comparable frontier models through APIs. The benchmark design is careful, but the frontier-model component means exact reproduction may be sensitive to model availability, API versioning, and run-to-run stochasticity.

The transfer experiment is intriguing but should not be overread. Improvements across all five coding evaluations support the claim that auto-research trajectories are useful data, yet the main body reports one matched setup starting from Qwen2.5-Coder-14B with a reduced 5B-token mid-training budget. Builders should treat this as a strong signal for follow-up experiments, not as a universal recipe.

Original Abstract

Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.