Accept 7.3/10 cs.AI trendtoknow-paper-summaries codex-pro/gpt-5.6-luna

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh ยท October 01, 2026 ยท cs.AI, cs.CL, cs.IR

๐Ÿ“Œ Highlights

ScholarCatalyst turns the informal notion of research taste into a measurable retrieval problem: given an early-stage research question, retrieve prior papers whose ideas could help the project progress. Its author-centered labels target inspiration that is often absent from citations and difficult to infer from the final paper.

The strongest evidence points to a candidate-coverage bottleneck. Embedding retrieval reaches 0.48 Recall@20, agentic search reaches 0.42 despite using retrieval tools, and Claude Fable 5.1 reaches only 0.51 in a post-cutoff comparison. The paper therefore motivates training retrieval systems on expert judgments about useful ideas, not only on semantic similarity or citation behavior.

  • 894 author-validated queries from 184 researchers and 207 projects.
  • Approximately 191K temporally filtered computer-science papers in the evaluation corpus.
  • 43.6% of positive papers are uncited by the source paper.
  • Agentic search does not outperform embedding retrieval because many catalyst papers never enter the candidate pool.

๐ŸŽฏ Introduction

The paper argues that scientific progress is cumulative but that current AI systems mostly accelerate already-defined subproblems. Choosing which existing idea a vague research question should build on remains a human-centered capability: a prior paper may reveal a limitation, suggest an adaptable approach, or turn an open question into a tractable project. ScholarCatalyst calls such prior work catalyst papers and evaluates whether models can retrieve it from an early-stage question.

The key difficulty is that completed papers and bibliographies do not fully record the intellectual path that produced them. Existing benchmarks focus on claim support, facet-level relevance, or citation relations, whereas ScholarCatalyst asks whether a paper could have advanced the project before its solution was known. Its benchmark distinction is summarized in the comparison with prior literature-retrieval datasets.

tab:benchmark_comparison

Caption: Comparison with literature retrieval benchmarks. ScholarCatalyst combines queries that precede a project's solution (Pre-discovery query), direct author annotation (Author annotation), and relevance judgment beyond what is published in the completed source paper (Undocumented relevance).

Content:
BenchmarkRelevancePre-discovery queryAuthor annotationUndocumented relevanceQueriesCorpus
SciFact (Wadden et al., 2020)Claim supportโœ—โœ—โœ—1.4K5.2K
DORIS-MAE (Wang et al., 2023)Multi-aspect relevanceโœ—โœ—โœ—100100
ScholarQABench (Asai et al., 2024)Claim supportโœ—โœ—โœ—2,96745M
LitSearch (Ajith et al., 2024)Citation relationโœ—โœ“โœ—59764K
MIR (Garikaparthi et al., 2025)Methodological citation intentโœ“โœ—โœ—1394.7K
ScholarCatalyst (ours)Author-judged inspirationโœ“โœ“โœ“894191K

Why It Matters: The comparison positions ScholarCatalyst as combining pre-discovery queries, direct author annotation, and relevance judgments that are not recoverable from the completed paper.

๐Ÿ”ฌ Methodology

Each task instance is built from a source paper S, an early-stage query q, positive catalyst papers, and author-reviewed hard negatives. Core Research Queries address the central research question; Subfield-specific Queries approach that question from a particular research subfield. The retrieval corpus contains only papers published before the source paper was completed, and external web access is disabled. Systems are evaluated with Recall@5, Recall@20, Recall@100, and nDCG@20.

The data-collection pipeline first obtains source-paper metadata and files through arXiv, parses the full text, and resolves bibliography entries into a one-hop citation graph. It combines resolved references with 181K additional arXiv papers from 2020โ€“2024, yielding 190,896 papers. Gemini 3.1 Pro drafts queries and rationales; BM25 and Qwen3-Embedding-8B retrieve additional uncited candidates; Gemini 3.6 Flash reranks the pooled candidates; and authors validate or rewrite queries and label each candidate. The pipeline retains 10 candidates per query for review.

tab:corpus_stats

Caption: Dataset statistics. For each query type, we report the number of queries, the average number of positive papers (|D+q|) and hard negatives (|Dโˆ’q|), and the average query length in tokens.

Content:
Query typeQueriesAvg.D+qCited / Uncited Pos. (%)Avg.Dโˆ’qAvg. Length
Core Research Query2073.1995.5 / 4.510.01131.8
Subfield-specific Query6876.0956.4 / 43.66.8863.5

Why It Matters: The dataset has 207 broad queries and 687 subfield-specific queries, with different positive-set sizes, citation rates, hard-negative counts, and query lengths.

๐Ÿ“Š Experiments

The benchmark evaluates BM25, LateOn, seven dense retrievers including Qwen3-Embedding-4B and 8B, Gemini Embedding 2, SPECTER2, OpenScholar, ReasonEmbed, and Inf-Retriever-v1-Pro, plus GPT-4.1 and o3 search agents. The main evaluation uses a temporally filtered local corpus and reports nDCG@20 and Recall@5, Recall@20, and Recall@100. Claude Fable 5.1 is reported separately as a post-cutoff comparison because its training data may include the source papers.

All systems struggle with catalyst retrieval. General-purpose dense retrievers place 39% and 51% of author-identified inspiration papers in the top 20 for the two query types, while scientific-document retrievers trail by at least 16 and 27 Recall@20 points on those query types. BM25 and LateOn trail by 15โ€“16 and 18โ€“25 points. Agentic search does not solve the problem: the abstract reports 0.42 Recall@20 for agentic search versus 0.48 for embedding retrieval, while the grep agent encounters only 8% of gold queryโ€“paper pairs during search compared with 46% for the tool-calling agent. Fable 5.1 reaches 0.51 Recall@20 despite its potential training-data advantage.

The ablations identify coverage rather than document comprehension as the main bottleneck. Adding source title and abstract raises Recall@20 by 0.11 for one query type and 0.02 for the other; the source bibliography, which contains 95% of gold papers, raises it by 0.35. Without search tools, title-and-abstract context reaches only 0.31 Recall@20. Reading beyond abstracts changes Recall@20 by at most 0.04, and the best reading policy still trails the retriever at 0.50 versus 0.55. Query expansion and multi-query generation do not meaningfully improve recall, while HyDE variants reduce it.

tab:corpus_stats

Caption: Dataset statistics. For each query type, we report the number of queries, the average number of positive papers (|D+q|) and hard negatives (|Dโˆ’q|), and the average query length in tokens.

Content:
Query typeQueriesAvg.D+qCited / Uncited Pos. (%)Avg.Dโˆ’qAvg. Length
Core Research Query2073.1995.5 / 4.510.01131.8
Subfield-specific Query6876.096.0956.4 / 43.66.8863.5

Why It Matters: The dataset has 207 broad queries and 687 subfield-specific queries, with different positive-set sizes, citation rates, hard-negative counts, and query lengths.

fig:benchmark_task

fig:benchmark_task

Caption: asks whether a system can retrieve prior papers that inspire progress on a research project. Such a paper may share little semantic similarity with the question yet offer an insight that advances the project (right), while topically related work may not (left).

Why It Matters: This figure captures the benchmark's core distinction between topical relevance and research-inspiring usefulness.

fig:benchmark_example

fig:benchmark_example

Caption: An example instance. An author of the source paper $S$ provides a research question $q$, expressed as a or , along with positive and hard negative papers $D_q^+, D_q^-$ and a rationale for these judgments.

Why It Matters: It shows the unit of supervision: an early-stage question, author-judged positives, hard negatives, and explanatory rationales.

fig:benchmark_construction

fig:benchmark_construction

Caption: construction process. Our automated data-collection pipeline constructs a retrieval corpus and author-review materials, including structured paper data and candidate papers (sec:bench:pipeline), which authors review through our interface to validate and revise queries, judge candidate-paper relevance, and provide rationales (sec:bench:author_data).

Why It Matters: The pipeline is the scaling mechanism that turns author knowledge about completed projects into benchmark supervision.

fig:agent_coverage

fig:agent_coverage

Caption: Grep rarely surfaces gold papers. Share of gold papers that each GPT-4.1 agent encounters during search (Found) and returns in its final top 20.

Why It Matters: It visualizes why agentic reasoning is limited: agents cannot select papers they never encounter during search.

๐Ÿ”ฎ Conclusion

ScholarCatalyst shows that identifying papers capable of advancing a research project is substantially harder than retrieving papers that support a claim or resemble a query. Current systems recover only a fraction of author-credited catalyst papers, and agentic search remains close to embedding retrieval because useful papers are often absent from the candidate pool.

The authors identify two longer-term milestones: expert-level retrieval within a field and retrieval with expert-level intuition across fields. The second may be especially valuable because 49.5% of credited papers come from outside the project's subject area, while 43.6% of subfield-specific positives were never cited by the source paper.

๐Ÿ› ๏ธ Future Research Improvements

The most direct improvement is to train retrieval models on the benchmark's author judgments and rationales, with explicit supervision for uncited positives and hard negatives that are topically related but not useful. The reported gold-injection analysis suggests that stronger rerankers can benefit substantially when missing catalyst papers are inserted into the candidate pool, so future systems should jointly improve first-stage coverage and reranking.

Future evaluations should expand beyond the current computer-science corpus and test whether expert-level retrieval transfers across disciplines. They should also establish a more reliable performance ceiling through carefully designed coauthor or domain-expert studies, while recognizing that outsider relabeling would measure informed inference rather than the original author's research process.

๐Ÿญ Potential Industry Use Scenarios

A credible application is an AI research assistant that receives a half-formed technical question and surfaces cross-disciplinary papers that could reshape the project, rather than returning only papers with matching terminology. Such a system could support research planning, prior-art exploration, and hypothesis formation when the relevant idea is not explicitly cited or described in the query.

The benchmark also supports evaluation infrastructure for enterprise research agents: teams could measure candidate coverage, uncited-paper recall, and expert-judged usefulness before deploying agents for literature review. These are proposed use scenarios rather than capabilities demonstrated in a production system; the paper's results indicate that reliable deployment would require improved retrieval coverage and expert-aligned training.

๐Ÿ’ฌ Critical Analysis

The strongest design choice is the use of firsthand author supervision combined with temporal filtering. This targets a scientifically meaningful label that ordinary citation graphs omit and prevents systems from relying on later accounts of the completed project. The automated pipeline is also practically important because it turns a source-paper identifier into author-review materials and allows the benchmark to be refreshed as newer models absorb older papers.

The main weakness is that the target is retrospective and expensive to validate, so the benchmark may reflect authors' remembered intellectual histories rather than an independently auditable ground truth. The restricted 191K-paper corpus also understates the open-literature search problem. Finally, the experiments indicate that current agents are bottlenecked before reasoning begins: better reading, query rewriting, or iterative planning cannot recover papers that lexical or embedding retrieval never surfaces.

Original Abstract

What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.