๐ Highlights
ScholarCatalyst turns the informal notion of research taste into a measurable retrieval problem: given an early-stage research question, retrieve prior papers whose ideas could help the project progress. Its author-centered labels target inspiration that is often absent from citations and difficult to infer from the final paper.
The strongest evidence points to a candidate-coverage bottleneck. Embedding retrieval reaches 0.48 Recall@20, agentic search reaches 0.42 despite using retrieval tools, and Claude Fable 5.1 reaches only 0.51 in a post-cutoff comparison. The paper therefore motivates training retrieval systems on expert judgments about useful ideas, not only on semantic similarity or citation behavior.
- 894 author-validated queries from 184 researchers and 207 projects.
- Approximately 191K temporally filtered computer-science papers in the evaluation corpus.
- 43.6% of positive papers are uncited by the source paper.
- Agentic search does not outperform embedding retrieval because many catalyst papers never enter the candidate pool.
๐ฏ Introduction
The paper argues that scientific progress is cumulative but that current AI systems mostly accelerate already-defined subproblems. Choosing which existing idea a vague research question should build on remains a human-centered capability: a prior paper may reveal a limitation, suggest an adaptable approach, or turn an open question into a tractable project. ScholarCatalyst calls such prior work catalyst papers and evaluates whether models can retrieve it from an early-stage question.
The key difficulty is that completed papers and bibliographies do not fully record the intellectual path that produced them. Existing benchmarks focus on claim support, facet-level relevance, or citation relations, whereas ScholarCatalyst asks whether a paper could have advanced the project before its solution was known. Its benchmark distinction is summarized in the comparison with prior literature-retrieval datasets.
tab:benchmark_comparison
Caption: Comparison with literature retrieval benchmarks. ScholarCatalyst combines queries that precede a project's solution (Pre-discovery query), direct author annotation (Author annotation), and relevance judgment beyond what is published in the completed source paper (Undocumented relevance).
| Benchmark | Relevance | Pre-discovery query | Author annotation | Undocumented relevance | Queries | Corpus |
|---|---|---|---|---|---|---|
| SciFact (Wadden et al., 2020) | Claim support | โ | โ | โ | 1.4K | 5.2K |
| DORIS-MAE (Wang et al., 2023) | Multi-aspect relevance | โ | โ | โ | 100 | 100 |
| ScholarQABench (Asai et al., 2024) | Claim support | โ | โ | โ | 2,967 | 45M |
| LitSearch (Ajith et al., 2024) | Citation relation | โ | โ | โ | 597 | 64K |
| MIR (Garikaparthi et al., 2025) | Methodological citation intent | โ | โ | โ | 139 | 4.7K |
| ScholarCatalyst (ours) | Author-judged inspiration | โ | โ | โ | 894 | 191K |
Why It Matters: The comparison positions ScholarCatalyst as combining pre-discovery queries, direct author annotation, and relevance judgments that are not recoverable from the completed paper.
๐ฌ Methodology
Each task instance is built from a source paper S, an early-stage query q, positive catalyst papers, and author-reviewed hard negatives. Core Research Queries address the central research question; Subfield-specific Queries approach that question from a particular research subfield. The retrieval corpus contains only papers published before the source paper was completed, and external web access is disabled. Systems are evaluated with Recall@5, Recall@20, Recall@100, and nDCG@20.
The data-collection pipeline first obtains source-paper metadata and files through arXiv, parses the full text, and resolves bibliography entries into a one-hop citation graph. It combines resolved references with 181K additional arXiv papers from 2020โ2024, yielding 190,896 papers. Gemini 3.1 Pro drafts queries and rationales; BM25 and Qwen3-Embedding-8B retrieve additional uncited candidates; Gemini 3.6 Flash reranks the pooled candidates; and authors validate or rewrite queries and label each candidate. The pipeline retains 10 candidates per query for review.
tab:corpus_stats
Caption: Dataset statistics. For each query type, we report the number of queries, the average number of positive papers (|D+q|) and hard negatives (|Dโq|), and the average query length in tokens.
| Query type | Queries | Avg. | D+q | Cited / Uncited Pos. (%) | Avg. | Dโq | Avg. Length | ||
|---|---|---|---|---|---|---|---|---|---|
| Core Research Query | 207 | 3.19 | 95.5 / 4.5 | 10.01 | 131.8 | ||||
| Subfield-specific Query | 687 | 6.09 | 56.4 / 43.6 | 6.88 | 63.5 |
Why It Matters: The dataset has 207 broad queries and 687 subfield-specific queries, with different positive-set sizes, citation rates, hard-negative counts, and query lengths.
๐ Experiments
The benchmark evaluates BM25, LateOn, seven dense retrievers including Qwen3-Embedding-4B and 8B, Gemini Embedding 2, SPECTER2, OpenScholar, ReasonEmbed, and Inf-Retriever-v1-Pro, plus GPT-4.1 and o3 search agents. The main evaluation uses a temporally filtered local corpus and reports nDCG@20 and Recall@5, Recall@20, and Recall@100. Claude Fable 5.1 is reported separately as a post-cutoff comparison because its training data may include the source papers.
All systems struggle with catalyst retrieval. General-purpose dense retrievers place 39% and 51% of author-identified inspiration papers in the top 20 for the two query types, while scientific-document retrievers trail by at least 16 and 27 Recall@20 points on those query types. BM25 and LateOn trail by 15โ16 and 18โ25 points. Agentic search does not solve the problem: the abstract reports 0.42 Recall@20 for agentic search versus 0.48 for embedding retrieval, while the grep agent encounters only 8% of gold queryโpaper pairs during search compared with 46% for the tool-calling agent. Fable 5.1 reaches 0.51 Recall@20 despite its potential training-data advantage.
The ablations identify coverage rather than document comprehension as the main bottleneck. Adding source title and abstract raises Recall@20 by 0.11 for one query type and 0.02 for the other; the source bibliography, which contains 95% of gold papers, raises it by 0.35. Without search tools, title-and-abstract context reaches only 0.31 Recall@20. Reading beyond abstracts changes Recall@20 by at most 0.04, and the best reading policy still trails the retriever at 0.50 versus 0.55. Query expansion and multi-query generation do not meaningfully improve recall, while HyDE variants reduce it.
tab:corpus_stats
Caption: Dataset statistics. For each query type, we report the number of queries, the average number of positive papers (|D+q|) and hard negatives (|Dโq|), and the average query length in tokens.
| Query type | Queries | Avg. | D+q | Cited / Uncited Pos. (%) | Avg. | Dโq | Avg. Length | ||
|---|---|---|---|---|---|---|---|---|---|
| Core Research Query | 207 | 3.19 | 95.5 / 4.5 | 10.01 | 131.8 | ||||
| Subfield-specific Query | 687 | 6.09 | 6.09 | 56.4 / 43.6 | 6.88 | 63.5 |
Why It Matters: The dataset has 207 broad queries and 687 subfield-specific queries, with different positive-set sizes, citation rates, hard-negative counts, and query lengths.
fig:benchmark_task

Caption: asks whether a system can retrieve prior papers that inspire progress on a research project. Such a paper may share little semantic similarity with the question yet offer an insight that advances the project (right), while topically related work may not (left).
Why It Matters: This figure captures the benchmark's core distinction between topical relevance and research-inspiring usefulness.
fig:benchmark_example

Caption: An example instance. An author of the source paper $S$ provides a research question $q$, expressed as a or , along with positive and hard negative papers $D_q^+, D_q^-$ and a rationale for these judgments.
Why It Matters: It shows the unit of supervision: an early-stage question, author-judged positives, hard negatives, and explanatory rationales.
fig:benchmark_construction

Caption: construction process. Our automated data-collection pipeline constructs a retrieval corpus and author-review materials, including structured paper data and candidate papers (sec:bench:pipeline), which authors review through our interface to validate and revise queries, judge candidate-paper relevance, and provide rationales (sec:bench:author_data).
Why It Matters: The pipeline is the scaling mechanism that turns author knowledge about completed projects into benchmark supervision.
fig:agent_coverage

Caption: Grep rarely surfaces gold papers. Share of gold papers that each GPT-4.1 agent encounters during search (Found) and returns in its final top 20.
Why It Matters: It visualizes why agentic reasoning is limited: agents cannot select papers they never encounter during search.
๐ฎ Conclusion
ScholarCatalyst shows that identifying papers capable of advancing a research project is substantially harder than retrieving papers that support a claim or resemble a query. Current systems recover only a fraction of author-credited catalyst papers, and agentic search remains close to embedding retrieval because useful papers are often absent from the candidate pool.
The authors identify two longer-term milestones: expert-level retrieval within a field and retrieval with expert-level intuition across fields. The second may be especially valuable because 49.5% of credited papers come from outside the project's subject area, while 43.6% of subfield-specific positives were never cited by the source paper.
๐ ๏ธ Future Research Improvements
The most direct improvement is to train retrieval models on the benchmark's author judgments and rationales, with explicit supervision for uncited positives and hard negatives that are topically related but not useful. The reported gold-injection analysis suggests that stronger rerankers can benefit substantially when missing catalyst papers are inserted into the candidate pool, so future systems should jointly improve first-stage coverage and reranking.
Future evaluations should expand beyond the current computer-science corpus and test whether expert-level retrieval transfers across disciplines. They should also establish a more reliable performance ceiling through carefully designed coauthor or domain-expert studies, while recognizing that outsider relabeling would measure informed inference rather than the original author's research process.
๐ญ Potential Industry Use Scenarios
A credible application is an AI research assistant that receives a half-formed technical question and surfaces cross-disciplinary papers that could reshape the project, rather than returning only papers with matching terminology. Such a system could support research planning, prior-art exploration, and hypothesis formation when the relevant idea is not explicitly cited or described in the query.
The benchmark also supports evaluation infrastructure for enterprise research agents: teams could measure candidate coverage, uncited-paper recall, and expert-judged usefulness before deploying agents for literature review. These are proposed use scenarios rather than capabilities demonstrated in a production system; the paper's results indicate that reliable deployment would require improved retrieval coverage and expert-aligned training.
๐ฌ Critical Analysis
The strongest design choice is the use of firsthand author supervision combined with temporal filtering. This targets a scientifically meaningful label that ordinary citation graphs omit and prevents systems from relying on later accounts of the completed project. The automated pipeline is also practically important because it turns a source-paper identifier into author-review materials and allows the benchmark to be refreshed as newer models absorb older papers.
The main weakness is that the target is retrospective and expensive to validate, so the benchmark may reflect authors' remembered intellectual histories rather than an independently auditable ground truth. The restricted 191K-paper corpus also understates the open-literature search problem. Finally, the experiments indicate that current agents are bottlenecked before reasoning begins: better reading, query rewriting, or iterative planning cannot recover papers that lexical or embedding retrieval never surfaces.