Skip to content

Finding a Paper You Can Put to Work

Asked to find earlier work useful to completed computer-science projects, ScholarCatalyst’s GPT-4.1 and o3 search agents recovered fewer than half the authors’ recorded helpful papers in their first 20 results, averaged across questions.

HomeBench checked whether a language model could turn instructions into device commands that met predefined specifications. In a released retrospective account, a contributor to a later smart-home benchmark, SimuHome, describes what made that rule useful: a limitation to address.

The contributor explains that different valid sets of actions can sometimes produce the desired home conditions. They credit HomeBench with helping motivate a different evaluation: check whether the agent reached the intended state, rather than matching its actions to a predefined answer.

The published SimuHome system checks achieved states for feasible tasks involving changes in its simulated home. It still requires certain preliminary actions, and uses a language-model judge for other requests. The earlier paper had helped identify a particular design choice to change.

That account appears in ScholarCatalyst, a new preprint that makes researchers’ reasons for valuing earlier work part of a search test. Its 184 contributors judged papers against 207 completed computer-science projects. They identified work that did or could have helped, so the collection includes remembered influences and possibilities recognized afterward.

Across 894 reconstructed research questions, the strongest embedding retriever, Qwen3-Embedding-8B, recovered about 48% of the recorded helpful papers in its first 20 results, averaged over questions. An embedding retriever uses learned numerical representations of text to rank related papers. The best of the three main search agents recovered about 42%; those agents used GPT-4.1 and o3, with stated training cutoffs before all the completed source papers.

These are fractions of the authors’ recorded helpful sets, not the useful fraction of each result list. Missing half that set does not mean half the searches found nothing. The experiment also had no matched comparison with people performing the search.

The result concerns a recognizable research task: find earlier work that helps with the question in front of you. The SimuHome account makes clear why the help needs a description. A researcher can value a paper for a feature they intend to replace.

A comparison can be the point

Another contributor needed almost the opposite relationship to an earlier method: keep it simple enough to use as a comparison.

The project, Benefits and Limitations of Communication in Multi-Agent Reasoning, studied what language models could accomplish by communicating. In the released account, its contributor credits Self-Consistency: generate several reasoning paths independently, then choose the most common answer.

Here the useful feature was that the models did not communicate while producing those paths. It gave the researcher a straightforward alternative to the more elaborate arrangements they were studying. Their published experiments use self-consistency with majority voting as the baseline for a task that asks models to return the value paired with a given identifier.

One account concerns a limitation to address; the other, a comparison to preserve. Neither requires a leap between remote fields. HomeBench and SimuHome even concern the same kind of agent. A topical match can be a good beginning without specifying whether the reader should borrow a method, compare against it, or change it.

The different purposes also appear in an independent study of 20 academic and industry data scientists, conducted in 2022 through interviews and think-aloud searches. One participant sought applications to data or use cases like their own. Another wanted existing approaches to use as quantitative comparisons. Those readers needed different things from an account of how a method worked.

Calling both papers “relevant” leaves that distinction for the person opening them to resolve.

The question is still taking shape

Finding the right purpose can be part of the search itself. In the data-scientist study, one participant described returning to a search over several weeks before finding the vocabulary used by the relevant community. They had an idea, but did not yet know how that community described it.

A list of related papers can help someone learn how to ask. That is a different starting point from a finished paragraph describing the research problem.

ScholarCatalyst builds its questions by looking backward. A model reads a completed paper and proposes questions and explanations of why earlier papers might help. Project authors can validate, revise or rewrite them. The resulting question may describe the original problem faithfully without reproducing the original researcher’s knowledge or language.

The helpful-paper sets need not be complete. Authors assessed proposed references and a pool partly assembled by systems later evaluated, including Qwen3’s retriever. Useful work outside those judgments could be returned without receiving credit. The score measures recovery of this particular record of usefulness.

The comparison between systems is similarly specific. Qwen3’s approximately 48% is the strongest pooled embedding result. The agent’s actual search tool, Gemini Embedding 2, scored about 44% on its own. The agents had limited search budgets and searched a restricted corpus without web access or citation-following in the main protocol. No uncertainty intervals accompany the main comparison. The reported results do not establish that adding a search agent can never help.

What the benchmark contributes is an inspectable target. For each recommendation judged helpful, it records a reason that can be checked against the earlier work and the question it is meant to help answer.

The paper can leave work unfinished

A reason is not a guarantee. In the SimuHome case, HomeBench’s actual evaluation description supports the claim about predefined specifications. The later simulator documents its different checks. In the comparison case, the earlier voting method and the later experiment establish what was used. Those checks substantiate the technical relationship; they do not independently verify the contributor’s remembered discovery process.

A returned title gives the reader somewhere to look. An explanation of its intended use identifies what they need to check.

For the person trying to apply a method, another obstacle can appear after the search has succeeded. The data-scientist interviews describe missing implementation details, including parameter choices absent from papers. Some readers turned to code to understand the method; one described asking authors whether code was available.

The short think-aloud study did not follow those requests through to completed implementations. But the reported next action is concrete. A reader can have found a useful paper and still need to ask its author for the code.

A note on the dates

The released retrieval code filters by source-paper publication month, not a measured project-start date. A code-and-data check found 15 labels linking questions to helpful papers excluded by that rule but retained in the scoring denominator. Removing them could raise pooled recall by at most about 0.28 percentage points with rankings unchanged. This small inconsistency cannot explain the principal result; the original model runs were not reproduced.

Sources