arXiv:2511.04921cs.CL2025-11被引 2

用论文引用网络自动推荐实验数据和基线,提升科研效率。

AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent

  • 基于论文引用网络构建数据与基线的联合检索框架
  • 在顶会数据上召回率提升5.85%,命中率提升8.30%
  • 支持可解释推理,适合需要高效实验设计的研究者

大型语言模型代理在信息检索、复杂推理等网络任务中表现日益出色,推动了其在科学探索中的应用。其中,自动化实验设计通过智能检索数据集与基线成为关键方向。然而,现有方法存在数据覆盖不足的问题:推荐数据主要来自公开门户,遗漏了大量实际发表论文中使用的数据;且过度依赖内容相似性,导致模型偏向表面匹配而忽略实验适用性。为此,我们提出一种融合学术引用网络集体认知的完整推荐框架。首先,设计自动化数据采集管道,将约十万篇已接受论文与其实际使用过的基线和数据集进行关联。其次,提出增强型检索器,通过拼接自描述与聚合引用上下文来表征每个数据集或基线在学术网络中的位置,并微调嵌入模型实现高效候选召回。最后,构建推理增强重排序器,生成显式推理链,并微调大模型以输出可解释的说明与优化排名。所构建数据集覆盖过去五年顶会中85%的数据集与基线。在该数据集上,所提方法相比最强基线,平均在Recall@20上提升5.85%,HitRate@5提升8.30%。结果表明,本工作显著提升了实验设计自动化的可靠性与可解释性。

原文摘要 · Abstract (English)

Large language model agents are becoming increasingly capable at web-centric tasks such as information retrieval, complex reasoning. These emerging capabilities have given rise to surge research interests in developing LLM agent for facilitating scientific quest. One key application in AI research is to automate experiment design through agentic dataset and baseline retrieval. However, prior efforts suffer from limited data coverage, as recommendation datasets primarily harvest candidates from public portals and omit many datasets actually used in published papers, and from an overreliance on content similarity that biases model toward superficial similarity and overlooks experimental suitability. Harnessing collective perception embedded in the baseline and dataset citation network, we present a comprehensive framework for baseline and dataset recommendation. First, we design an automated data-collection pipeline that links roughly one hundred thousand accepted papers to the baselines and datasets they actually used. Second, we propose a collective perception enhanced retriever. To represent the position of each dataset or baseline within the scholarly network, it concatenates self-descriptions with aggregated citation contexts. To achieve efficient candidate recall, we finetune an embedding model on these representations. Finally, we develop a reasoning-augmented reranker that exact interaction chains to construct explicit reasoning chains and finetunes a large language model to produce interpretable justifications and refined rankings. The dataset we curated covers 85\% of the datasets and baselines used at top AI conferences over the past five years. On our dataset, the proposed method outperforms the strongest prior baseline with average gains of +5.85\% in Recall@20, +8.30\% in HitRate@5. Taken together, our results advance reliable, interpretable automation of experimental design.

实验设计智能代理推荐系统引用网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。