构建首个大规模专家标注阅读清单数据集,助力高效文献学习。
ACL-rlg: A Dataset for Reading List Generation
- 构建开源专家标注的阅读清单数据集ACL-rlg
- 验证传统检索工具在阅读清单生成上表现不佳
- 适合文献综述、科研入门与AI辅助学习研究者
了解新科学领域及其现有文献可能令人望而生畏,因可用论文数量庞大。由专家整理的学术参考文献清单(即阅读清单)可提供结构化路径,帮助全面掌握某一领域或具体科学问题。本文提出ACL-rlg,目前最大规模的公开专家标注阅读清单数据集,并为阅读清单生成任务提供多个基线方法,正式将其定义为一项检索任务。定性研究表明,传统学术搜索引擎和索引方法在此任务上表现较差;尽管GPT-4o表现更优,但存在潜在数据污染迹象。
原文摘要 · Abstract (English)
Familiarizing oneself with a new scientific field and its existing literature can be daunting due to the large amount of available articles. Curated lists of academic references, or reading lists, compiled by experts, offer a structured way to gain a comprehensive overview of a domain or a specific scientific challenge. In this work, we introduce ACL-rlg, the largest open expert-annotated reading list dataset. We also provide multiple baselines for evaluating reading list generation and formally define it as a retrieval task. Our qualitative study highlights the fact that traditional scholarly search engines and indexing methods perform poorly on this task, and GPT-4o, despite showing better results, exhibits signs of potential data contamination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。