arXiv:2510.24478cs.CL2025-10

构建首个科学演讲引文预测数据集,助力语音内容自动匹配相关文献。

Talk2Ref: A Dataset for Reference Prediction from Scientific Talks

  • 提出从长篇无结构演讲中预测相关论文的新任务
  • 构建包含6279个演讲与43429篇引用论文的数据集,平均每讲26篇
  • 模型在演讲语音内容上微调后引文预测性能显著提升

科学演讲正成为科研成果传播的重要形式,自动识别能支撑或丰富演讲内容的相关文献对研究者和学生极具价值。本文提出「从演讲中预测引文」(Reference Prediction from Talks, RPT)这一新任务,旨在将长篇、非结构化的科学演讲映射到相关论文。为支持该任务研究,我们构建了首个大规模数据集 Talk2Ref,包含6,279个演讲和43,429篇被引用论文(平均每篇演讲对应26篇),其相关性以演讲对应论文的引用文献为准。我们评估了主流文本嵌入模型在零样本检索中的表现,并提出一种基于 Talk2Ref 训练的双编码器架构。进一步探索了处理长文本转录稿的方法及领域自适应训练策略。实验表明,在 Talk2Ref 上微调可显著提升引文预测性能,验证了该任务的挑战性与数据集在学习口语科学内容语义表征方面的有效性。数据集与训练模型已开源,推动语音科学交流与引文推荐系统的融合研究。

原文摘要 · Abstract (English)

Scientific talks are a growing medium for disseminating research, and automatically identifying relevant literature that grounds or enriches a talk would be highly valuable for researchers and students alike. We introduce Reference Prediction from Talks (RPT), a new task that maps long, and unstructured scientific presentations to relevant papers. To support research on RPT, we present Talk2Ref, the first large-scale dataset of its kind, containing 6,279 talks and 43,429 cited papers (26 per talk on average), where relevance is approximated by the papers cited in the talk's corresponding source publication. We establish strong baselines by evaluating state-of-the-art text embedding models in zero-shot retrieval scenarios, and propose a dual-encoder architecture trained on Talk2Ref. We further explore strategies for handling long transcripts, as well as training for domain adaptation. Our results show that fine-tuning on Talk2Ref significantly improves citation prediction performance, demonstrating both the challenges of the task and the effectiveness of our dataset for learning semantic representations from spoken scientific content. The dataset and trained models are released under an open license to foster future research on integrating spoken scientific communication into citation recommendation systems.

引文预测语音理解科学传播数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。