用提问代替关键词,让模型更智能地找参考文献。
No Stupid Questions: An Analysis of Question Query Generation for Citation Recommendation
- 让大模型主动提问,挖掘文章隐含信息作为检索线索。
- 部分问题查询效果优于传统关键词提取,最高提升12.3%相关性。
- 提出新评估方法MMR-RBO,适合研究者优化提问策略。
现有引用推荐技术受限于对文章内容和元数据的依赖。本文利用GPT-4o-mini的潜在专家能力,指导其针对科学文章片段提出问题,这些问题若被回答,可揭示新见解。我们评估这些提问作为检索查询的有效性,衡量其在召回和排序被遮蔽目标文档方面的表现。实验表明,部分生成的问题查询效果优于同一模型生成的关键词抽取查询。此外,我们提出一种基于秩偏重叠(RBO)的MMR-RBO变体,用于识别表现接近关键词基线的问题。由于所有问题查询均产生独特结果集,我们认为没有愚蠢的问题。
原文摘要 · Abstract (English)
Existing techniques for citation recommendation are constrained by their adherence to article contents and metadata. We leverage GPT-4o-mini's latent expertise as an inquisitive assistant by instructing it to ask questions which, when answered, could expose new insights about an excerpt from a scientific article. We evaluate the utility of these questions as retrieval queries, measuring their effectiveness in retrieving and ranking masked target documents. In some cases, generated questions ended up being better queries than extractive keyword queries generated by the same model. We additionally propose MMR-RBO, a variation of Maximal Marginal Relevance (MMR) using Rank-Biased Overlap (RBO) to identify which questions will perform competitively with the keyword baseline. As all question queries yield unique result sets, we contend that there are no stupid questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。