arXiv:2507.23404cs.CL2025-07被引 2

针对阿拉伯语复杂性,提出注意力相关性评分提升文本检索准确率

Enhanced Arabic Text Retrieval with Attentive Relevance Scoring

  • 用自适应注意力机制替代传统交互方式,精准捕捉问句与段落语义关联
  • 在阿拉伯语问答数据集上显著提升排名准确率,优于标准DPR模型
  • 适合研究阿拉伯语NLP或信息检索的学者与开发者使用

阿拉伯语因复杂的形态变化、可选的符号标记以及现代标准阿拉伯语与方言并存,在自然语言处理和信息检索中面临特殊挑战。尽管阿拉伯语全球重要性日益提升,但在NLP研究与基准资源中仍被严重低估。本文提出专为阿拉伯语设计的增强型密集段落检索(DPR)框架,核心是新颖的注意力相关性评分(ARS)机制,取代传统交互方式,采用自适应评分函数更有效地建模问题与段落间的语义相关性。该方法结合预训练阿拉伯语语言模型与架构优化,显著提升阿拉伯语问答任务中的检索性能与排名准确性。代码已公开于GitHub。

原文摘要 · Abstract (English)

Arabic poses a particular challenge for natural language processing (NLP) and information retrieval (IR) due to its complex morphology, optional diacritics and the coexistence of Modern Standard Arabic (MSA) and various dialects. Despite the growing global significance of Arabic, it is still underrepresented in NLP research and benchmark resources. In this paper, we present an enhanced Dense Passage Retrieval (DPR) framework developed specifically for Arabic. At the core of our approach is a novel Attentive Relevance Scoring (ARS) that replaces standard interaction mechanisms with an adaptive scoring function that more effectively models the semantic relevance between questions and passages. Our method integrates pre-trained Arabic language models and architectural refinements to improve retrieval performance and significantly increase ranking accuracy when answering Arabic questions. The code is made publicly available at \href{https://github.com/Bekhouche/APR}{GitHub}.

阿拉伯语文本检索注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。