arXiv:2510.15087cs.IRcs.AI2025-10ACL被引 2

专为灾害管理设计的文本检索模型,提升灾情搜索准确率。

DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management

  • 分三阶段训练:双向注意力适配+无监督对比预训练+难度感知指令微调
  • 六类灾情搜索任务全超基准,596M模型胜过13.3倍大的基线
  • 小模型仅用7.6%参数就超越大模型,适合资源受限场景

灾害管理中高效获取相关信息至关重要。然而,当前缺乏针对灾害管理领域的专用检索模型,通用模型难以应对灾害场景中的多样化查询意图,导致性能不一致且不可靠。为此,我们提出DMRetriever,首个面向该领域的密集检索模型系列(33M至7.6B参数)。其通过创新的三阶段框架训练:双向注意力适配、无监督对比预训练及难度感知渐进式指令微调,并利用先进数据清洗流水线生成高质量数据。全面实验表明,DMRetriever在所有六类搜索意图下均达到最优性能,且在各规模模型上表现领先。此外,模型极具参数效率:596M模型性能超过13.3倍大的基线,33M模型仅用7.6%参数即超越基线。代码、数据与检查点已开源。

原文摘要 · Abstract (English)

Effective and efficient access to relevant information is essential for disaster management. However, no retrieval model is specialized for disaster management, and existing general-domain models fail to handle the varied search intents inherent to disaster management scenarios, resulting in inconsistent and unreliable performance. To this end, we introduce DMRetriever, the first series of dense retrieval models (33M to 7.6B) tailored for this domain. It is trained through a novel three-stage framework of bidirectional attention adaptation, unsupervised contrastive pre-training, and difficulty-aware progressive instruction fine-tuning, using high-quality data generated through an advanced data refinement pipeline. Comprehensive experiments demonstrate that DMRetriever achieves state-of-the-art (SOTA) performance across all six search intents at every model scale. Moreover, DMRetriever is highly parameter-efficient, with 596M model outperforming baselines over 13.3 X larger and 33M model exceeding baselines with only 7.6% of their parameters. All codes, data, and checkpoints are available at https://github.com/KaiYin97/DMRETRIEVER

信息检索灾害管理稠密检索参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。