首个专为灾害管理设计的检索评估基准,助力精准救灾决策。
DisastIR: A Comprehensive Information Retrieval Benchmark for Disaster Management
- 构建涵盖48项任务的灾害管理信息检索基准
- 测试30个主流模型,发现无一模型在所有任务上均优
- 适合灾害应急、AI救援系统研发人员参考
高效灾害管理依赖及时获取准确且情境相关的资讯。现有信息检索(IR)基准多聚焦通用或特定领域(如医疗、金融),忽视灾害场景中独特的语言复杂性与多元信息需求。为此,我们提出首个专为灾害管理设计的综合信息检索评估基准——DisastIR。该基准包含9,600条多样化用户查询及超过130万条标注的查询-段落对,覆盖由六种搜索意图和八个通用灾害类别衍生出的48项检索任务,涵盖301种具体事件类型。对30个先进检索模型的评估显示,各任务间性能差异显著,无模型能全面领先。对比分析进一步揭示通用领域与灾害专用任务间存在显著性能差距,凸显构建灾害管理专用基准对指导模型选择、支持灾害决策的重要性。所有源码与DisastIR数据集已公开于https://github.com/KaiYin97/Disaster_IR。
原文摘要 · Abstract (English)
Effective disaster management requires timely access to accurate and contextually relevant information. Existing Information Retrieval (IR) benchmarks, however, focus primarily on general or specialized domains, such as medicine or finance, neglecting the unique linguistic complexity and diverse information needs encountered in disaster management scenarios. To bridge this gap, we introduce DisastIR, the first comprehensive IR evaluation benchmark specifically tailored for disaster management. DisastIR comprises 9,600 diverse user queries and more than 1.3 million labeled query-passage pairs, covering 48 distinct retrieval tasks derived from six search intents and eight general disaster categories that include 301 specific event types. Our evaluations of 30 state-of-the-art retrieval models demonstrate significant performance variances across tasks, with no single model excelling universally. Furthermore, comparative analyses reveal significant performance gaps between general-domain and disaster management-specific tasks, highlighting the necessity of disaster management-specific benchmarks for guiding IR model selection to support effective decision-making in disaster management scenarios. All source codes and DisastIR are available at https://github.com/KaiYin97/Disaster_IR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。