用大模型从人道主义报告中提取因果证据并验证其可靠性
Causal Evidence Extraction and Triangulation in Crisis Reports using Large Language Models: A ReliefWeb-based Study

- 分两阶段提取干预与结果的结构化因果关系,带方向和强度属性
- 在100份报告上达到94.15%加权F1,现金援助效果呈现强正向共识
- 通过跨场景证据聚合生成可信度评分,适合政策评估与应急决策
人道主义报告冗长、嘈杂且多主题,难以整合关键因果证据。我们基于ReliefWeb(2000–2024)开展研究,提出一个两阶段大语言模型(LLM)流水线,可提取带有方向与强度属性的干预-结果记录。查询条件抽取限制输出至特定干预类别,减少过量提取;片段溯源将每条关系关联原文段落,保障可审计性与分类准确性。在100份专家标注报告上,最佳闭源模型达90.73%加权F1,成本高效;经监督微调的Llama-3.1-8B模型达94.15%加权F1。我们进一步提出保持上下文的三角验证方法:在灾害×来源单元内聚合加权证据,采用拉普拉斯平滑并等权重单元,以‘证据等级’量化跨情境一致性。应用于现金援助时,食物相关结果呈现强正向收敛(LoE=0.865),且长期趋势稳定。
原文摘要 · Abstract (English)
Humanitarian reports are long, noisy, and multi-topic, making it difficult to consolidate decision-relevant causal evidence. We present a ReliefWeb study (2000-2024) and a two-stage Large Language Model (LLM) pipeline that extracts structured intervention-outcome records with direction and strength attributes. Query-conditioned extraction restricts output to a specified intervention class, reducing retrieval-induced over-extraction, while snippet grounding links each relation to supporting text for auditability and classification. In an expert-annotated dataset of 100 reports, the best closed-source LLM achieved a weighted F1 score of 90.73% with strong cost-efficiency, while Llama-3.1-8B with supervised fine-tuning reached 94.15% weighted F1 score. We further propose context-preserving triangulation that aggregates strength-weighted evidence within disaster$\times$source cells, applies Laplace smoothing and equally weights cells to quantify cross-context convergence via a Level-of-Evidence score. Applied to cash assistance, food-related outcomes show strong positive convergence (LoE=0.865) and stable long-horizon trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。