构建非洲经济报告问答基准,挑战模型精准定位数值与时间信息的能力。
AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports
- 从220份世行报告中构建需精确定位数值与时间的问答数据集
- 零样本下最佳模型F1仅0.545,列表与跨期比较仍存显著误差
- 适合研究长文档理解、量化推理与时间建模的AI系统评测
在长篇机构文档上实现可靠的问答,不仅需要主题检索,更需精确定位支持性段落,并在指标反复出现于不同年份、国家和预测周期时保持精确的数值与时间细节。现有问答基准很少在文档规模下测试这种精确锚定与时间精度的结合。我们提出AfriEconQA,一个基于220份世界银行非洲经济报告的文档级问答基准,其内容关联具体财政周期、国家和预测状态。该数据集包含4,309个与证据关联的问答实例,涵盖五类推理:事实型(1,093)、列表型(879)、选择题(944)、综合型(710)和对比型(683),分别针对数值提取、集合恢复、判别、因果整合及跨期或跨国推理。每个实例均附有支持证据与来源溯源,通过代理生成流程构建并经分层人工验证确保标注质量。我们在保留测试集(n=862)上评估Qwen 3.6 35B、DeepSeek v4-pro与Gemma 4 12B IT在零样本、最优提示和混合检索增强生成(RAG)条件下的表现。三模型均显示检索带来显著提升,但最佳RAG系统F1仅为0.545,残余错误集中于列表提取、综合推理和时间限定对比任务。AfriEconQA因此为长篇机构经济报告中的精确数值与时间锚定提出了严峻挑战。
原文摘要 · Abstract (English)
Reliable question answering over long institutional documents requires more than topical retrieval: a system must localize the exact passage that supports a claim and preserve precise numerical and temporal detail when the same indicator recurs across years, countries, and projection horizons in dense, repetitive prose. Existing question-answering benchmarks rarely test this combination of exact grounding and temporal precision at document scale. We introduce AfriEconQA, a benchmark for document-grounded question answering built from 220 World Bank economic reports on African economies, a corpus whose claims are tied to specific fiscal periods, countries, and projection states. AfriEconQA contains 4,309 evidence-linked QA instances across five reasoning categories: Factoid (1,093), List (879), Multiple Choice (944), Synthesis (710), and Comparison (683), targeting quantitative extraction, set recovery, discrimination, causal integration, and cross-period or cross-country reasoning. Each instance carries sup- porting evidence and source provenance and is constructed through an agentic generation pipeline with evidence-grounding checks and a stratified human-validation subset for gold-label auditing. We evaluate Qwen 3.6 35B, DeepSeek v4-pro, and Gemma 4 12B IT under zero-shot, oracle, and hybrid retrieval-augmented generation (RAG) conditions on the held-out test split (n = 862). Across all three models, retrieval yields substantial gains, yet the best RAG system reaches only 0.545 F1, with residual errors concentrated in list extraction, synthesis, and temporally scoped comparison. AfriEconQA therefore poses a hard open challenge for exact numerical and temporal grounding over long institutional economic reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。