arXiv:2608.27481cs.CLcs.AI2026-08

构建跨语言多跳问答基准,揭示语言边界对知识推理的干扰。

XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

论文配图:XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering
图 1 · 摘自论文原文
  • 设计含语言标签的证据依赖图,模拟跨语言推理链
  • 验证集99.8%问题与答案在不同语言间传递,95.6%使用异语证据
  • 暴露模型在跨语言匹配和不同脚本下的显著性能下降

知识密集型多跳问答要求系统选择证据并组合相关事实,但现有多语言基准通常将整例翻译为单一语言,掩盖了推理链中语言边界处的失败。我们提出XHotpotQA,一个针对混合语言证据上跨语言知识组合的受控基准。每个实例建模为证据依赖图,其问题、桥接证据、答案承载证据及干扰项均明确标注语言。该审计资源包含15,661个训练与7,405个验证实例,提供句级支持监督与已知干扰项。验证集中99.81%的样本跨越问题到黄金证据的语言界面,95.60%使用不同语言的黄金段落。在三个阅读器模型中,完全不匹配问题-证据语言导致Unicode感知答案F1降低10.25至15.79点,不同脚本文本导致11.98至23.70点下降;相应适配选择器对比仅差1.71与1.78点。在该候选供给设计下,评估阅读器的条件相关缺陷远大于选择器。XHotpotQA提供角色感知诊断、模块化评估与审计测试环境,适用于需跨语言整合证据的知识系统。

原文摘要 · Abstract (English)

Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries inside the reasoning chain. We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence. Each instance is modeled as an evidence-dependency graph whose question, bridge evidence, answer-bearing evidence, and distractors have explicit language assignments. The audited resource contains 15,661 training and 7,405 validation instances, with sentence-level support supervision and supplied distractors. In validation, 99.81% of items cross the question-to-gold-evidence language interface and 95.60% use gold paragraphs in different languages. Across three reader artifacts, full question-evidence mismatch is associated with 10.25 to 15.79 lower Unicode-aware answer F1 than partial alignment, and different-script evidence with deficits of 11.98 to 23.70 points; the corresponding adapted-selector contrasts are 1.71 and 1.78 points. Under this supplied-candidate design, the evaluated readers therefore show substantially larger condition-associated deficits than the selector. XHotpotQA provides role-aware diagnostics, modular evaluation, and an audited test bed for knowledge-based systems that must integrate evidence across languages.

多跳问答跨语言知识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。