跨语言检索挑战大,研究智能体处理外语证据能力明显下降
Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus
- 构建双语评测基准XBCP,测试智能体在多语言证据下的表现
- 即使提供全部正确证据,跨语言任务准确率仍显著降低
- 发现检索与推理双瓶颈,适合关注多语言AI系统的开发者
深度研究智能体的评估通常依赖于用户查询与证据同语种的假设。本文提出XBCP(Cross-lingual BrowseComp-Plus)基准,保留原英文问答空间,但将支持性文档设置为不同语言。该基准包含两种场景:跨语言场景下每题对应单一指定语言的证据;多语言场景下12种语言(涵盖高/低资源语言)等量分布全语料。评估四类智能体搭配稀疏与密集多语言检索器,测量准确率、证据召回率、搜索行为、校准度、引用可信度及理想检索表现。结果显示,当证据被翻译后,模型性能大幅下降,强稠密检索器也出现证据召回率下滑,智能体校准度降低且引用可靠性变差。值得注意的是,即便直接提供所有黄金证据,准确率仍低于同语种水平。这表明跨语言深度研究不仅存在检索失败,还暴露出智能体自身整合语言不匹配信息的独立困难。
原文摘要 · Abstract (English)
Deep research agents are increasingly evaluated on their ability to search for evidence, reason over retrieved sources, and produce grounded answers. Existing browsing benchmarks, however, largely assume that the user's query and the supporting evidence are written in the same language, leaving open whether agentic search systems can operate when relevant evidence appears in another language. We introduce XBCP (Cross-lingual BrowseComp-Plus), a controlled benchmark that preserves the English question-and-answer space of BrowseComp-Plus but varies the languages of the supporting documents. XBCP instantiates two complementary settings: in the cross-lingual setting, each query is paired with evidence in a single assigned language. In the multilingual setting, the full evidence corpus is distributed equally and randomly across 12 languages spanning high-resource and low-resource regimes. We evaluate four deep research agents using sparse and dense multilingual retrievers, measuring answer accuracy, evidence recall, search behavior, calibration, citation fidelity, and oracle retrieval. Results reveal substantial degradation when evidence is translated. Even strong, dense retrievers lose evidence recall, and agents become less calibrated and cite evidence less reliably. Notably, accuracy remains lower even when all gold evidence is supplied directly. These findings suggest that cross-lingual deep research exposes both retrieval failures and an independent, agent-side difficulty in integrating language-mismatched evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。