评测多模态模型在中英双语专业文档中的推理能力,填补语言和领域空白。
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

- 构建包含1000道人工标注题的中英双语专业文档推理基准
- 最强模型在该基准上仍存在明显性能差距
- 首次系统评估幻觉检测方法在真实专业场景中的可靠性
尽管多模态大语言模型(MLLMs)在视觉理解方面取得显著进展,但其对文本密集型专业文档的推理能力仍缺乏充分评估。现有基准多聚焦信息提取,依赖外部知识,或仅将专业文档作为众多场景之一,且主要以英语或中文为中心,导致其他语言特别是俄语严重缺失。为解决这些问题,我们提出BEAR-Bench(双语企业与学术推理基准),一个自包含、复杂的中英双语基准,包含1000道基于文本丰富的商业与科学文档的人工标注问题。我们评估了16个专有及开源权重的MLLMs,包括Gemini 3.1 Pro和Qwen3.5-397B,发现即使最强系统在该基准上仍存在明显性能差距。最后,利用模型输出对比现有幻觉检测方法,不仅评估模型在BEAR-Bench上的错误频率,还检验这些错误被识别的可靠性。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。