arXiv:2606.08034cs.CVcs.AI2026-06

构建多语言视觉融合的STEM问题评测集,揭示大模型跨语言鲁棒性差异。

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems

论文配图:Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems
图 1 · 摘自论文原文
  • 设计5学科7语言共4242个可生成变体的问题模板,每题含执行代码与解题步骤。
  • 17个视觉语言模型在最坏情况下的准确率显著低于平均准确率,差距明显。
  • 小模型跨语言性能下降严重,大模型和闭源模型更稳定,适合评估鲁棒性。

符号化评测基准已成为评估模型在微小修改下对科学、技术、工程、数学问题鲁棒性的重要方法。然而,现有基准大多局限于数学推理,缺乏视觉信息支持,且以英语为主。本文提出Sci-Rho(Science Rhobustness),一个动态的多语言、视觉融合式STEM评测集,涵盖五个学科和七种语言,包含4,242个由领域专家(包括奥赛奖牌得主)设计的问题模板(每语言606个)。每个模板为可执行的Python代码,通过改变数值、视觉模式、几何形状、颜色方案和函数类型,生成42,420个等价问题实例,每个实例配有推理步骤与真值答案。我们评估了17个前沿视觉语言模型,发现最坏情况准确率(即模型在所有变体中均正确回答的问题模板比例)显著低于平均准确率。小模型在跨语言任务中表现明显下滑,而大型及专有模型则保持稳健。步骤级评估也显示平均F1与最坏情况F1之间存在显著差距。进一步分析表明,模型注意力头在不同语言间对图像与文本标记的关注度差异显著。本工作强调需超越静态评测集,以动态多变评测作为衡量视觉语言模型质量的关键指标。

原文摘要 · Abstract (English)

Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions. However, existing symbolic benchmarks mostly remain limited to mathematical reasoning, lack visual grounding, and are predominantly in English. In this work, we introduce Sci-Rho (Science Rhobustness), a dynamic benchmark for visually-grounded STEM problems spanning five subjects and seven languages, comprising 4,242 problem templates (606 per language) crafted by domain experts, including Olympiad medalists. Each template is implemented as executable Python code that generates diverse but equivalent problem instances by varying numerical values, visual patterns, geometric shapes, color schemes, and function types, resulting in 42,420 instances in total, each paired with reasoning steps and ground-truth solutions. We evaluated 17 state-of-the-art VLMs and discovered a noticeable gap between worst-case accuracy (defined as the proportion of problem templates that a model answers correctly across every generated variation) and average accuracy. We also discovered that smaller models show noticeable performance degradation across languages, whereas proprietary and larger models remain robust. Step-level evaluation reflects this same trend, revealing a significant gap between average F1 and worst-case F1 scores. Finally, our inspection of attention heads of a VLM reveals substantial cross-lingual variation in the relative attention allocated to image tokens compared to text tokens. Our work highlights the importance of evaluation beyond static benchmarks as a metric to measure the quality of VLMs.

视觉语言模型多语言评测鲁棒性评估STEM教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。