评测大模型在图文可验证问题上的推理能力,发布新基准ROME。
FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions
- 构建图文可验证的推理评测基准ROME,避免数据污染。
- 发现当前大模型在视觉线索推理上表现有限,尤其面对复杂场景时。
- 适合关注多模态推理评估的开发者与研究者使用。
我们对当前大型推理模型(LRMs)进行了中等规模、一定程度上无污染的评估,得出一些初步发现。同时发布了 ROME——一个面向视觉语言模型的评测基准,旨在测试模型从视觉线索中进行推理的能力。相关基准、评测数据及其他更新已发布于官网:https://flageval-baai.github.io/LRM-Eval/。
原文摘要 · Abstract (English)
We conduct a moderate-scale contamination-free (to some extent) evaluation of current large reasoning models (LRMs) with some preliminary findings. We also release ROME, our evaluation benchmark for vision language models intended to test reasoning from visual clues. We attach links to the benchmark, evaluation data, and other updates on this website: https://flageval-baai.github.io/LRM-Eval/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。