arXiv:2509.14142cs.CV2025-09ICCV被引 4

聚焦真实场景的多模态推理挑战,发布两个新数据集并开放竞赛结果。

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

  • 构建真实世界与广告视频两类多模态推理数据集
  • 覆盖12种日常场景,评估40+模型在三类任务中的表现
  • 适合关注多模态应用落地的研究者和工业界开发者

本文回顾了2025年多模态推理挑战赛MARS2。我们通过大型基准测试整合多模态机器学习与大语言模型的不同方法,助力研究者跟踪该快速发展的领域。今年挑战赛聚焦真实世界与专用场景,拓展多模态大语言模型的应用范围。组委会发布了两个定制数据集Lens(支持12种日常场景的一般推理)和AdsQA(支持广告视频的领域特定推理)。共评估40+基线模型,涵盖通用型与任务专用模型,并设立三个竞赛赛道:真实场景视觉定位(VG-RS)、具空间意识的视觉问答(VQA-SA)以及创意广告视频中的视觉推理(VR-Ads)。来自知名学术与工业机构的76支团队注册,1200+提交中筛选出40+有效参赛作品进入排名。所有数据集、代码(40+基线及15+参赛方法)与排名均公开于MARS2研讨会官网及GitHub仓库 https://github.com/mars2workshop/,后续更新与活动预告将持续发布。

原文摘要 · Abstract (English)

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the state-of-the-art in this very dynamic area. Meanwhile, a growing number of testbeds have boosted the evolution of general-purpose large language models. Thus, this year's MARS2 focuses on real-world and specialized scenarios to broaden the multimodal reasoning applications of MLLMs. Our organizing team released two tailored datasets Lens and AdsQA as test sets, which support general reasoning in 12 daily scenarios and domain-specific reasoning in advertisement videos, respectively. We evaluated 40+ baselines that include both generalist MLLMs and task-specific models, and opened up three competition tracks, i.e., Visual Grounding in Real-world Scenarios (VG-RS), Visual Question Answering with Spatial Awareness (VQA-SA), and Visual Reasoning in Creative Advertisement Videos (VR-Ads). Finally, 76 teams from the renowned academic and industrial institutions have registered and 40+ valid submissions (out of 1200+) have been included in our ranking lists. Our datasets, code sets (40+ baselines and 15+ participants' methods), and rankings are publicly available on the MARS2 workshop website and our GitHub organization page https://github.com/mars2workshop/, where our updates and announcements of upcoming events will be continuously provided.

多模态推理大模型评测竞赛数据集视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。