跨域评估多模态思维链推理在视觉问答中的表现
Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
- 分两阶段生成理由并融合视觉特征,用T5模型实现多模态推理
- 视觉信息显著降低推理幻觉,但常识类问题仍难应对
- 适合研究多模态推理系统泛化能力的学者参考
尽管近期工作已将思维链(CoT)拓展至多模态场景,并在ScienceQA等科学问答基准上取得领先效果,但这些方法在不同领域的泛化能力仍缺乏深入探索。本文对多模态思维链(Multimodal-CoT)进行系统评估,涵盖A-OKVQA、OKVQA和ChartQA三个数据集,需超越科学推理的广泛常识与世界知识。采用Zhang等人[3]提出的两阶段框架,分离理由生成与答案推断,通过门控融合机制整合视觉特征,并基于T5语言模型实现推理。通过系统的消融实验,分析视觉特征、理由质量及架构选择的影响。结果表明,视觉信息显著减少理由生成中的幻觉,但不同题型下CoT效果差异明显,尤其常识推理面临挑战。本研究为多模态推理系统实现提供实用洞见,并指明跨域泛化改进的关键方向。
原文摘要 · Abstract (English)
While recent work has extended CoT to multimodal settings, achieving state-of-the-art results on science question answering benchmarks like ScienceQA, the generalizability of these approaches across diverse domains remains underexplored. This work presents a comprehensive analysis of Multimodal Chain-of-Thought (Multimodal-CoT) reasoning, evaluating its effectiveness on the A-OKVQA, OKVQA and ChartQA datasets, which requires broad commonsense and world knowledge beyond scientific reasoning. We implement the two-stage framework proposed by Zhang et al. [3], which separates rationale generation from answer inference and integrates vision features through a gated fusion mechanism with T5-based language models. Through systematic ablation studies, we analyze the contributions of vision features, rationale quality, and architectural choices. Our findings reveal that while vision integration significantly reduces hallucination in rationale generation, the effectiveness of CoT reasoning varies substantially across question types, with commonsense reasoning presenting particular challenges. This work provides practical insights for researchers implementing multimodal reasoning systems and identifies key areas for future improvement in cross-domain generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。