提升日文文档多选题理解能力,解决模型偏见与复杂布局难题
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- 构建分层视觉语言推理框架,结合子问题分解增强语义解析
- 在十选一测试中性能显著优于主流模型,尤其在日文场景下提升明显
- 适合需要处理复杂排版日文文档的多模态问答应用
多模态大语言模型(MLLM)在视觉问答任务中展现出强大的多模态理解能力,但面对十选一选择题评估范式时,现有方法在处理具有复杂布局和长篇内容的PDF文档时仍存在显著局限。当前主流模型对英语训练数据存在强依赖性,导致日语等其他语言场景表现不佳。为此,本文提出一种新型日文PDF文档理解框架,融合多模态分层推理机制与Colqwen优化的检索方法,并创新性引入通过子问题分解实现的语义验证策略。实验表明,该框架不仅显著提升了模型对复杂文档的深层语义解析能力,还在实际应用中表现出更强的鲁棒性。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal understanding capabilities in Visual Question Answering (VQA) tasks by integrating visual and textual features. However, under the challenging ten-choice question evaluation paradigm, existing methods still exhibit significant limitations when processing PDF documents with complex layouts and lengthy content. Notably, current mainstream models suffer from a strong bias toward English training data, resulting in suboptimal performance for Japanese and other language scenarios. To address these challenges, this paper proposes a novel Japanese PDF document understanding framework that combines multimodal hierarchical reasoning mechanisms with Colqwen-optimized retrieval methods, while innovatively introducing a semantic verification strategy through sub-question decomposition. Experimental results demonstrate that our framework not only significantly enhances the model's deep semantic parsing capability for complex documents, but also exhibits superior robustness in practical application scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。