arXiv:2504.04781cs.CV2025-04中稿 · the Multimodal Alg…

用3D监督和思维链引导,提升大模型对遮挡物体的识别能力。

OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance

  • 融合3D重建专家与多模态大模型,实现遮挡物体理解。
  • 构建11万样本的遮挡物体思维链数据集,提升推理能力。
  • 在多个模型上实现最高16.98%的识别性能提升,适合视觉语言研究者。

现有大规模视觉-语言多模态模型对遮挡物体的理解研究不足。当前最先进的多模态大模型依赖通用视觉编码器和监督学习策略,在理解遮挡物体时表现不佳。为此,我们提出OCC-MLLM-CoT-Alpha,一个结合3D感知监督与思维链引导的多模态大视觉语言框架。首先,构建由大型多模态视觉语言模型与3D重建专家模型组成的框架;其次,通过监督与强化学习相结合的方式,学习多模态思维链,使模型能借助思维链指导增强识别能力;最后,构建了一个包含11万条手持遮挡物体样本的大规模多模态思维链推理数据集。实验表明,该方法在多种先进模型的不同设置下,决策得分分别提升了15.75%、15.30%、16.98%、14.62%以及4.42%、3.63%、6.94%、10.70%。

原文摘要 · Abstract (English)

Comprehending occluded objects are not well studied in existing large-scale visual-language multi-modal models. Current state-of-the-art multi-modal large models struggles to provide satisfactory results in understanding occluded objects through universal visual encoders and supervised learning strategies. Therefore, we propose OCC-MLLM-CoT-Alpha, a multi-modal large vision language framework that integrates 3D-aware supervision and Chain-of-Thoughts guidance. Particularly, (1) we build a multi-modal large vision-language model framework which is consisted of a large multi-modal vision-language model and a 3D reconstruction expert model. (2) the corresponding multi-modal Chain-of-Thoughts is learned through a combination of supervised and reinforcement training strategies, allowing the multi-modal vision-language model to enhance the recognition ability with learned multi-modal chain-of-thoughts guidance. (3) A large-scale multi-modal chain-of-thoughts reasoning dataset, consisting of $110k$ samples of occluded objects held in hand, is built. In the evaluation, the proposed methods demonstrate decision score improvement of 15.75%,15.30%,16.98%,14.62%, and 4.42%,3.63%,6.94%,10.70% for two settings of a variety of state-of-the-art models.

多模态遮挡识别思维链3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。