让AI理解物体方向更贴近人视角,提升多模态模型的感知准确性。
Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning
- 基于用户第一视角构建统一标注标准,优化模型方向理解。
- 在跨域图像上测试,方向识别准确率显著提升,且不影响模型整体性能。
- 适合关注视觉-语言对齐与真实场景交互的研究者使用。
多模态大语言模型(MLLM)在多模态应用中充当人与AI交互的重要接口。然而,当前MLLM在图像中物体方向的理解上存在偏差,源于训练数据中方向标注不一致,阻碍了连贯的方向认知发展。为此,本文提出第一人称指令微调方法,基于用户第一人称视角建立一致的标注标准,使模型方向理解与用户视角对齐。首先利用MLLM识别物体细节并结合先验知识生成第一人称指令数据;再通过指令微调增强模型的方向解析能力。此外,我们构建EgoOrientBench基准,涵盖三个任务,使用来自多样领域的图像进行评估。实验结果表明,该方法显著提升方向理解能力,且不损害模型整体性能。指令数据与基准数据集已开源,详见项目页 https://github.com/jhCOR/EgoOrientBench。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) act as essential interfaces, connecting humans with AI technologies in multimodal applications. However, current MLLMs face challenges in accurately interpreting object orientation in images due to inconsistent orientation annotations in training data, hindering the development of a coherent orientation understanding. To overcome this, we propose egocentric instruction tuning, which aligns MLLMs' orientation understanding with the user's perspective, based on a consistent annotation standard derived from the user's egocentric viewpoint. We first generate egocentric instruction data that leverages MLLMs' ability to recognize object details and applies prior knowledge for orientation understanding. Using this data, we perform instruction tuning to enhance the model's capability for accurate orientation interpretation. In addition, we introduce EgoOrientBench, a benchmark that evaluates MLLMs' orientation understanding across three tasks using images collected from diverse domains. Experimental results on this benchmark show that egocentric instruction tuning significantly improves orientation understanding without compromising overall MLLM performance. The instruction data and benchmark dataset are available on our project page at https://github.com/jhCOR/EgoOrientBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。