arXiv:2411.19106cs.CV2024-11被引 3

让AI描述物体时可自由控制细节维度,更精准满足用户需求。

Detailed Object Description with Controllable Dimensions

  • 通过提取、擦除、补充三步分解描述,聚焦用户指定维度
  • 在多个MLLM上提升物体细节描述质量,效果稳定
  • 无需训练,适用于视障人士等需要精准描述的场景

物体描述对视障人群理解与比较物体至关重要。近期多模态大模型虽具强大感知能力,但生成的描述常包含无关内容或遗漏关键维度细节。特定场景下,用户仅需关注某些维度。本文提出无需训练的描述优化流程Dimension Tailor,包含维度提取、擦除与补充三步,将描述分解为用户指定维度。该方法不仅能提升物体细节质量,还可根据用户偏好灵活增减维度。大量实验表明,该流程能持续提升当前主流多模态大模型在可控物体描述上的表现。代码已开源:https://github.com/xin-ran-w/ControllableObjectDescription。

原文摘要 · Abstract (English)

Object description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models(MLLMs) exhibit powerful perceptual abilities and demonstrate impressive potential for generating object-centric descriptions. However, the descriptions generated by such models may still usually contain a lot of content that is not relevant to the user intent or miss some important object dimension details. Under special scenarios, users may only need the details of certain dimensions of an object. In this paper, we propose a training-free object description refinement pipeline, Dimension Tailor, designed to enhance user-specified details in object descriptions. This pipeline includes three steps: dimension extracting, erasing, and supplementing, which decompose the description into user-specified dimensions. Dimension Tailor can not only improve the quality of object details but also offer flexibility in including or excluding specific dimensions based on user preferences. We conducted extensive experiments to demonstrate the effectiveness of Dimension Tailor on controllable object descriptions. Notably, the proposed pipeline can consistently improve the performance of the recent MLLMs. The code is currently accessible at https://github.com/xin-ran-w/ControllableObjectDescription.

物体描述多模态可控生成视障辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。