让大模型精准理解图像关键点的语义与位置
KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
- 先识别关键点语义,再用思维链精确定位
- 在多个基准上超越现有方法,提升关键点理解能力
- 适合需要细粒度视觉理解的研究与应用
多模态大模型在图像理解方面取得显著进展,但对像素级语义细节(如物体关键点)仍难以把握。为此,我们提出「语义关键点理解」新挑战,涵盖关键点语义理解、视觉提示与文本提示下的关键点检测。我们提出KptLLM,一种统一的多模态模型,采用‘识别-定位’策略:先解析关键点语义,再通过思维链逐步精确定位。其多个精心设计模块可处理多种模态输入,有效融合语义内容与空间位置信息。大量实验表明,KptLLM在多个关键点检测基准上表现卓越,且具备独特的语义理解能力。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have greatly improved their abilities in image understanding. However, these models often struggle with grasping pixel-level semantic details, e.g., the keypoints of an object. To bridge this gap, we introduce the novel challenge of Semantic Keypoint Comprehension, which aims to comprehend keypoints across different task scenarios, including keypoint semantic understanding, visual prompt-based keypoint detection, and textual prompt-based keypoint detection. Moreover, we introduce KptLLM, a unified multimodal model that utilizes an identify-then-detect strategy to effectively address these challenges. KptLLM underscores the initial discernment of semantics in keypoints, followed by the precise determination of their positions through a chain-of-thought process. With several carefully designed modules, KptLLM adeptly handles various modality inputs, facilitating the interpretation of both semantic contents and keypoint locations. Our extensive experiments demonstrate KptLLM's superiority in various keypoint detection benchmarks and its unique semantic capabilities in interpreting keypoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。