让大模型精准理解图像中物体的关键点,提升人机协作能力。
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
- 先理解关键点语义,再通过思维链定位位置。
- 训练数据超50万样本,覆盖多种物体与复杂遮挡场景。
- 在多个基准上达到顶尖性能,适合细粒度图像分析任务。
多模态大语言模型(MLLMs)通过融合文本与视觉模态革新了图像理解。然而,这类模型在捕捉细粒度语义信息(如物体关键点的精确识别与分析)方面仍存在不足。关键点作为结构感知、像素级且紧凑的物体表示,尤其在可动物体中至关重要,广泛应用于细粒度图像分析、物体检索和行为识别。本文提出KptLLM++,一种专为通用关键点理解设计的多模态大模型,通过用户定义指令引导的多模态输入融合实现。该模型采用创新的“先识别后检测”范式,先解析关键点语义,再通过结构化思维链推理精确定位。为突破性能瓶颈,训练数据规模扩展至超过50万样本,涵盖多样物体、关键点类别、图像风格及复杂遮挡场景。大规模训练使模型充分释放潜力,在多个关键点检测基准上展现出卓越准确率与泛化能力,验证其作为细粒度图像理解统一解决方案的可行性,并推动人机协作的变革。
原文摘要 · Abstract (English)
The emergence of Multimodal Large Language Models (MLLMs) has revolutionized image understanding by bridging textual and visual modalities. However, these models often struggle with capturing fine-grained semantic information, such as the precise identification and analysis of object keypoints. Keypoints, as structure-aware, pixel-level, and compact representations of objects, particularly articulated ones, play a crucial role in applications such as fine-grained image analysis, object retrieval, and behavior recognition. In this paper, we propose KptLLM++, a novel multimodal large language model that specifically designed for generic keypoint comprehension through the integration of diverse input modalities guided by user-defined instructions. By unifying keypoint detection across varied contexts, KptLLM++ establishes itself as an advanced interface, fostering more effective human-AI collaboration. The model is built upon a novel identify-then-detect paradigm, which first interprets keypoint semantics and subsequently localizes their precise positions through a structured chain-of-thought reasoning mechanism. To push the boundaries of performance, we have scaled up the training dataset to over 500K samples, encompassing diverse objects, keypoint categories, image styles, and scenarios with complex occlusions. This extensive scaling enables KptLLM++ to unlock its potential, achieving remarkable accuracy and generalization. Comprehensive experiments on multiple keypoint detection benchmarks demonstrate its state-of-the-art performance, underscoring its potential as a unified solution for fine-grained image understanding and its transformative implications for human-AI interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。