针对人像描述中空间定位错误问题,提出新评估基准与训练方法。
Human-Centric Image Captioning with Subject-Centered Spatial Understanding

- 通过人体部位精确定位生成空间提示,指导分阶段重写图像描述。
- 在新基准上显著提升模型对主体空间关系的准确性,尤其改善左右侧定位。
- 适合需要精准人体动作与姿态理解的应用场景,如虚拟角色生成。
尽管多模态大语言模型在通用图像描述任务上表现优异,但在以人类为中心的场景中常出现结构幻觉。准确建模人类主体是实现精确虚拟角色/视频/图像生成及细粒度人体动作理解的关键,但这些任务要求高度精确的主体中心空间定位,例如区分第一视角左右方向、保持身体部位与物体的正确绑定关系。然而,这些局部空间错位虽严重影响结构完整性,却常被现有评估指标所掩盖。为此,我们提出SPACE(Subject-centric Poses, Appearance, and Characteristics Evaluation)基准,系统化暴露并量化该瓶颈。在SPACE上发现,当前MLLMs虽具备强大泛化感知能力,却普遍无法在主体内在参考系中正确锚定描述。为填补这一差距,我们设计了专用数据构建与对齐流程:首先从细粒度肢体部位定位中提取结构化空间提示,引导两阶段描述重写,生成高空间保真度训练数据;进一步设计基于评分标准的奖励机制,用于组相对策略优化(GRPO),显式惩罚关键空间错误。大量实验表明,该框架显著提升人像描述质量,尤其在主体中心空间推理方面,性能媲美强闭源模型。基准与代码已开源。
原文摘要 · Abstract (English)
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, they frequently suffer from structural hallucinations in human-centric scenarios. Accurately modeling human subjects is foundational for critical downstream applications, such as accurate avatar/video/image generation and fine-grained human action understanding. However, these tasks require highly precise subject-centered spatial grounding, such as distinguishing egocentric left/right laterality and maintaining correct anatomical-object bindings. Although catastrophic for structural integrity, these localized spatial inversions are often overshadowed by overall descriptive metrics in existing benchmarks. To systematically expose and quantify this bottleneck, we introduce SPACE (Subject-centric Poses, Appearance, and Characteristics Evaluation), a benchmark designed to evaluate subject-centered spatial understanding. On SPACE, we reveal that despite strong generic perception, current MLLMs consistently fail to ground descriptions in the subject's intrinsic frame of reference. To bridge this gap, we propose a specialized data construction and alignment pipeline. We first extract structured spatial hints from fine-grained body-part localization to guide a two-stage caption rewriting process, yielding highly spatially-faithful training data. Furthermore, we design a rubric-based reward for Group Relative Policy Optimization (GRPO) that explicitly penalizes structurally critical spatial errors during alignment. Extensive experiments on SPACE demonstrate our framework significantly improves human-centric caption quality, particularly in subject-centered spatial reasoning, achieving performance competitive with strong closed-source models. Our benchmark and code are available at https://github.com/JHang2020/SPACE-Eval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。