用人体姿态动态信息提升动作描述精度,减少错误生成。
Towards Fine-Grained Human Motion Video Captioning
- 引入人体网格恢复的运动表征,显式建模身体动态
- 在11.5万数据上实现更准确的动作细节与时间变化描述
- 适合关注动作理解与视频描述的科研人员
生成视频中人类动作的准确描述仍是视频字幕模型的挑战。现有方法常难以捕捉细微动作细节,导致描述模糊或语义不一致。本文提出运动增强字幕模型(M-ACM),通过人体网格恢复获取的运动表征,显式强化人体动态建模,减少幻觉,提升生成字幕的语义保真度与空间对齐性。为支持该领域研究,我们构建了包含11.5万视频-描述对的Human Motion Insight(HMI)数据集,并推出专用于评估运动感知字幕的HMI-Bench基准。实验表明,M-ACM显著优于以往方法,在复杂人体动作与微小时间变化描述上表现更优,树立了以运动为核心的视频字幕新标准。
原文摘要 · Abstract (English)
Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motion details, resulting in vague or semantically inconsistent captions. In this work, we introduce the Motion-Augmented Caption Model (M-ACM), a novel generative framework that enhances caption quality by incorporating motion-aware decoding. At its core, M-ACM leverages motion representations derived from human mesh recovery to explicitly highlight human body dynamics, thereby reducing hallucinations and improving both semantic fidelity and spatial alignment in the generated captions. To support research in this area, we present the Human Motion Insight (HMI) Dataset, comprising 115K video-description pairs focused on human movement, along with HMI-Bench, a dedicated benchmark for evaluating motion-focused video captioning. Experimental results demonstrate that M-ACM significantly outperforms previous methods in accurately describing complex human motions and subtle temporal variations, setting a new standard for motion-centric video captioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。