引入物理力信息可显著提升人体动作理解的准确性与鲁棒性。
Beyond Motion Pattern: An Empirical Study of Physical Forces for Human Motion Understanding
- 将生物力学中的关节力信息融入主流动作理解模型中。
- 在多个数据集上实现显著性能提升,如步态识别最高增益达3.0%。
- 尤其在遮挡、视角变化等复杂条件下效果更优,适合做鲁棒性研究。
人体动作理解虽在视觉识别、跟踪和描述方面取得进展,但多数方法忽视了生物力学中基础的关节作用力等物理线索。本研究系统探讨了物理推断力是否以及何时能增强动作理解:通过将力信息融入现有框架,在三个主要任务(步态识别、动作识别、细粒度视频描述)上评估其影响。在8个基准测试中,加入力信息均带来一致性能提升。例如,在CASIA-B数据集上,步态识别准确率从89.52%提升至90.39%(+0.87),在穿外套或侧视条件下分别提升+2.7%和+3.0%;在Gait3D上从46.0%升至47.3%(+1.3)。动作识别中,CTR-GCN在Penn Action上提升+2.00%,高用力动作如击打/拍打类提升+6.96%。视频描述方面,Qwen2.5-VL的ROUGE-L得分从0.310增至0.339(+0.029),表明力信息增强了时间定位与语义丰富性。结果表明,力线索在动态、遮挡或外观变化条件下能有效补充视觉与运动学特征。
原文摘要 · Abstract (English)
Human motion understanding has advanced rapidly through vision-based progress in recognition, tracking, and captioning. However, most existing methods overlook physical cues such as joint actuation forces that are fundamental in biomechanics. This gap motivates our study: if and when do physically inferred forces enhance motion understanding? By incorporating forces into established motion understanding pipelines, we systematically evaluate their impact across baseline models on 3 major tasks: gait recognition, action recognition, and fine-grained video captioning. Across 8 benchmarks, incorporating forces yields consistent performance gains; for example, on CASIA-B, Rank-1 gait recognition accuracy improved from 89.52% to 90.39% (+0.87), with larger gain observed under challenging conditions: +2.7% when wearing a coat and +3.0% at the side view. On Gait3D, performance also increases from 46.0% to 47.3% (+1.3). In action recognition, CTR-GCN achieved +2.00% on Penn Action, while high-exertion classes like punching/slapping improved by +6.96%. Even in video captioning, Qwen2.5-VL's ROUGE-L score rose from 0.310 to 0.339 (+0.029), indicating that physics-inferred forces enhance temporal grounding and semantic richness. These results demonstrate that force cues can substantially complement visual and kinematic features under dynamic, occluded, or appearance-varying conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。