无需临床标签,通过物理仿真学习跌倒风险表征
Bridging the Visual-to-Physical Gap: Physically Aligned Representations for Fall Risk Analysis
- 用视频与仿真对齐,学习受物理规律约束的运动表征
- 在4个公开数据集上提升风险表征质量,跌倒检测性能不变
- 零样本自动生成严重程度排序,可解释性强
基于视觉的跌倒分析进展迅速,但核心瓶颈在于:外观相似的动作可能对应截然不同的物理后果,因接触力学和保护性反应差异难以仅从视觉推断。现有方法依赖可靠伤情标签进行监督预测,但真实伤情事件稀少且无法安全模拟,视频证据常受遮挡和视角限制,导致标注噪声大。本文提出PHARL(物理感知对齐表示学习),在无需临床结果标签的情况下,学习具有物理意义的跌倒表征。PHARL通过两个互补约束正则化运动嵌入:(1) 轨迹级时间一致性,保证表征稳定;(2) 多类别物理对齐,利用仿真生成的接触结果塑造嵌入几何结构。通过将视频片段与时间对齐的仿真描述符配对,PHARL捕捉局部影响相关动力学,同时保持纯前馈推理。在四个公开数据集上的实验表明,相比纯视觉基线,PHARL持续提升风险对齐表征质量,同时保持强跌倒检测性能。值得注意的是,PHARL展现出零样本序数性:无需显式序数监督,即可自发形成可解释的严重程度结构(头部 > 躯干 > 支持状态)。
原文摘要 · Abstract (English)
Vision-based fall analysis has advanced rapidly, but a key bottleneck remains: visually similarmotions can correspond to very different physical outcomes because small differences in contactmechanics and protective responses are hard to infer from appearance alone. Most existingapproaches handle this by supervised injury prediction, which depends on reliable injury labels.In practice, such labels are difficult to obtain: video evidence is often ambiguous (occlusion,viewpoint limits), and true injury events are rare and cannot be safely staged, leading to noisysupervision. We address this problem with PHARL (PHysics-aware Alignment RepresentationLearning), which learns physically meaningful fall representations without requiring clinicaloutcome labels. PHARL regularizes motion embeddings with two complementary constraints:(1) trajectory-level temporal consistency for stable representation learning, and (2) multi-classphysics alignment, where simulation-derived contact outcomes shape embedding geometry. Bypairing video windows with temporally aligned simulation descriptors, PHARL captures localimpact-relevant dynamics while keeping inference purely feed-forward. Experiments on fourpublic datasets show that PHARL consistently improves risk-aligned representation quality overvisual-only baselines while maintaining strong fall-detection performance. Notably, PHARL alsoexhibits zero-shot ordinality: an interpretable severity structure (Head > Trunk > Supported)emerges without explicit ordinal supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。