arXiv:2607.15400cs.CVcs.LG2026-07

用无监督关键点实现隐私保护的实时跌倒检测,抗遮挡能力更强。

Unsupervised Keypoints for Real-Time Fall Detection: Comparative Analysis Under Real-world Conditions with Predictive Bandwidth Reduction

论文配图:Unsupervised Keypoints for Real-Time Fall Detection: Comparative Analysis Under Real-world Conditions with Predictive Bandwidth Reduction
图 1 · 摘自论文原文
  • 用无监督关键点替代RGB视频,减少带宽并保护隐私。
  • 在遮挡场景下,无监督方法跌倒检出率比有监督高近一倍。
  • 适合身体可见性差的老人监护场景,如居家或光照不良环境。

老年人跌倒是一大安全挑战,但持续监测难以维持。视频可捕捉跌倒时的姿态与动作,但部署受限于隐私、计算和带宽问题。有监督姿态估计具有解剖可解释性,但易受遮挡和部分可见影响。本文提出一种隐私保护框架,将RGB传输替换为基于无监督关键点与预测性时序建模的紧凑运动表征。本地处理完成分割与关键点提取;变分递归预测与序列分类从观测与预测的运动中检测跌倒。在UR跌倒检测和人类跌倒数据集上,采用随机、主体不重叠及遮挡相关划分进行评估。随机划分下,两种表征无明显优劣,提示标准协议可能掩盖真实差异。主体不重叠评估中,有监督关键点表现显著更优,但性能随个体变化:当解剖标志可见时效果更好;而无监督关键点对遮挡和部分可见更具鲁棒性,尽管复杂活动误报更多。遮挡评估下,有监督方法遗漏近半数跌倒,而无监督方法保持强敏感性,显著优于前者。其解剖无关性使空间锚点能自适应可见身体结构,而非因缺失标志失败。带宽受限时,有监督定位误差在时序模型中累积,差距进一步扩大。结果表明,表征选择应反映预期视觉条件,当身体可见性受损时,无监督关键点更具优势。

原文摘要 · Abstract (English)

Falls among older adults are a major safety challenge, but continuous monitoring is difficult to sustain. Video captures fall-related posture and motion, yet deployment is limited by privacy, computation, and bandwidth. Supervised pose estimation is anatomically interpretable but vulnerable to occlusion and partial body visibility. We propose a privacy-preserving framework that replaces RGB transmission with compact motion representations based on unsupervised keypoints and predictive temporal modeling. Local processing performs segmentation and keypoint extraction; variational recurrent prediction and sequence classification then detect falls from observed and forecasted motion. We evaluate the framework on the UR Fall Detection and Human Fall datasets using random, subject-disjoint, and occlusion-based splits. Under random splits, neither representation consistently dominates, suggesting that standard protocols may hide meaningful differences. Under subject-disjoint evaluation, supervised keypoints show a statistically significant advantage, but performance varies by subject: they perform better when anatomical landmarks are visible, whereas unsupervised keypoints are more robust to occlusion and partial visibility, though they produce more false positives for complex activities. Under occlusion-based evaluation, supervised keypoints miss nearly half of all falls, while unsupervised keypoints retain strong sensitivity and substantially outperform them. Their anatomical independence allows spatial anchors to adapt to visible body structure rather than fail on absent landmarks. The gap widens under bandwidth constraints, where supervised localization errors compound through the temporal model. These findings show that representation choice should reflect expected visual conditions and that unsupervised keypoints offer an advantage when body visibility is compromised.

跌倒检测无监督学习隐私保护边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。