arXiv:2504.10808cs.CVcs.HC2025-04被引 1

在保护隐私的前提下,用视频特征统计量实现高精度共情检测。

Privacy-Preserving Empathy Detection in Video Interactions

  • 用表格基础模型处理视频特征摘要,实现强隐私下的共情识别。
  • 跨被试评估下准确率从59.0%提升至73.0%,显著增强泛化能力。
  • 适合注重数据隐私与伦理合规的智能人机交互系统开发。

从视频交互中检测共情的应用日益增多,但因隐私与伦理限制,原始视频难以用于训练AI模型。现有公开基准仅提供预提取特征,形成隐私受限的学习环境,其隐私-效用权衡尚不明确。本文定义了视频行为预测的三个隐私层级:无隐私(原始视频)、部分隐私(如面部关键点、动作单元、眼动等时间视觉特征)和强隐私(这些特征的汇总统计量),并探究在强隐私条件下是否可实现具备主体泛化能力的共情检测。提出TFMPathy框架,基于两种最新表格基础模型(TabPFN v2 和 TabICL),采用上下文学习与微调范式。在公开的人机交互基准上,TFMPathy在强隐私下仍表现优异,显著优于已有基线。为评估鲁棒性并促进公平安全部署,引入此前缺失的跨被试评估协议。在此协议下,TFM微调使准确率从0.590提升至0.730,AUC从0.564升至0.669,显著提升泛化性能。将时间特征聚合为统计量还能抑制个体及人口统计学特征线索,符合数据最小化原则。因此,TFMPathy为在治理、同意或政策限制下使用人类中心视频构建AI系统提供了可行路径。代码将在录用后公开于https://github.com/hasan-rakibul/TFMPathy。

原文摘要 · Abstract (English)

Detecting empathy from video interactions has emerging applications, yet raw videos that could be used for training AI models are rarely available due to privacy and ethical constraints. Public benchmarks are consequently released only as pre-extracted features, creating a privacy-constrained learning regime whose privacy-utility trade-off is poorly characterised. We formalise three levels of privacy for video-based behavioural prediction -- no privacy (raw video), partial privacy (temporal visual features such as facial landmarks, action units and eye gaze) and strong privacy (summary statistics of those features) -- and ask whether strong, subject-generalisable empathy detection is achievable at the strong-privacy level. We propose TFMPathy, instantiated with two recent Tabular Foundation Models (TFMs) (TabPFN v2 and TabICL), under both in-context learning and fine-tuning paradigms. On a public human-robot interaction benchmark, TFMPathy achieves strong utility under strong privacy, outperforming established baselines by a substantial margin. To assess robustness and facilitate fair, safe deployment, we introduce a cross-subject evaluation protocol that was previously lacking in this benchmark. Under this protocol, TFM fine-tuning improves generalisation capacity substantially (accuracy: $0.590 \rightarrow 0.730$; AUC: $0.564 \rightarrow 0.669$). Aggregating temporal features into summary statistics also suppresses subject-specific and demographic cues, aligning TFMPathy with data-minimisation principles. TFMPathy, therefore, offers a practical route to building AI systems that depend on human-centred video when governance, consent or institutional policies restrict the sharing of raw video. Code will be released upon acceptance at https://github.com/hasan-rakibul/TFMPathy.

共情检测隐私保护表格模型人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。