arXiv:2601.17194cs.CV2026-01

通过动作识别心理状态,为建筑设计提供隐私保护的互动评估方法

Decoding Psychological States Through Movement: Inferring Human Kinesic Functions with Application to Built Environments

  • 将埃克曼动作分类体系转化为可量化的行为词汇,用于分析人际互动
  • 在12组双人互动中实现高精度动作功能识别,跨场景通用性好
  • 无需人工标注词典,基于骨骼数据直接推断沟通意图,适合城市设计研究

社交基础设施及其他建成环境日益需要通过促进社会互动来支持福祉与社区韧性。然而,在土木与建成环境研究中,尚无一致且隐私友好的方式来表征和度量这些空间中的社会意义互动,导致研究对“互动”的操作化定义各不相同,限制了从业者评估设计干预是否改变了社会资本理论预测的关键互动形式的能力。为填补这一领域级方法论空白,我们提出了双人用户参与数据集(DUET)及嵌入式动作识别框架,将埃克曼与弗里森的动作分类体系作为与社会资本相关行为(如互惠与注意力协调)对应的功能级互动词汇。DUET涵盖四种传感模态、三个建成环境场景下的12组双人互动,涵盖五类动作功能——象征、描绘、情感表达、调节与调控,支持通过运动实现隐私保护的沟通意图分析。对六种主流开源人体活动识别模型的基准测试表明,通信功能识别难度高,且普遍的单人动作识别模型在扩展至双人社会性互动时存在局限。基于DUET,我们的识别框架直接从隐私保护的骨骼运动中推断沟通功能,无需手工构建动作-功能词典;采用迁移学习架构,揭示了动作功能的结构化聚类,并发现表示质量与分类性能强相关,且在不同被试和场景间具有良好泛化能力。

原文摘要 · Abstract (English)

Social infrastructure and other built environments are increasingly expected to support well-being and community resilience by enabling social interaction. Yet in civil and built-environment research, there is no consistent and privacy-preserving way to represent and measure socially meaningful interaction in these spaces, leaving studies to operationalize "interaction" differently across contexts and limiting practitioners' ability to evaluate whether design interventions are changing the forms of interaction that social capital theory predicts should matter. To address this field-level and methodological gap, we introduce the Dyadic User Engagement DataseT (DUET) dataset and an embedded kinesics recognition framework that operationalize Ekman and Friesen's kinesics taxonomy as a function-level interaction vocabulary aligned with social capital-relevant behaviors (e.g., reciprocity and attention coordination). DUET captures 12 dyadic interactions spanning all five kinesic functions-emblems, illustrators, affect displays, adaptors, and regulators-across four sensing modalities and three built-environment contexts, enabling privacy-preserving analysis of communicative intent through movement. Benchmarking six open-source, state-of-the-art human activity recognition models quantifies the difficulty of communicative-function recognition on DUET and highlights the limitations of ubiquitous monadic, action-level recognition when extended to dyadic, socially grounded interaction measurement. Building on DUET, our recognition framework infers communicative function directly from privacy-preserving skeletal motion without handcrafted action-to-function dictionaries; using a transfer-learning architecture, it reveals structured clustering of kinesic functions and a strong association between representation quality and classification performance while generalizing across subjects and contexts.

动作识别社会互动隐私保护城市设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。