arXiv:2410.08470cs.HCcs.CV2024-10中稿 · ACM Multimedia 202…被引 9

用音视频融合与对话感知机制,提升对话中参与度估计精度。

DAT: Dialogue-Aware Transformer with Modality-Group Fusion for Human Engagement Estimation

  • 分模态独立融合音视频特征,再全局整合信息
  • 在多数据集上达成0.76的CCC得分,平均0.64
  • 适合关注人机交互与情感计算的研究者

参与度估计在理解人类社交行为中至关重要,日益受到情感计算与人机交互领域的关注。本文提出一种仅依赖音视频输入、语言无关的对话感知变换器框架(DAT),结合模态分组融合(MGF)方法,用于估计对话中的参与度。该方法先在每个模态内独立融合音频与视觉特征,再推断整体音视频内容,显著提升了模型性能与鲁棒性。此外,引入的对话感知变换器同时考虑目标参与者自身行为及其对话伙伴的线索,以更准确估计其参与水平。本方法在MultiMediate'24举办的多领域参与度估计挑战赛中进行了严格测试,相比基线模型,在参与度回归精度上表现优异。在NoXi Base测试集上取得0.76的一致性相关系数(CCC),在NoXi Base、NoXi-Add与MPIIGI三个测试集上的平均CCC达0.64。

原文摘要 · Abstract (English)

Engagement estimation plays a crucial role in understanding human social behaviors, attracting increasing research interests in fields such as affective computing and human-computer interaction. In this paper, we propose a Dialogue-Aware Transformer framework (DAT) with Modality-Group Fusion (MGF), which relies solely on audio-visual input and is language-independent, for estimating human engagement in conversations. Specifically, our method employs a modality-group fusion strategy that independently fuses audio and visual features within each modality for each person before inferring the entire audio-visual content. This strategy significantly enhances the model's performance and robustness. Additionally, to better estimate the target participant's engagement levels, the introduced Dialogue-Aware Transformer considers both the participant's behavior and cues from their conversational partners. Our method was rigorously tested in the Multi-Domain Engagement Estimation Challenge held by MultiMediate'24, demonstrating notable improvements in engagement-level regression precision over the baseline model. Notably, our approach achieves a CCC score of 0.76 on the NoXi Base test set and an average CCC of 0.64 across the NoXi Base, NoXi-Add, and MPIIGI test sets.

参与度估计多模态融合对话感知情感计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。