融合视觉信号提升对话中修复请求的识别准确率
Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach

- 引入眼神、表情等视觉特征增强对话修复检测
- 在两个语料库上均显著优于纯文本音频模型
- 适合研究人机交互与多模态对话系统的开发者
其他发起的自我修复(OIR)是对话互动中的关键机制,接收方通过言语或非言语信号指出说话、听觉或理解问题,促使前一说话者进行修正。对于对话系统而言,准确识别此类修复启动行为至关重要。尽管已有研究发现OIR常伴随视线转移、面部表情、身体姿态和手势等非语言信号,但现有计算方法主要依赖文本和语音。本文提出一种新型多模态模型,融合来自对话分析的视觉特征,用于OIR检测与分类。我们在两个具有不同语言和交互场景的语料库上评估该方法。结果表明,视觉信息能持续提升性能,优于仅使用文本和音频的基线模型,并揭示了跨模态特征在两组数据中的贡献差异。
原文摘要 · Abstract (English)
Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a recipient signals a problem in speaking, hearing, or understanding, prompting the previous speaker to resolve it. In the case of conversational agents, it is essential to accurately identify these repair initiation strategies to address communication breakdowns efficiently. While conversational analysis studies have shown that OIR initiation is accompanied by both verbal and non-verbal signals such as gaze shifts, facial expressions, body postures, and hand gestures, existing computational approaches rely mainly on text and audio. This paper introduces a novel multimodal model for OIR detection and classification, incorporating a set of visual features drawn from conversation analysis. We evaluate our approach on two corpora with distinct languages and interaction settings. Results demonstrate that visual information consistently improves performance over text and audio baselines, and provide insights into cross-modal feature contributions across two corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。