arXiv:2607.17452cs.CL2026-07中稿 · , to appear in Pro…

评估对话中多模态信号的可靠性,发现互动特征最稳定。

How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

论文配图:How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks
图 1 · 摘自论文原文
  • 构建三维评估框架,测试预测准确、跨任务泛化和重测信度。
  • 语言特征预测认知负荷最准但无法跨任务通用,声学特征受说话人影响大。
  • 只有互动特征在控制说话人后仍可靠,适合用于鲁棒对话系统设计。

测量对话状态如认知负荷和对话权力需依赖既具预测性又跨任务稳定的特征。我们提出一个三维评估框架,评估交互、声学与语言特征在视频协作任务中的预测准确性、跨任务泛化性和重测信度(基于AVCAffe数据集,53对参与者,9项任务)。结果表明:无单一特征家族在三维度均占优。语言特征在认知负荷预测中表现最佳,但在跨任务评估中崩溃,显示对任务特定词汇敏感;声学特征的可靠性在控制说话人身份后大幅下降,说明其主要反映语音特性而非对话状态。仅交互特征在说话人归一化后保持稳定。分类权力角色始终接近随机水平,表明任务级聚合行为难以预测对话权力。研究揭示三点:(1) 语言特征预测强但泛化差;(2) 声学可靠性在控制说话人后趋近零,挑战现有评估范式;(3) 互动特征是唯一真正可靠的信号,且能预测同组内认知负荷不对称。结论支持在对话系统中采用说话人归一化与多维评估以实现上下文感知的稳健特征选择。

原文摘要 · Abstract (English)

Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts. We present a three-dimensional evaluation framework assessing predictive accuracy, cross-task generalizability, and test-retest reliability, applied to interactional, acoustic, and linguistic features extracted from dyadic conversations during collaborative tasks performed over a video-conferencing platform (AVCAffe dataset; 53 dyads, 9 tasks). Our results show that no single feature family dominates all three dimensions. Linguistic features show the highest predictive accuracy for cognitive load but collapse under cross-task evaluation, revealing sensitivity to task-specific vocabulary. Additionally, acoustic reliability, often reported as evidence of feature stability, degrades once speaker identity is controlled, confirming that standard prosodic features measure vocal characteristics rather than conversational state. Interaction features provide the only genuinely reliable signal, unchanged after speaker normalization. Interestingly, classifying power role remained near chance baseline across all conditions, indicating limitations of task-level aggregated behavior for predicting power role in conversation. Our findings reveal three insights: (1) linguistic features predict best but generalize poorly across task contexts; (2) acoustic reliability collapses to near-zero once speaker identity is controlled, challenging standard evaluation practice; and (3) interaction features provide the only genuinely reliable signal, with floor dominance predicting within-dyad cognitive load asymmetry. These results argue for speaker normalization and multi-dimensional evaluation as prerequisites for context-aware, robust multimodal feature selection in conversational systems.

多模态分析对话状态可靠性评估交互特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。