arXiv:2512.02558cs.AI2025-12

融合多模态与监督文档,提升心理咨询中的共情预测准确率

Empathy Level Prediction in Multi-Modal Scenario with Supervisory Documentation Assistance

  • 用视频、音频、文本多模态特征融合预测共情水平
  • 引入监督文档作为训练时的额外信息,提升文本特征质量
  • 适合心理辅导、人机共情系统研究者参考

现有共情预测方法多聚焦单一模态(如文本),忽视多模态融合与额外优势信息。为此,本文提出一种结合视频、音频与文本的多模态共情预测方法,包含多模态共情预测与监督文档辅助训练两部分。采用预训练网络提取各模态特征,并通过跨模态融合生成统一表示,用于共情标签预测。为增强文本特征提取,引入监督文档作为训练阶段的特权信息——由督导人员编写,聚焦咨询主题与咨询师共情表现。利用隐狄利克雷分布(LDA)识别潜在话题分布,对文本特征施加约束。该特权信息仅在训练时可用,预测阶段不可见。在多模态对话共情数据集上的实验表明,本方法优于现有主流方法。

原文摘要 · Abstract (English)

Prevalent empathy prediction techniques primarily concentrate on a singular modality, typically textual, thus neglecting multi-modal processing capabilities. They also overlook the utilization of certain privileged information, which may encompass additional empathetic content. In response, we introduce an advanced multi-modal empathy prediction method integrating video, audio, and text information. The method comprises the Multi-Modal Empathy Prediction and Supervisory Documentation Assisted Training. We use pre-trained networks in the empathy prediction network to extract features from various modalities, followed by a cross-modal fusion. This process yields a multi-modal feature representation, which is employed to predict empathy labels. To enhance the extraction of text features, we incorporate supervisory documents as privileged information during the assisted training phase. Specifically, we apply the Latent Dirichlet Allocation model to identify potential topic distributions to constrain text features. These supervisory documents, created by supervisors, focus on the counseling topics and the counselor's display of empathy. Notably, this privileged information is only available during training and is not accessible during the prediction phase. Experimental results on the multi-modal and dialogue empathy datasets demonstrate that our approach is superior to the existing methods.

共情预测多模态学习监督文档心理咨询

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。