arXiv:2605.12987cs.CL2026-05被引 1

用多模态自我一致性推理提升酒精干预对话自动编码准确率

Leveraging Multimodal Self-Consistency Reasoning in Coding Motivational Interviewing for Alcohol Use Reduction

  • 融合语音语调、语言内容等多模态信号,通过12条推理路径投票决策
  • 在5个录音会话上实现46.40%的宏平均F1分数,优于单次推理方法
  • 适合需要高可靠性自动分析心理治疗对话的研究者和临床团队

动机访谈(MI)会话的编码对理解患者行为和预测干预效果至关重要,但依赖专业人员耗时耗力。本文利用音频-语言模型(ALM)分析原始音频,提出一种基于多模态自我一致性推理的自动编码方法。针对每个话语,采用四种互补提示:语言线索分析、声调感知、证据评分与对比推理,每种生成3个随机采样推理路径,共12条独立推理轨迹,最终通过多数投票确定结果。在5段去标识化录音会上测试,该方法达到52.56%准确率、54.03%精确率、47.45%召回率,宏平均F1为46.40%,显著优于基线。系统性消融实验表明,移除任一模块均导致性能下降。

原文摘要 · Abstract (English)

BACKGROUND: Coding Motivational Interviewing (MI) sessions is essential for understanding client behaviors and predicting outcomes, but it requires substantial time and labor from trained MI professionals. Recent advances in audio-language models (ALMs) offer new opportunities to automate MI coding by capturing multimodal behavioral signals. OBJECTIVE: This study aims to develop an automatic MI coding approach based on ALMs that analyzes raw audio input and integrates predictions from multiple reasoning trajectories using self-consistency to improve coding robustness. METHODS: We experimented with five recorded sessions from de-identified MI audio tapes. We deployed ALMs with four complementary analytic prompts to support utterance-level reasoning: analytic prompting for verbal cues, prosody-aware prompting for acoustic cues, evidence-scoring prompting for quantitative hypothesis testing, and comparative prompting for contrastive reasoning. Three stochastic samples were drawn for each prompt, generating 12 independent reasoning trajectories per utterance. Final predictions were determined by majority voting across all trajectories. RESULTS: Performance was evaluated using accuracy, precision, recall, and macro-F1 scores. The proposed multimodal self-consistency approach achieved 52.56% accuracy, 54.03% precision, 47.45% recall, and a macro-F1 score of 46.40%, exceeding baseline methods. Systematic ablation experiments that removed individual modules consistently degraded performance on the primary metrics. CONCLUSIONS: Multimodal self-consistency outperforms single-pass baseline prompting approaches for MI coding. These findings suggest that incorporating both what clients say and how they say it can support more reliable automatic MI coding.

动机访谈多模态自洽推理语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。