arXiv:2602.12714cs.LG2026-02被引 1

让语音模型像侦探一样通过证据推理情绪,更准且可解释。

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools -- From Consensus Learning to Ambiguity-Driven Emotion Reasoning

  • 将情绪判断转为多轮探查过程,动态调用语义与声学工具。
  • 在多数场景下提升主情绪识别准确率,显著改善小众情绪识别。
  • 把少数人标注当信号而非噪声,适合需要可解释性的场景。

语音大模型(SLLMs)能进行高层次情绪推理,但常因缺乏可验证的声学证据而产生脱离实际的文本偏见判断。相比之下,自监督语音编码器如WavLM虽具备强声学表征,但判别能力不透明、可解释性差。为此,我们提出ADEPT(基于证据探针工具的情绪代理解码框架),将情绪识别重构为多轮探究过程,而非单次预测。ADEPT将SLLM转化为智能体,持续维护候选情绪集,并在结构化流程中自适应调用语义与声学探针工具,完成候选生成、证据收集与裁决。关键在于,它实现了从共识学习到模糊驱动情绪推理的范式转变:由于人类情感本具复杂性且常共现,我们不再丢弃少数标注,而是将其视为有信息量的感知信号。最后,通过引入群体相对策略优化(GRPO)与证据可信门控机制,显式关联工具使用行为与预测质量,强制实现基于证据的推理。实验表明,ADEPT在多数设置下提升了主情绪准确率,大幅改善了小众情绪的刻画能力,且推理过程可审计、基于可验证的声学与语义证据。

原文摘要 · Abstract (English)

Speech Large Language Models (SLLMs) enable high-level emotion reasoning but often produce ungrounded, text-biased judgments without verifiable acoustic evidence. In contrast, self-supervised speech encoders such as WavLM provide strong acoustic representations yet remain opaque discriminative models with limited interpretability. To bridge this gap, we introduce ADEPT (Agentic Decoding of Emotion via Evidence Probing Tools), a framework that reframes emotion recognition as a multi-turn inquiry process rather than a single-pass prediction. ADEPT transforms an SLLM into an agent that maintains an evolving candidate emotion set and adaptively invokes dedicated semantic and acoustic probing tools within a structured pipeline of candidate generation, evidence collection, and adjudication. Crucially, ADEPT enables a paradigm shift from consensus learning to ambiguity-driven emotion reasoning. Since human affect exhibits inherent complexity and frequent co-occurrence of emotions, we treat minority annotations as informative perceptual signals rather than discarding them as noise. Finally, we integrate Group Relative Policy Optimization (GRPO) with an Evidence Trust Gate to explicitly couple tool-usage behaviors with prediction quality and enforce evidence-grounded reasoning. Experiments show that ADEPT improves primary emotion accuracy in most settings while substantially improving minor emotion characterization, producing explanations grounded in auditable acoustic and semantic evidence.

情绪识别语音模型可解释性智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。