arXiv:2604.01711cs.CL2026-04被引 1

让大模型结合人类判断,提升越南语语音情绪识别准确率

Human-Guided Reasoning with Large Language Models for Vietnamese Speech Emotion Recognition

  • 用大模型推理+人类标注规则,区分易判和难判样本
  • 在2764样本数据集上达86.59%准确率,对模糊情绪更敏感
  • 适合低资源语言情绪识别,不依赖特定模型架构

越南语语音情绪识别(SER)因声学特征模糊且缺乏可靠标注数据而困难,尤其在真实场景中情绪边界不清晰。本文提出一种人机协同框架,将人类知识融入学习过程,而非仅依赖数据驱动模型。核心为基于大语言模型(LLM)的推理机制,利用基于声学特征的模型提供置信度和特征级证据作为辅助信号。引入基于置信度的路由机制,区分简单与模糊样本,将不确定案例交由大模型根据人类标注行为提炼的结构化规则进行深度推理。同时采用迭代优化策略,通过错误分析与规则更新持续提升性能。在包含2,764个样本、三个情绪类别(平静、愤怒、恐慌)的越南语语音数据集上进行实验,标注者间一致性高(Fleiss Kappa = 0.8574),确保真实标签可信。所提方法表现优异,准确率达86.59%,宏平均F1在0.85-0.86之间,有效处理模糊及难分类样本。整体表明,结合数据驱动模型与人类推理可构建鲁棒、模型无关的低资源语音情绪识别方案。

原文摘要 · Abstract (English)

Vietnamese Speech Emotion Recognition (SER) remains challenging due to ambiguous acoustic patterns and the lack of reliable annotated data, especially in real-world conditions where emotional boundaries are not clearly separable. To address this problem, this paper proposes a human-machine collaborative framework that integrates human knowledge into the learning process rather than relying solely on data-driven models. The proposed framework is centered around LLM-based reasoning, where acoustic feature-based models are used to provide auxiliary signals such as confidence and feature-level evidence. A confidence-based routing mechanism is introduced to distinguish between easy and ambiguous samples, allowing uncertain cases to be delegated to LLMs for deeper reasoning guided by structured rules derived from human annotation behavior. In addition, an iterative refinement strategy is employed to continuously improve system performance through error analysis and rule updates. Experiments are conducted on a Vietnamese speech dataset of 2,764 samples across three emotion classes (calm, angry, panic), with high inter-annotator agreement (Fleiss Kappa = 0.8574), ensuring reliable ground truth. The proposed method achieves strong performance, reaching up to 86.59% accuracy and Macro F1 around 0.85-0.86, demonstrating its effectiveness in handling ambiguous and hard-to-classify cases. Overall, this work highlights the importance of combining data-driven models with human reasoning, providing a robust and model-agnostic approach for speech emotion recognition in low-resource settings.

语音情绪识别大模型人机协同低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。