用大模型从虚拟现实语音中自动生成情绪标注,解决人工标注难的问题。
LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning

- 基于大模型的上下文学习,用少量语音样本生成情绪标签。
- 在多用户虚拟环境中实现连续情绪状态的自动标注,准确率超基准方法。
- 适合做人机交互、团队协作研究的研究者参考。
理解人类状态与互动动态是人机交互的核心目标。随着交互方式日益沉浸化,虚拟现实(VR)已成为研究协作工作的重要平台。在此类场景中,评估团队协作状态(如团队表现与韧性)需要从多模态传感器数据(如语音信号)中持续可靠地推断团队层面的隐含认知与情感状态。然而,由于传感器噪声、情境差异及专家标注稀疏,生成这些隐状态的真实标签仍具挑战。传统自报告方法仅提供静态、延迟的测量,难以捕捉语音数据中的动态过程。本文提出一种基于大语言模型(LLM)的智能推理流程,通过上下文学习(ICL)从流式语音数据中自动化生成情绪相关的合成真实标签。利用LLM的泛化能力,采用少样本示范(配对语音与转录文本)进行任务适配,其性能接近微调但无需参数更新开销。为构建有效提示,采用基于检索的选择策略,根据声学特征空间相似性动态选取相关示范音频。
原文摘要 · Abstract (English)
Understanding human states and interaction dynamics is a core goal of human-computer interaction (HCI). As interaction paradigms become more immersive, virtual reality (VR) has emerged as a powerful platform for studying collaborative work. In such settings, evaluating team collaboration states, including team performance and team resilience, requires continuous and reliable inference of latent team-level cognitive and affective states from multi-modal sensor data, such as speech signals. However, generating ground truth labels for these latent states remains challenging due to sensor-induced noise, contextual variability, and sparse expert annotations. Traditional self-reporting approaches provide only static and delayed measurements and are therefore insufficient for capturing dynamic team processes reflected in continuous speech data. In this work, we propose a large language model (LLM)-driven, agentic inference workflow for automated emotion-related synthetic ground truth generation from streaming speech data in multi-user VR environments. Leveraging the generalization capabilities of LLMs, we use In-Context Learning (ICL) with few-shot demonstrations of paired audio-based samples and their corresponding transcriptions. ICL tends to achieve task adaptation comparable to model fine-tuning while circumventing the computational overhead of parameter updates. To construct informative and robust in-context prompts, we adopt a retrieval-based selection strategy that dynamically identifies relevant audio demonstrations based on similarity in the acoustic feature space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。