发现语音生成模型中可直接操控情绪的神经元,实现无需训练的情绪控制。
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
- 通过激活聚合筛选出特定情绪响应的微型神经元
- 干预后情绪匹配度提升,且保持语言准确性
- 适用于多种模型,适合需要实时情绪调整的场景
大型音频-语言模型(LALMs)能生成富有表现力的语音,但可靠的情绪控制仍难以实现:转换常偏离目标情感,且可能因拒绝、幻觉或改写导致语言保真度下降。本文首次开展语音生成LALMs的神经元级情绪控制研究,证明紧凑的情绪敏感神经元(ESNs)具有因果可操作性,可在推理阶段实现无需训练的情绪调节。ESNs通过强化情感实现与内容保留的成功过滤激活聚合方法识别。在三个LALMs(Qwen2.5-Omni-7B、MiniCPM-o 4.5、Kimi-Audio)上,ESN干预在未见说话人上均取得情绪特异性提升,经自动与人工评估验证。可控性受选择器设计、掩码稀疏度、过滤策略及干预强度影响。研究建立了无需训练的情绪控制机制框架。
原文摘要 · Abstract (English)
Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic fidelity through refusals, hallucinations, or paraphrase. We present, to our knowledge, the first neuron-level study of emotion control in speech-generative LALMs and demonstrate that compact emotion-sensitive neurons (ESNs) are causally actionable, enabling training-free emotion steering at inference time. ESNs are identified via success-filtered activation aggregation enforcing both emotion realization and content preservation. Across three LALMs (Qwen2.5-Omni-7B, MiniCPM-o 4.5, Kimi-Audio), ESN interventions yield emotion-specific gains that generalize to unseen speakers and are supported by automatic and human evaluation. Controllability depends on selector design, mask sparsity, filtering, and intervention strength. Our results establish a mechanistic framework for training-free emotion control in speech generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。