arXiv:2602.23387cs.SDcs.AI2026-02

让AI对话更像真人,有情感有语气

Hello-Chat: Towards Realistic Social Audio Interactions

  • 用真实对话数据训练,穿插多模态学习提升自然度
  • 在语音韵律和情绪匹配上超越现有模型
  • 适合开发有共情能力的智能客服或虚拟助手

近年来,大音频语言模型(LALMs)在语音识别和翻译任务中表现卓越。然而,现有模型常因感知与表达脱节,生成机械化的“朗读式”语音,缺乏真实社交互动中的自发性与情感共鸣。本文提出Hello-Chat,一个面向真实社交场景的端到端音频语言模型。通过利用大规模真实对话数据集,并采用模态交错训练策略,该模型在拟人化生成方面取得突破。实验表明,Hello-Chat不仅在特定音频理解任务上达到当前最优(SOTA)水平,还在语调自然度和情绪一致性上显著优于现有基线模型,为下一代共情型AI代理的发展铺平道路。

原文摘要 · Abstract (English)

Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer from a disconnect between perception and expression, resulting in a robotic "read-speech" style that lacks the spontaneity and emotional resonance of real human interaction. In this report, we introduce Hello-Chat, an end-to-end audio language model designed for realistic social scenarios. By leveraging a massive dataset of real-life conversations and employing a modality-interleaved training strategy, Hello-Chat achieves a breakthrough in anthropomorphic generation. Experimental results show that our model not only reaches state-of-the-art (SOTA) performance on specific audio understanding tasks but also significantly outperforms existing baselines in prosodic naturalness and emotional alignment, paving the way for the next generation of empathetic AI agents.

语音生成共情AI音频模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。