arXiv:2601.12289cs.SDcs.LG2026-01AAAI被引 2

提出统一框架ParaMETA,从语音中解耦并控制说话风格。

ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech

  • 将语音映射到独立子空间,实现情绪、年龄等风格的解耦表征。
  • 在多任务分类上准确率优于基线,生成语音更自然且可精细调控。
  • 适合需要风格可控语音生成的场景,如人机交互与内容创作。

学习不同说话风格(如情绪、年龄、性别)的代表性嵌入,对识别任务(如认知计算和人机交互)与生成任务(如风格可控语音合成)至关重要。本文提出ParaMETA,一种统一且灵活的框架,直接从语音中学习并控制说话风格。不同于依赖单任务模型或跨模态对齐的现有方法,ParaMETA通过将语音投影到每个风格类型的专用子空间,实现解耦的、任务特定的嵌入表示。该设计减少了任务间干扰,缓解了负迁移问题,使单一模型可同时处理情绪、性别、年龄和语言分类等多类副语言任务。除识别外,ParaMETA还可用于文本到语音生成模型中的细粒度风格控制,支持语音或文本提示,并可在修改某一风格时保持其他风格不变。大量实验表明,ParaMETA在分类准确率上超越强基线,生成语音更自然、更具表现力,同时保持轻量高效,适用于实际应用。

原文摘要 · Abstract (English)

Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and human-computer interaction) and generative tasks (e.g., style-controllable speech generation). In this work, we introduce ParaMETA, a unified and flexible framework for learning and controlling speaking styles directly from speech. Unlike existing methods that rely on single-task models or cross-modal alignment, ParaMETA learns disentangled, task-specific embeddings by projecting speech into dedicated subspaces for each type of style. This design reduces inter-task interference, mitigates negative transfer, and allows a single model to handle multiple paralinguistic tasks such as emotion, gender, age, and language classification. Beyond recognition, ParaMETA enables fine-grained style control in Text-To-Speech (TTS) generative models. It supports both speech- and text-based prompting and allows users to modify one speaking styles while preserving others. Extensive experiments demonstrate that ParaMETA outperforms strong baselines in classification accuracy and generates more natural and expressive speech, while maintaining a lightweight and efficient model suitable for real-world applications.

语音生成风格控制解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。