arXiv:2506.02443cs.SDeess.AS2025-06被引 3

提出无需文本的音频到音频智能框架,服务7亿音频文盲人群

Breaking the Barriers of Text-Hungry and Audio-Deficient AI

  • 构建纯音频输入输出模型,绕过文本依赖
  • 用多尺度音频语义表征实现高保真语音生成
  • 适合无文字语言或偏好语音交互的用户

全球有超过7164种已确认语言,但当前主流机器智能架构仍严重偏向书面文本,导致超过7亿居住在偏远地区的人群(主要为音频识读者)被排除在外。本文提出一种全文本无依赖的音频到音频智能框架,专为这一未被覆盖群体及追求音频效率的用户设计。创新性地开发了谱图、小波、时频标量和音素单元等多种音频到音频翻译架构。核心是多尺度音频语义变换(MAST),可编码音调、语调、说话人特征与表达风格。进一步将MAST融入基于分数布朗运动的均值场型分数扩散框架,实现无需文本监督的高保真、语义一致语音生成。系统可直接从原始音频学习,适用于无文字或极少数字化的语言。本工作标志着向音频原生智能系统的根本性转变,拓展了长期被排除在现有智能生态外社区的语义技术接入能力。

原文摘要 · Abstract (English)

While global linguistic diversity spans more than 7164 recognized languages, the current dominant architecture of machine intelligence remains fundamentally biased toward written text. This bias excludes over 700 million people particularly in rural and remote regions who are audio-literate. In this work, we introduce a fully textless, audio-to-audio machine intelligence framework designed to serve this underserved population, and all the people who prefer audio-efficiency. Our contributions include novel Audio-to-Audio translation architectures that bypass text entirely, including spectrogram-, scalogram-, wavelet-, and unit-based models. Central to our approach is the Multiscale Audio-Semantic Transform (MAST), a representation that encodes tonal, prosodic, speaker, and expressive features. We further integrate MAST into a fractional diffusion of mean-field-type framework powered by fractional Brownian motion. It enables the generation of high-fidelity, semantically consistent speech without reliance on textual supervision. The result is a robust and scalable system capable of learning directly from raw audio, even in languages that are unwritten or rarely digitized. This work represents a fundamental shift toward audio-native machine intelligence systems, expanding access to language technologies for communities historically left out of the current machine intelligence ecosystem.

音频智能无文本语音生成语言平等

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。