用自然语言传输声音,让文字既是描述也是音频载体。
Communicating Sound Through Natural Language
- 用预训练大模型分析声音并转为英文词汇编码
- 传输文本保留可测量的声学结构,音质损失可控
- 适合需要可读、可编辑音频的应用场景
自然语言常用于描述或控制音频系统,却极少作为音频本身的传输载体。我们提出词法声学编码(LAC)框架,由预训练大语言模型充当发送端和接收端,仅通过自然语言进行声音传输。在固定系统提示下,发送端将输入波形分解为可解释的非学习声学特征,用特定词汇表量化每项特征,并转化为英文句子;接收端解析句子为声学约束,通过闭环优化还原波形。传输文本兼具丰富描述与音频载体功能。我们将LAC视为有限速率有损量化器,揭示词汇量、码率与保真度间的权衡。短时声音和符号音乐传输实验表明,纯文本能保留可测量的声学结构,同时保持可读性、可编辑性,并适配大语言模型通信流程。
原文摘要 · Abstract (English)
Natural language is widely used to describe, prompt, and control audio systems, but rarely serves as the representation carrying audio itself. We introduce lexical acoustic coding (LAC), a framework in which pre-trained LLM sender and receiver agents transmit sound through natural language. Under fixed system prompts, the agents write their own analysis and synthesis code, communicating only through a lexical sentence, shared vocabulary, and optional symbolic music structure. The sender analyzes an input waveform into interpretable, non-learned acoustic descriptors, quantizes each with a feature-specific interval vocabulary, and verbalizes the lexical code as English. The receiver parses the sentence back into lexical-acoustic constraints and renders a waveform through closed-loop refinement. The transmitted text serves as both a rich caption and as the transport representation itself. We frame LAC as a finite-rate lossy quantizer, exposing trade-offs between vocabulary size, rate, and fidelity. Experiments on short sounds and symbolic music transfer show that plain text preserves measurable acoustic structure while remaining interpretable, editable, and native to LLM-mediated communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。