arXiv:2601.09448cs.SDcs.AI2026-01中稿 · publication in the…

用自然语言控制音效均衡,让系统自动适应不同听感需求。

One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization

论文配图:One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization
图 1 · 摘自论文原文
  • 用大模型将文本提示转为音效参数,实现对话式调节。
  • 在实验中显著提升对人群偏好分布的匹配度,优于随机和固定设置。
  • 适合想轻松调音的普通用户或需要个性化音效的专业场景。

传统音频均衡是静态过程,需手动调整以适应不同听觉情境(如心情、位置或社交环境)。本文提出一种基于大语言模型(LLM)的方法,将自然语言文本提示映射为均衡参数,实现对话式音效控制。利用受控听音实验收集的数据,模型通过上下文学习和参数高效微调,可靠地对齐群体偏好的均衡设置。评估采用捕捉用户偏好差异的分布度量方法,结果显示其在分布对齐上显著优于随机采样和静态预设基线。结果表明,LLM可作为“人工均衡器”,推动更易用、上下文感知且达到专家水平的音频调校技术发展。

原文摘要 · Abstract (English)

Conventional audio equalization is a static process that requires manual and cumbersome adjustments to adapt to changing listening contexts (e.g., mood, location, or social setting). In this paper, we introduce a Large Language Model (LLM)-based alternative that maps natural language text prompts to equalization settings. This enables a conversational approach to sound system control. By utilizing data collected from a controlled listening experiment, our models exploit in-context learning and parameter-efficient fine-tuning techniques to reliably align with population-preferred equalization settings. Our evaluation methods, which leverage distributional metrics that capture users' varied preferences, show statistically significant improvements in distributional alignment over random sampling and static preset baselines. These results indicate that LLMs could function as ``artificial equalizers," contributing to the development of more accessible, context-aware, and expert-level audio tuning methods.

音频处理大模型应用人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。