arXiv:2409.13507cs.GRcs.CL2024-09SIGGRAPH被引 2

用声音模仿声音,让机器像人一样‘画声’。

Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation

论文配图:Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation
图 1 · 摘自论文原文
  • 基于人声声道模型,调节参数生成类人声音模仿。
  • 加入沟通策略建模后,更符合人类直觉。
  • 为计算机图形中的表征研究提供新视角。

我们提出一种自动生成类人声音模仿的方法:相当于用声音‘素描’,而非视觉。从人声声道的模拟模型出发,首先通过调整控制参数,使合成声音在感知显著的听觉特征上匹配目标声音。接着,为更贴近人类认知,引入沟通认知理论,考虑说话者对听众的策略性推理。最后,通过多组实验和用户研究验证,加入这种沟通推理机制后,生成结果比仅匹配听觉特征更符合人类直觉。该发现对计算机图形学中表征问题的研究具有广泛意义。

原文摘要 · Abstract (English)

We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try generating vocal imitations by tuning the model's control parameters to make the synthesized vocalization match the target sound in terms of perceptually-salient auditory features. Then, to better match human intuitions, we apply a cognitive theory of communication to take into account how human speakers reason strategically about their listeners. Finally, we show through several experiments and user studies that when we add this type of communicative reasoning to our method, it aligns with human intuitions better than matching auditory features alone does. This observation has broad implications for the study of depiction in computer graphics.

声音模仿人机交互音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。