arXiv:2605.17509eess.AS2026-05

让漫画中的拟声词图像能自动匹配对应声音,反之亦然。

Audio-Image Cross-Modal Retrieval with Onomatopoeic Images

论文配图:Audio-Image Cross-Modal Retrieval with Onomatopoeic Images
图 1 · 摘自论文原文
  • 用特定投影头对齐拟声图像与声音的嵌入表示。
  • 在50类声音事件上实现跨模态检索性能显著提升。
  • 适合漫画、动画等视觉化声音表达场景使用。

在多媒体制作中,寻找能准确传达创作者意图的声音效果仍主要依赖人工操作,尤其在漫画等视觉媒介中,通过文字形状、笔画、布局和装饰图案表现听觉印象的拟声词图像尤为常见。然而,拟声图像与通用声音之间的跨模态检索尚未被深入研究。本文提出一个双向检索框架,连接拟声图像与对应音效片段。不同于直接比较预训练图像与音频编码器提取的嵌入,我们训练了特定模态的投影头,以重新对齐视觉拟声与对应声音的表示。为此构建了多模态图像-音频拟声数据集(MIAO),包含50个声音事件类别的配对数据。实验表明,所提方法显著优于使用预训练CLIP和CLAP嵌入的零样本基线。结果证明,适配预训练表示可有效支持双向检索:从拟声图像到声音,以及从声音到拟声图像。

原文摘要 · Abstract (English)

Finding sound effects or environmental sounds that match a creator's intended impression remains a largely manual process in multimedia production. This is especially relevant for comics and other visual media, where visually stylized onomatopoeic expressions convey auditory impressions through letter shapes, strokes, layouts, and decorative patterns. However, cross-modal retrieval between onomatopoeic images and general sounds has been largely unexplored. This paper thus introduces a bidirectional retrieval framework between onomatopoeic images and the corresponding sound clips. Instead of directly comparing embeddings extracted from pretrained image and audio encoder, we train modality-specific projection heads that re-align the embeddings for visual onomatopoeia and corresponding sounds. We then construct the Multimodal Image-Audio Onomatopoeia dataset (MIAO), which contains paired onomatopoeic images and sound clips across 50 sound event classes. Experimental results show that the proposed method substantially outperforms a zero-shot baseline using pretrained CLIP and CLAP embeddings. These results demonstrate that adapting pretrained representations enables effective retrieval in both directions: from onomatopoeic images to sounds and from sounds to onomatopoeic images.

跨模态检索拟声词音频图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。