arXiv:2608.19174cs.SDcs.AI2026-08
用声音模仿检索音效,提出两种高效微调方法。
Finetuning Strategies for Querying Sounds by Vocal Imitation
- 采用冻结预训练的CED编码器做对比学习。
- 使用MobileNetV3结合三元组与半硬负样本联合优化。
- 适合音效检索、人声模仿交互等应用开发。
本技术报告介绍了我们在AES AIMLA 2025挑战赛中关于通过人声模仿检索音效的获胜提交方案。我们研究了两种互补的微调策略:一是使用冻结的预训练CED编码器进行对比学习;二是采用MobileNetV3编码器,结合带半硬负样本的对比-三元组联合学习。报告已更新,包含挑战赛后公布的细节。
原文摘要 · Abstract (English)
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.
音效检索声音模仿微调策略
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。