评测音频编码器在多任务真实场景下的能力,推动音频模型进步。
The ICME 2025 Audio Encoder Capability Challenge
- 提交可将原始波形转为连续嵌入的预训练音频编码器
- 分参数化与无参数两赛道,测试语音、环境音、音乐等多任务表现
- 适合关注音频模型实用性和多任务泛化能力的研究者
本挑战旨在评估音频编码器的能力,尤其在多任务学习和真实应用场景中的表现。参赛者需提交将原始波形映射为连续嵌入的预训练音频编码器,这些编码器将在包括语音、环境声音和音乐在内的多种任务上进行测试,重点考察其在真实场景中的可用性。挑战设两个赛道:Track A为参数化评估,Track B为参数自由评估。该挑战为评估和推进音频编码器设计的最先进水平提供了一个平台。
原文摘要 · Abstract (English)
This challenge aims to evaluate the capabilities of audio encoders, especially in the context of multi-task learning and real-world applications. Participants are invited to submit pre-trained audio encoders that map raw waveforms to continuous embeddings. These encoders will be tested across diverse tasks including speech, environmental sounds, and music, with a focus on real-world usability. The challenge features two tracks: Track A for parameterized evaluation, and Track B for parameter-free evaluation. This challenge provides a platform for evaluating and advancing the state-of-the-art in audio encoder design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。