arXiv:2602.19976cs.SD2026-02中稿 · ICLR被引 4

用自适应调制生成翻唱歌曲,参数少效果好。

SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation

  • 通过元素级线性调制实现旋律精准控制。
  • 生成质量优于现有方法,参数量不足30%。
  • 适合音乐生成与创作研究者使用。

翻唱歌曲是音乐文化的重要组成部分,既保留原曲主旋律,又通过重新诠释赋予新情感和主题。尽管已有研究探索了基于旋律条件的文本到音乐生成,但翻唱歌曲生成仍鲜有涉及。本文将翻唱生成建模为条件生成任务,同时根据原始人声旋律和文本提示生成新的人声与伴奏。为此提出SongEcho框架,采用实例自适应逐元素线性调制(IA-EiLM),改进条件注入机制与条件表示。为增强条件注入,将特征线性调制(FiLM)扩展为逐元素线性调制(EiLM),实现更精确的时序对齐;为优化条件表示,提出实例自适应条件精炼(IACR),通过与生成模型隐藏状态交互,获得实例自适应的条件特征。此外,针对高质量全曲数据集稀缺问题,构建Sun70k,一个标注全面的高质量AI歌曲数据集。多数据集实验表明,所提方法生成效果优于现有方法,且参数量少于30%。代码、数据集与演示已公开。

原文摘要 · Abstract (English)

Cover songs constitute a vital aspect of musical culture, preserving the core melody of an original composition while reinterpreting it to infuse novel emotional depth and thematic emphasis. Although prior research has explored the reinterpretation of instrumental music through melody-conditioned text-to-music models, the task of cover song generation remains largely unaddressed. In this work, we reformulate our cover song generation as a conditional generation, which simultaneously generates new vocals and accompaniment conditioned on the original vocal melody and text prompts. To this end, we present SongEcho, which leverages Instance-Adaptive Element-wise Linear Modulation (IA-EiLM), a framework that incorporates controllable generation by improving both conditioning injection mechanism and conditional representation. To enhance the conditioning injection mechanism, we extend Feature-wise Linear Modulation (FiLM) to an Element-wise Linear Modulation (EiLM), to facilitate precise temporal alignment in melody control. For conditional representations, we propose Instance-Adaptive Condition Refinement (IACR), which refines conditioning features by interacting with the hidden states of the generative model, yielding instance-adaptive conditioning. Additionally, to address the scarcity of large-scale, open-source full-song datasets, we construct Suno70k, a high-quality AI song dataset enriched with comprehensive annotations. Experimental results across multiple datasets demonstrate that our approach generates superior cover songs compared to existing methods, while requiring fewer than 30% of the trainable parameters. The code, dataset, and demos are available at https://github.com/lsfhuihuiff/SongEcho_ICLR2026.

翻唱生成音乐生成条件生成模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。