让用户用自然语言主动描述听歌场景,实现更智能的音乐发现。
MuChator: Enabling Active Music Discovery via Conversational Music LLMs in Douyin Music

- 构建音乐知识预训练三阶段框架,融合客观、主观与个性化偏好。
- 通过上下文感知指令微调,精准理解用户模糊的听歌意图。
- 支持真实场景下口语化表达,适合想主动探索音乐的用户。
抖音音乐作为日活数百万的大规模平台,采用沉浸式信息流推荐模式,用户被动浏览音乐内容。该模式虽有效但限制了用户主动表达听歌意图的能力。真实场景中的音乐发现常为情境化、口语化,需求模糊且不明确。现有大模型在音乐领域存在知识不足、查询推理能力弱及个性化理解浅的问题。为此,我们提出 MuChator,一个基于 MusicLLM 的交互式音乐发现框架。其包含三个核心组件:(1) 音乐知识预训练,分三阶段逐步注入客观音乐知识、主观音乐感受和个性化偏好;(2) 上下文感知指令微调,通过自动化合成高质量用户-查询-音乐三元组,对齐模型与情境化意图;(3) 混合奖励建模的偏好对齐,联合建模意图相关性、个性化偏好与基本约束,并使用 GRPO 增强学习优化。在工业级音乐推荐数据集上的评估显示,MuChator 超越主流私有模型(如 Gemini-3-Pro)。已在字节跳动抖音音乐应用上线,线上 A/B 测试中用户活跃天数提升 46.49%。
原文摘要 · Abstract (English)
Douyin Music, a large-scale platform with millions of daily users, adopts an immersive, feed-based discovery paradigm, where users passively explore music through continuous recommendations. While effective for passive music discovery, this paradigm restricts users to recommendation results and provides limited support for explicitly specifying listening intents. Unlike conventional search, where users express well-defined intents through explicit queries such as specific songs or artists, real-world active music discovery is often situational and colloquial, involving vague or underspecified requests. While LLMs enable natural language interaction, their direct use in music discovery remains limited by insufficient music-domain knowledge, lack of music-query collaborative reasoning, and shallow understanding of personalized preferences. To address these challenges, we introduce MuChator, an interactive MusicLLM-based framework that enables users to actively express situational music intents in natural language. MuChator incorporates three key components: (1) Music Knowledge Pre-training, a three-stage scheme that incrementally injects objective music knowledge, subjective music knowledge, and personalized music preferences into LLMs; (2) Context-aware Instruction Tuning, which constructs high-quality user-query-music triplets through an automated synthesis pipeline to align LLMs with active and situational user intents; and (3) Preference Alignment with Hybrid RM, which jointly models intent relevance, personalized preferences, and basic constraints, and is optimized using GRPO-based reinforcement learning. Extensive evaluations on industrial music recommendation datasets demonstrate that MuChator outperforms leading proprietary models, such as Gemini-3-Pro. The model has been deployed on Douyin Music App within ByteDance, with 46.49\% improvement of user active days in online A/B test.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。