用智能体生成多模态音乐推荐对话数据,助力生成式推荐模型训练。
TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
- 设计多角色大模型智能体模拟用户与推荐系统对话
- 生成含音频图像的多模态对话数据,覆盖多种场景
- 适合研究生成式音乐推荐与多模态对话系统的学者
我们提出 TalkPlayData 2,一个由智能体数据流水线生成的多模态对话式音乐推荐合成数据集。该流水线中,多个大语言模型(LLM)代理在不同角色下运行,通过专用提示词和信息访问权限,记录听众代理与推荐系统代理之间的对话。为覆盖多样对话场景,每个对话中的听众代理均基于微调后的对话目标进行条件化。所有代理均为多模态,支持音频与图像输入,实现对多模态推荐与对话的仿真。在 LLM 作为裁判和主观评估实验中,TalkPlayData 2 在多个方面达成生成式音乐推荐模型训练目标。数据集及生成代码已公开于 https://talkpl-ai.github.io。
原文摘要 · Abstract (English)
We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In the proposed pipeline, multiple large language model (LLM) agents are created under various roles with specialized prompts and access to different parts of information, and the chat data is acquired by logging the conversation between the Listener LLM and the Recsys LLM. To cover various conversation scenarios, for each conversation, the Listener LLM is conditioned on a finetuned conversation goal. Finally, all the LLMs are multimodal with audio and images, allowing a simulation of multimodal recommendation and conversation. In the LLM-as-a-judge and subjective evaluation experiments, TalkPlayData 2 achieved the proposed goal in various aspects related to training a generative recommendation model for music. TalkPlayData 2 and its generation code are released at https://talkpl-ai.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。