用合成数据训练大模型,让用户能用自然语言描述复杂观影偏好。
Optimizing Recommendations using Fine-Tuned LLMs
- 通过模拟真实互动生成对话式合成数据,支持复杂偏好表达。
- 实现包含主题、氛围等多维度的自然语言查询,提升推荐精准度。
- 适合构建下一代对话式影视推荐系统,推动个性化体验升级。
随着数字媒体平台不断满足用户日益变化的期望,提供高度个性化且直观的电影与媒体推荐已成为吸引和留住观众的关键。传统系统通常依赖关键词搜索与推荐技术,使用户受限于特定关键词及其组合。本文提出一种方法,通过建模真实用户互动生成合成数据集,创建反映多样化偏好的对话式数据。这使得用户能以更丰富的方式表达复杂偏好,如情绪、剧情细节和主题元素,而不仅限于类型、片名或演员等传统标准。当前搜索系统无法支持类似‘寻找一部奇幻电影,主角是冰原上的恶狼,设定在严酷的冰封世界,主题为忠诚与生存’的查询。基于此,我们评估了合成数据集在多样性与训练及基准测试中的有效性,尤其填补了传统数据集常缺失的领域。该方法通过支持表达性强、自然的用户查询,增强了个性化与准确性,为下一代基于对话式AI的数字娱乐搜索与推荐系统奠定了基础。
原文摘要 · Abstract (English)
As digital media platforms strive to meet evolving user expectations, delivering highly personalized and intuitive movies and media recommendations has become essential for attracting and retaining audiences. Traditional systems often rely on keyword-based search and recommendation techniques, which limit users to specific keywords and a combination of keywords. This paper proposes an approach that generates synthetic datasets by modeling real-world user interactions, creating complex chat-style data reflective of diverse preferences. This allows users to express more information with complex preferences, such as mood, plot details, and thematic elements, in addition to conventional criteria like genre, title, and actor-based searches. In today's search space, users cannot write queries like ``Looking for a fantasy movie featuring dire wolves, ideally set in a harsh frozen world with themes of loyalty and survival.'' Building on these contributions, we evaluate synthetic datasets for diversity and effectiveness in training and benchmarking models, particularly in areas often absent from traditional datasets. This approach enhances personalization and accuracy by enabling expressive and natural user queries. It establishes a foundation for the next generation of conversational AI-driven search and recommendation systems in digital entertainment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。