测试多语言提示对大模型推荐系统的影响,发现非英语提示性能下降。
Multilingual Prompts in LLM-Based Recommenders: Performance Across Languages
- 扩展英文提示模板为西语和土耳其语,评估跨语言表现。
- 土耳其语提示性能显著低于英语,但多语言训练使各语言表现更均衡。
- 适合关注多语言推荐、模型公平性的研究者参考。
大型语言模型(LLMs)在自然语言处理任务中应用日益广泛。推荐系统传统上采用协同过滤、矩阵分解等方法,近年也引入深度学习与强化学习技术。尽管已有研究将语言模型用于推荐,但近期趋势聚焦于利用LLM的生成能力实现更个性化的推荐。现有研究多集中于资源丰富的英语,本文探讨非英语提示对推荐性能的影响。通过OpenP5平台,将英文提示模板扩展至西班牙语和土耳其语,并在ML1M、LastFM和Amazon-Beauty三个真实数据集上进行评估。结果显示,使用非英语提示通常会降低性能,尤其在资源较少的土耳其语中表现更差。此外,我们使用多语言提示重新训练了基于LLM的推荐模型,结果表明多语言训练可提升各语言间的性能平衡性,但略微降低了英语表现。本研究强调了在基于LLM的推荐系统中支持多样化语言的重要性,并建议未来研究应构建多语言评估数据集、使用更新模型及更多语言。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in natural language processing tasks. Recommender systems traditionally use methods such as collaborative filtering and matrix factorization, as well as advanced techniques like deep learning and reinforcement learning. Although language models have been applied in recommendation, the recent trend have focused on leveraging the generative capabilities of LLMs for more personalized suggestions. While current research focuses on English due to its resource richness, this work explores the impact of non-English prompts on recommendation performance. Using OpenP5, a platform for developing and evaluating LLM-based recommendations, we expanded its English prompt templates to include Spanish and Turkish. Evaluation on three real-world datasets, namely ML1M, LastFM, and Amazon-Beauty, showed that usage of non-English prompts generally reduce performance, especially in less-resourced languages like Turkish. We also retrained an LLM-based recommender model with multilingual prompts to analyze performance variations. Retraining with multilingual prompts resulted in more balanced performance across languages, but slightly reduced English performance. This work highlights the need for diverse language support in LLM-based recommenders and suggests future research on creating evaluation datasets, using newer models and additional languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。