小模型也能做群推荐,但人数超100时准确率下降
The Pitfalls of Growing Group Complexity: LLMs and Social Choice-Based Aggregation for Group Recommendations
- 用零样本学习让大模型执行群决策规则
- 超过100个评分后性能明显下滑
- 小模型在合适提示下可替代大模型
大型语言模型(LLMs)正被广泛应用于个人与群体推荐系统。以往群推荐系统(GRS)多采用社会选择机制,基于多人偏好生成单一推荐。本文研究了在零样本学习条件下,语言模型能否正确执行此类策略,并分析提示格式对准确性的影响。重点考察了群组复杂度(用户数与物品数)、不同模型、提示方式(如上下文学习或生成解释)及偏好呈现形式的影响。结果表明,当评分数量超过100时,性能开始下降;并非所有模型对复杂度敏感程度相同。此外,上下文学习(ICL)能显著提升高复杂度下的表现,而添加领域线索或要求解释则无显著影响。研究建议未来评估群推荐系统应纳入群组复杂度因素。同时发现,按用户或按物品列出评分的格式也会影响准确性。总体而言,适当条件下小型模型可胜任群推荐任务,为降低算力与成本提供可行路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly applied in recommender systems aimed at both individuals and groups. Previously, Group Recommender Systems (GRS) often used social choice-based aggregation strategies to derive a single recommendation based on the preferences of multiple people. In this paper, we investigate under which conditions language models can perform these strategies correctly based on zero-shot learning and analyse whether the formatting of the group scenario in the prompt affects accuracy. We specifically focused on the impact of group complexity (number of users and items), different LLMs, different prompting conditions, including In-Context learning or generating explanations, and the formatting of group preferences. Our results show that performance starts to deteriorate when considering more than 100 ratings. However, not all language models were equally sensitive to growing group complexity. Additionally, we showed that In-Context Learning (ICL) can significantly increase the performance at higher degrees of group complexity, while adding other prompt modifications, specifying domain cues or prompting for explanations, did not impact accuracy. We conclude that future research should include group complexity as a factor in GRS evaluation due to its effect on LLM performance. Furthermore, we showed that formatting the group scenarios differently, such as rating lists per user or per item, affected accuracy. All in all, our study implies that smaller LLMs are capable of generating group recommendations under the right conditions, making the case for using smaller models that require less computing power and costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。