提出功能型标签定义互补推荐关系,提升标注准确性和模型判断力。
Function-based Labels for Complementary Recommendation: Definition, Annotation, and LLM-as-a-Judge
- 以功能差异定义互补关系,摆脱用户行为数据依赖
- 2759对商品标注数据验证标签覆盖全面且歧义低
- LLM在新标签体系下与人工判断一致性达0.989
互补推荐通过推荐与查询商品功能不同但常一起购买的商品来提升用户体验。推断或评估两个商品是否存在互补关系需要互补关系标签,但此类关系本身存在固有模糊性。基于用户历史行为日志的标签往往不一致且不可靠。近期研究引入大语言模型(LLMs)推断互补关系,但仅提供二元分类,缺乏细粒度理解。本文提出功能型标签(FBLs),一种独立于用户购买记录和LLM黑箱决策过程的新定义。我们构建了一个包含2,759个商品对的人工标注FBL数据集,证明其能覆盖多种商品关系并最小化歧义。进一步评估了使用标注FBL的机器学习方法在未见商品对上的标签推断能力,以及LLM生成标签与人类感知的一致性。结果表明,即使数据有限,逻辑回归和SVM等模型也能达到约0.82的宏平均F1分数。此外,gpt-4o-mini等LLM在FBL定义下表现出0.989的一致性和0.849的分类准确率,表明其可作为有效模拟人类判断的自动标注工具。整体而言,本研究将FBLs作为互补关系的清晰定义,支持更精准的推断与自动化标注。
原文摘要 · Abstract (English)
Complementary recommendations enhance the user experience by suggesting items that are frequently purchased together while serving different functions from the query item. Inferring or evaluating whether two items have a complementary relationship requires complementary relationship labels; however, defining these labels is challenging because of the inherent ambiguity of such relationships. Complementary labels based on user historical behavior logs attempt to capture these relationships, but often produce inconsistent and unreliable results. Recent efforts have introduced large language models (LLMs) to infer these relationships. However, these approaches provide a binary classification without a nuanced understanding of complementary relationships. In this study, we address these challenges by introducing Function-Based Labels (FBLs), a novel definition of complementary relationships independent of user purchase logs and the opaque decision processes of LLMs. We constructed a human-annotated FBLs dataset comprising 2,759 item pairs and demonstrated that it covered possible item relationships and minimized ambiguity. We then evaluated whether some machine learning (ML) methods using annotated FBLs could accurately infer labels for unseen item pairs, and whether LLM-generated complementary labels align with human perception. Our results demonstrate that even with limited data, ML models, such as logistic regression and SVM achieve high macro-F1 scores (approximately 0.82). Furthermore, LLMs, such as gpt-4o-mini, demonstrated high consistency (0.989) and classification accuracy (0.849) under the detailed definition of FBLs, indicating their potential as effective annotators that mimic human judgment. Overall, our study presents FBLs as a clear definition of complementary relationships, enabling more accurate inferences and automated labeling of complementary recommendations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。