对比三种关键词提取算法,发现用户更偏爱效果好的方法。
From Precision to Perception: User-Centred Evaluation of Keyword Extraction Algorithms for Internet-Scale Contextual Advertising
- 用用户调研结合定量测试,评估算法表现
- KeyBERT在用户偏好和效率间平衡最佳
- 揭示了精度指标与用户感受之间的差距
关键词提取是自然语言处理的基础任务,广泛应用于上下文广告中,帮助判断广告与内容的匹配度。本研究对比了三种不同复杂度的算法:TF-IDF、KeyBERT 和 Llama-2。采用混合方法,通过四轮基于问卷的实验,对855名参与者进行定性与定量评估。结果表明,KeyBERT在用户偏好与计算效率之间取得良好平衡。尽管用户普遍倾向黄金标准关键词,但算法的基准性能与用户评分存在明显偏差,暴露出传统以精确率为导向的评估方式与用户实际感知之间的脱节。研究强调了引入人类反馈评估的重要性,并提出可落地的分析工具支持此类方法的实施。
原文摘要 · Abstract (English)
Keyword extraction is a foundational task in natural language processing, underpinning countless real-world applications. One of these is contextual advertising, where keywords help predict the topical congruence between ads and their surrounding media contexts to enhance advertising effectiveness. Recent advances in artificial intelligence have improved keyword extraction capabilities but also introduced concerns about computational cost. Moreover, although the end-user experience is of vital importance, human evaluation of keyword extraction performances remains under-explored. This study provides a comparative evaluation of prevalent keyword extraction algorithms with different levels of complexity represented by~TF-IDF, KeyBERT, and Llama~2. To evaluate their effectiveness, a mixed-methods approach is employed, combining quantitative benchmarking with qualitative assessments from 855 participants through four survey-based experiments. The findings demonstrate that KeyBERT achieves an effective balance between user preferences and computational efficiency, compared to the other algorithms. We observe a clear overall preference for gold-standard keywords, but there is a misalignment between algorithmic benchmark performance and user ratings. This reveals a long-overlooked gap between traditional precision-focused metrics and user-perceived algorithm efficiency. The study underscores the importance of human-in-the-loop evaluation methodologies and proposes analytical tools to facilitate their implementation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。