用真实互动数据评估大模型理解用户兴趣的能力,发现其在归因统计上存在短板。
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
- 通过行为数据验证用户兴趣,区分虚构与遗漏的偏好
- 小模型仅能正确识别42%的用户兴趣类别,大模型仍有误判
- 适合研究推荐系统与大模型交互的学者参考
我们提出GISTBench,一个用于评估大语言模型(LLMs)在推荐系统中理解用户历史行为能力的基准。不同于传统推荐系统侧重物品预测准确率,本基准关注模型从用户互动数据中提取并验证兴趣的能力。提出两种新指标:兴趣一致性(IG),分解为精确率和召回率以分别惩罚虚构兴趣类别和奖励覆盖范围;兴趣特异性(IS),评估验证后用户画像的独特性。我们构建了一个基于全球短视频平台真实用户互动的合成数据集,包含隐式与显式互动信号及丰富文本描述。通过用户调查验证数据保真度,并评估了8个参数量从7B到120B的开源大模型。结果揭示当前大模型在跨异构互动类型上准确计数与归因方面存在显著瓶颈。
原文摘要 · Abstract (English)
We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike traditional RecSys benchmarks that focus on item prediction accuracy, our benchmark evaluates how well LLMs can extract and verify user interests from engagement data. We propose two novel metric families: Interest Groundedness (IG), decomposed into precision and recall components to separately penalize hallucinated interest categories and reward coverage, and Interest Specificity (IS), which assesses the distinctiveness of verified LLM-predicted user profiles. We release a synthetic dataset constructed on real user interactions on a global short-form video platform. Our dataset contains both implicit and explicit engagement signals and rich textual descriptions. We validate our dataset fidelity against user surveys, and evaluate eight open-weight LLMs spanning 7B to 120B parameters. Our findings reveal performance bottlenecks in current LLMs, particularly their limited ability to accurately count and attribute engagement signals across heterogeneous interaction types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。