用用户行为数据构建个性化问答评估基准
CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors

- 基于用户与社区偏好差异,提炼六类认知因素作为评估维度
- 构建包含1985个用户画像的基准,实现细粒度个性化评估
- 适合研究个性化大模型与用户行为建模的学者使用
尽管大语言模型在问答任务中展现出巨大潜力,但个性化评估仍是关键瓶颈。现有方法多依赖词汇层面相似性或人工规则,缺乏充分的数据驱动验证。本文通过挖掘社区-个体偏好差异(CIPD),即个体选择超越共识的现象,提炼出六项关键个性化因素作为评估维度。据此提出CoPA基准,包含1,985个用户画像,可对模型输出与用户特定认知偏好之间的匹配度进行量化评估,提供比通用指标更全面、更具区分性的个性化问答评价标准。代码已公开于https://github.com/bjzgcai/CoPA。
原文摘要 · Abstract (English)
While LLMs have demonstrated remarkable potential in Question Answering (QA), evaluating personalization remains a critical bottleneck. Existing paradigms predominantly rely on lexical-level similarity or manual heuristics, often lacking sufficient data-driven validation. We address this by mining Community-Individual Preference Divergence (CIPD), where individual choices override consensus, to distill six key personalization factors as evaluative dimensions. Accordingly, we introduce CoPA, a benchmark with 1,985 user profiles for fine-grained, factor-level assessment. By quantifying the alignment between model outputs and user-specific cognitive preferences inferred from interaction patterns, CoPA provides a more comprehensive and discriminative standard for evaluating personalized QA than generic metrics. The code is available at https://github.com/bjzgcai/CoPA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。