首个基于用户长期行为日志的个性化大模型评测基准
LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

- 用多阶段合成方法构建跨领域行为日志数据集
- 实验证明行为日志是个性化关键,但需合理选择与整合证据
- 适合研究个性化大模型、行为建模与隐私保护的学者
现有个性化大模型评测主要依赖文本人格或孤立行为信号,难以评估跨领域行为个性化能力。为此,我们提出LUNAR,首个基于纵向应用交互历史、覆盖服装、饮食、住房、出行等通用生活领域的个性化大模型评测基准。为支持可扩展构建并缓解数据稀疏与隐私问题,LUNAR采用基于真实行为模式的多阶段粗到细合成流程。保真度分析显示,其行为分布更贴近真实数据。对19个主流大模型的实验表明:访问行为日志是深度个性化的必要条件,但非充分条件;更多上下文或更大模型并不保证更好性能;有效个性化依赖于跨域相关证据的选择与融合。直接检索细粒度行为记录优于压缩记忆,但更强个性化可能牺牲隐私。这些发现揭示了证据选择、跨域整合与隐私控制是个性化大模型的关键挑战。
原文摘要 · Abstract (English)
Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including clothing, food, housing, and mobility. To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns. Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks. Experiments on 19 mainstream LLMs show that access to behavioral logs is necessary but not sufficient for deep personalization: neither more context nor larger models guarantees better performance; effective personalization depends on selecting and integrating relevant evidence across domains. Direct retrieval of fine-grained behavioral records consistently outperforms compressed memory, while stronger personalization can come at the cost of privacy protection. These findings identify evidence selection, cross-domain integration, and privacy control as key challenges for personalized LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。