构建100人多日对话数据集,评测个性化文本生成效果
YNTP-100: A Benchmark for Your Next Token Prediction with 100 People
- 将个性化生成转为基于历史交互的逐标记预测任务
- 在100人多语言对话数据上验证,实现风格一致与内容匹配
- 适合研究个性化大模型、对话系统与隐私保护应用
通用大语言模型在进行下一项词预测时,常无法体现特定个体的表达风格。个性化对齐进展受限于真实个人通信数据收集难,主要受隐私约束。我们提出「你的下一个词预测」(YNTP)任务,将个性化回复生成建模为基于用户交互历史的逐标记预测。我们构建了「YNTP-100」基准,包含100名用户的多语言、多日人机对话数据,支持对用户特定响应行为的系统性评估。采用内容相似性与风格一致性指标,评估外部(参数不变)和内部(参数更新)对齐方法。数据集与结果已公开:https://github.com/AnonymousHub4Submissions/YNTP100。
原文摘要 · Abstract (English)
Large language models (LLMs) trained for general \textit{next-token prediction} often fail to generate responses that reflect how specific individuals communicate. Progress on personalized alignment is further limited by the difficulty of collecting real-world personal communication data due to privacy constraints. We propose Your Next Token Prediction (YNTP), a task that formulates personalized response generation as token-level prediction conditioned on user interaction history. We introduce \textbf{YNTP-100}, a benchmark built from multilingual multi-day human--agent conversations with 100 people, enabling systematic evaluation of user-specific response behavior. We evaluate external (parameter-preserving) and internal (parameter-updating) alignment methods using metrics of substance similarity and stylistic consistency. The dataset and results are publicly available at: https://github.com/AnonymousHub4Submissions/YNTP100.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。