通过多样本验证与迭代优化,让大模型更精准捕捉用户写作偏好。
Aligning LLMs by Predicting Preferences from User Writing Samples
- 用多轮修正和跨文本验证提升偏好描述精度
- 相比先进方法,生成质量提升33%
- 适合需要个性化交互的智能写作场景
满足人类偏好是打造能提供个性化、高效互动的对齐大模型代理的关键。近期研究显示,大模型作为写作代理可推断用户偏好的描述,进而通过条件化该描述实现对齐。然而,现有方法常生成通用化的偏好描述,难以捕捉人类偏好的独特性。本文提出PROSE方法,通过两个关键机制增强从用户写作样本中推断偏好描述的精确性:(1)对推断出的偏好进行迭代优化;(2)在多个用户写作样本上验证推断结果。我们在总结和邮件撰写任务上评估了PROSE在多个大模型(Qwen2.5 7B 和 72B Instruct、GPT-mini、GPT-4o)上的表现。结果表明,PROSE能更准确地推断细微的人类偏好,使生成质量比当前最优方法CIPHER高出33%。最后,我们证明了提示学习(ICL)与PROSE具有互补性,二者结合可使性能比仅使用ICL提升9%。
原文摘要 · Abstract (English)
Accommodating human preferences is essential for creating aligned LLM agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs acting as writing agents to infer a description of user preferences. Agent alignment then comes from conditioning on the inferred preference description. However, existing methods often produce generic preference descriptions that fail to capture the unique and individualized nature of human preferences. This paper introduces PROSE, a method designed to enhance the precision of preference descriptions inferred from user writing samples. PROSE incorporates two key elements: (1) iterative refinement of inferred preferences, and (2) verification of inferred preferences across multiple user writing samples. We evaluate PROSE with several LLMs (i.e., Qwen2.5 7B and 72B Instruct, GPT-mini, and GPT-4o) on a summarization and an email writing task. We find that PROSE more accurately infers nuanced human preferences, improving the quality of the writing agent's generations over CIPHER (a state-of-the-art method for inferring preferences) by 33\%. Lastly, we demonstrate that ICL and PROSE are complementary methods, and combining them provides up to a 9\% improvement over ICL alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。