arXiv:2502.11789cs.CL2025-02Conference of the …被引 1

用12个样本即可精准调整大模型人格,无需大量训练。

Personality Editing for Language Models through Adjusting Self-Referential Queries

  • 通过自指查询模拟心理构念,直接编辑模型人格
  • 仅需12个样本即显著提升人格一致性与平衡性
  • 适合需要快速定制对话风格的开发者和产品团队

大语言模型在对话代理和内容生成中广泛应用,精确控制其人格对保持语调一致性和用户参与度至关重要。现有基于提示或微调的方法或缺乏鲁棒性,或需大规模训练数据,成本高昂且不实用。本文提出PALETTE(Personality Adjustment by LLM Self-Targeted Queries),一种新型人格编辑方法。该方法引入调整查询,将基于心理学构念的自指陈述类比为事实知识,实现对人格相关回应的直接编辑。相较于微调,PALETTE仅需12个编辑样本即可在多个性格维度上实现显著的人格对齐提升。自动与人工评估结果均表明,该方法能实现更稳定、更均衡的人格控制。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are integral to applications such as conversational agents and content creation, where precise control over a model's personality is essential for maintaining tone, consistency, and user engagement. However, prevailing prompt-based or fine-tuning approaches either lack robustness or demand large-scale training data, making them costly and impractical. In this paper, we present PALETTE (Personality Adjustment by LLM SElf-TargeTed quEries), a novel method for personality editing in LLMs. Our approach introduces adjustment queries, where self-referential statements grounded in psychological constructs are treated analogously to factual knowledge, enabling direct editing of personality-related responses. Unlike fine-tuning, PALETTE requires only 12 editing samples to achieve substantial improvements in personality alignment across personality dimensions. Experimental results from both automatic and human evaluations demonstrate that our method enables more stable and well-balanced personality control in LLMs.

人格控制提示工程少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。