让大模型稳定扮演人类特质,无需微调即可跨任务保持一致。
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
- 通过自生成反思文本并优化,实现上下文内自我修正的特质诱导。
- 单次生成的反思可使模型在多任务中稳定表现目标特质,超越现有方法。
- 适合个性化对话、社会模拟等需一致人格表现的应用场景。
基于人类撰写的语料训练的大语言模型(LLMs)可通过提示词展现特定类人特质(如性格或价值观),在个性化模型和社交模拟中具有应用价值。然而,现有方法存在表层诱导问题:模型仅能模仿浅层、不稳定的风格模式,难以在多样化任务中精准且一致地体现目标特质。为此,我们提出 IROTE,一种无需微调的上下文内自我反思优化方法。受心理学理论启发,该方法自动生成并优化提示中的文本化自我反思,包含自我感知经验,以激发模型的特质驱动行为。优化过程通过迭代最大化信息论目标,增强模型行为与目标特质间的关联,同时减少反思中的冗余噪声。大量实验表明,单一 IROTE 生成的反思即可使模型在三种人类特质体系下,跨多种下游任务(不仅限于问卷回答)稳定呈现目标特质,性能持续优于现有强基线。
原文摘要 · Abstract (English)
Trained on various human-authored corpora, Large Language Models (LLMs) have demonstrated a certain capability of reflecting specific human-like traits (e.g., personality or values) by prompting, benefiting applications like personalized LLMs and social simulations. However, existing methods suffer from the superficial elicitation problem: LLMs can only be steered to mimic shallow and unstable stylistic patterns, failing to embody the desired traits precisely and consistently across diverse tasks like humans. To address this challenge, we propose IROTE, a novel in-context method for stable and transferable trait elicitation. Drawing on psychological theories suggesting that traits are formed through identity-related reflection, our method automatically generates and optimizes a textual self-reflection within prompts, which comprises self-perceived experience, to stimulate LLMs' trait-driven behavior. The optimization is performed by iteratively maximizing an information-theoretic objective that enhances the connections between LLMs' behavior and the target trait, while reducing noisy redundancy in reflection without any fine-tuning, leading to evocative and compact trait reflection. Extensive experiments across three human trait systems manifest that one single IROTE-generated self-reflection can induce LLMs' stable impersonation of the target trait across diverse downstream tasks beyond simple questionnaire answering, consistently outperforming existing strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。