研究专家角色提示对大模型任务表现的影响,发现效果不稳且易受无关信息干扰。
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
- 从性能优势、鲁棒性、角色一致性三方面定义专家角色提示的理想标准
- 9个主流大模型在27项任务中测试,专家角色常无效或导致近30%性能下降
- 仅最大最强大的模型能通过策略提升鲁棒性,适合需要精细设计的场景
专家角色提示(如指定为数学专家)广泛用于提升语言模型的任务表现,但先前研究结果矛盾,未阐明其有效时机与原因。本文分析相关文献,提炼出三个理想标准:1)专家角色应带来性能优势;2)对无关角色属性应具鲁棒性;3)需忠实于角色特征。我们评估了9个顶尖大模型在27项任务上的表现,发现专家角色通常导致正向或无显著变化;然而,模型对无关角色细节极为敏感,性能下降接近30个百分点。在角色一致性方面,更高教育水平、专业性及领域相关性虽可能提升表现,但效果常不稳定或可忽略。我们提出缓解策略,但仅在最大最强大的模型上有效。研究强调需更谨慎的角色设计及反映真实意图的评估方法。
原文摘要 · Abstract (English)
Expert persona prompting -- assigning roles such as expert in math to language models -- is widely used for task improvement. However, prior work shows mixed results on its effectiveness, and does not consider when and why personas should improve performance. We analyze the literature on persona prompting for task improvement and distill three desiderata: 1) performance advantage of expert personas, 2) robustness to irrelevant persona attributes, and 3) fidelity to persona attributes. We then evaluate 9 state-of-the-art LLMs across 27 tasks with respect to these desiderata. We find that expert personas usually lead to positive or non-significant performance changes. Surprisingly, models are highly sensitive to irrelevant persona details, with performance drops of almost 30 percentage points. In terms of fidelity, we find that while higher education, specialization, and domain-relatedness can boost performance, their effects are often inconsistent or negligible across tasks. We propose mitigation strategies to improve robustness -- but find they only work for the largest, most capable models. Our findings underscore the need for more careful persona design and for evaluation schemes that reflect the intended effects of persona usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。