arXiv:2605.29420cs.AIcs.LG2026-05被引 2

分析专家角色提示对大模型的影响,发现它提升专业深度但降低清晰度。

When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs

论文配图:When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs
图 1 · 摘自论文原文
  • 通过四种提示条件对比,测试专家角色注入效果
  • 角色提示使专业深度提升,但清晰度下降,结果依赖具体问题类型
  • 混合检索方法优于纯向量检索,适合需要精准角色匹配的场景

Persona prompting 广泛用于引导大语言模型,但其实际价值仍不明确。我们通过控制实验,对比四种提示条件在1,140个开放问题上的表现,覆盖38种专家角色和6个领域:无角色提示、通用领域专家提示、基于嵌入的角色检索,以及结合嵌入搜索与LLM选角的混合方法。总体平均得分差异很小,但细粒度分析揭示了被掩盖的系统性权衡:角色提示显著提升专业深度,但降低表达清晰度。该效应高度依赖任务类型——在医学、心理学等需结构化专家表述与风险沟通的咨询类问题中表现更优;而在金融、法律、科技等需简洁通俗解释的概念性问题上,基线提示反而更优。混合检索方法显著优于仅用嵌入的检索,但无法消除专业深度与清晰度之间的根本权衡。研究结论表明,persona prompting 主要改变响应特征而非普遍提升能力,多维度评估不可或缺。

原文摘要 · Abstract (English)

Persona prompting is widely used to steer large language models, yet its practical value remains unclear. Prior work often evaluates persona prompting using aggregate scores, making it difficult to determine whether expert-role prompting consistently improves response quality or instead changes responses along different quality dimensions. We study this question through a controlled comparison of four prompting conditions across 1,140 open-ended questions spanning 38 expert roles and six domains: no role prompt, a generic domain-expert prompt, embedding-based role retrieval, and a hybrid retrieval method combining embedding search with LLM-based role selection. Aggregate results show only small overall differences between conditions. However, metric-level analysis reveals a consistent tradeoff that aggregate averages obscure: role prompting systematically increases expertise depth while reducing clarity. These effects are highly conditional rather than universal. Role prompting performs best on advisory questions and in domains such as medicine and psychology, where structured expert framing and risk communication are intrinsically valuable. In contrast, baseline prompting performs better on conceptual and explanatory questions in finance, legal, science, and technology domains, where concise plain-language explanation is more important. We further show that hybrid retrieval significantly improves over embedding-only role selection, although better role retrieval does not eliminate the broader expertise-depth versus clarity tradeoff. Overall, our findings suggest that persona prompting primarily reshapes response characteristics rather than broadly improving capability, and that multi-metric evaluation is necessary for understanding its effects.

提示工程专家角色多维评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。