arXiv:2506.02659cs.CL2025-06EMNLP被引 4

研究大模型在不同角色下的表现一致性,发现任务越结构化,表现越稳定。

Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs

  • 构建新框架评估大模型在多任务中保持角色一致性的能力。
  • 结构化任务和额外上下文能显著提升角色一致性表现。
  • 模型设计与刻板印象会影响角色扮演的稳定性,适合角色扮演研究者参考。

个性化大语言模型(LLM)在各类应用中被赋予特定人物形象(如快乐的高中教师),以引导其回答。尽管已有研究关注模型在写作风格上对预设角色的遵循程度,但缺乏对不同角色和任务类型下一致性表现的全面分析。本文提出一种新的标准化框架,用于评估角色赋予模型的一致性,定义为同一角色在不同任务和运行中保持连贯响应的程度。该框架涵盖四类角色特征(幸福感、职业、性格、政治立场),覆盖五种任务维度(问卷写作、论文生成、社交媒体帖子生成、单轮对话、多轮对话)。研究发现,一致性受角色设定、刻板印象及模型设计选择的影响,且在更结构化的任务中更高,附加上下文可进一步增强。所有代码已开源。

原文摘要 · Abstract (English)

Personalized Large Language Models (LLMs) are increasingly used in diverse applications, where they are assigned a specific persona - such as a happy high school teacher - to guide their responses. While prior research has examined how well LLMs adhere to predefined personas in writing style, a comprehensive analysis of consistency across different personas and task types is lacking. In this paper, we introduce a new standardized framework to analyze consistency in persona-assigned LLMs. We define consistency as the extent to which a model maintains coherent responses when assigned the same persona across different tasks and runs. Our framework evaluates personas across four different categories (happiness, occupation, personality, and political stance) spanning multiple task dimensions (survey writing, essay generation, social media post generation, single turn, and multi-turn conversations). Our findings reveal that consistency is influenced by multiple factors, including the assigned persona, stereotypes, and model design choices. Consistency also varies across tasks, increasing with more structured tasks and additional context. All code is available on GitHub.

大模型角色一致性persona评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。