arXiv:2411.07965cs.CL2024-11ACL被引 5

提出新方法揭示角色扮演大模型的交互幻觉现象

SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing LLMs

  • 通过立场迁移定义交互幻觉,构建可泛化的分析框架
  • 实验证明主流角色扮演训练会掩盖知识导致行为单调
  • 适合研究大模型社会互动与幻觉机制的学者参考

大语言模型(LLMs)的角色扮演能力催生了丰富的交互场景,但现有研究忽视了交互中的幻觉问题,且存在泛化性差和角色一致性判断隐含的问题。受人类行为启发,我们提出一种通用且显式的分析范式,用于揭示跨多元世界观下LLM的交互模式。首先通过立场迁移定义交互幻觉,进而构建SHARP基准,该基准基于常识知识图谱提取关系,并利用LLM自身的幻觉特性模拟多角色交互。大量实验验证了该范式的有效性与稳定性,分析了影响指标的关键因素,并挑战了传统幻觉缓解方案。更广泛地,本工作揭示了当前主流角色扮演后训练方法的根本局限:倾向于将知识隐藏于风格之下,导致行为单调却看似人性化——即交互幻觉。

原文摘要 · Abstract (English)

The advanced role-playing capabilities of Large Language Models (LLMs) have enabled rich interactive scenarios, yet existing research in social interactions neglects hallucination while struggling with poor generalizability and implicit character fidelity judgments. To bridge this gap, motivated by human behaviour, we introduce a generalizable and explicit paradigm for uncovering interactive patterns of LLMs across diverse worldviews. Specifically, we first define interactive hallucination through stance transfer, then construct SHARP, a benchmark built by extracting relations from commonsense knowledge graphs and utilizing LLMs' inherent hallucination properties to simulate multi-role interactions. Extensive experiments confirm our paradigm's effectiveness and stability, examine the factors that influence these metrics, and challenge conventional hallucination mitigation solutions. More broadly, our work reveals a fundamental limitation in popular post-training methods for role-playing LLMs: the tendency to obscure knowledge beneath style, resulting in monotonous yet human-like behaviors - interactive hallucination.

角色扮演幻觉检测大模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。