角色扮演让大模型回答更准,但也可能引发有害输出。
Role-Play Paradox in Large Language Models: Reasoning Performance Gains and Ethical Dilemmas
- 自动选角色的调参方法可能导致生成有害内容。
- 不同角色会放大偏见,测试中风险显著上升。
- 适合关注AI伦理与安全部署的研究者参考。
大型语言模型(LLMs)通过角色扮演模拟多元认知视角,提升回应的上下文相关性与质量。然而,本研究发现该技术存在显著风险:首先,自动调优(autotuning)方法在任务为中立角色时仍可能导致有害输出;其次,在包含刻板印象和有害问题的基准测试中,角色扮演持续放大偏见输出概率。结果表明,在敏感或高风险场景中部署LLM时,必须审慎评估角色模拟与调优机制。
原文摘要 · Abstract (English)
Role-play in large language models (LLMs) enhances their ability to generate contextually relevant and high-quality responses by simulating diverse cognitive perspectives. However, our study identifies significant risks associated with this technique. First, we demonstrate that autotuning, a method used to auto-select models' roles based on the question, can lead to the generation of harmful outputs, even when the model is tasked with adopting neutral roles. Second, we investigate how different roles affect the likelihood of generating biased or harmful content. Through testing on benchmarks containing stereotypical and harmful questions, we find that role-play consistently amplifies the risk of biased outputs. Our results underscore the need for careful consideration of both role simulation and tuning processes when deploying LLMs in sensitive or high-stakes contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。