不同人物设定影响大模型代码生成,效果因模型而异。
The Librarian Who Refused to Code: Model-Dependent Identity Enactment in LLM Code Generation
- 用四种角色提示词测试代码生成,发现角色影响取决于具体模型。
- 研究型图书管理员角色使代码正确率从0.92降到0.67,引发大量拒绝响应。
- 极简工程师角色让输出减少30%,但未提升代码质量,适合控制输出长度。
生物性角色设定常用于系统提示词,但其对代码生成的影响极少在受控、预注册条件下评估。本研究测试了四种提示条件(无角色、两种工程师角色、一名研究型图书管理员角色),12个代码任务,两个前沿模型,每组5次运行(共480次生成)。角色效应在两模型间存在显著差异。预注册的混合效应分析显示,条件与模型的交互对提供者统计的输出词数有显著影响;事后可见字符量分析也呈现相同趋势。六次GPT-5.5生成因长度限制被截断,单独报告。在Claude Opus上,极简工程师角色使可见输出减少30%(提供者词数减少33%),但未提升正确性;详尽工程师角色增加输出,亦未提高正确性。探索性事后分析显示,图书管理员角色在60次Opus响应中引发55次角色化免责声明,12次真实无代码响应,导致平均正确率从0.92降至0.67。GPT-5.5在59次非截断响应中未表现出此类行为。结果表明,角色设定是模型依赖的行为策略偏见,而非普适的质量干预。研究释放原始生成内容、衍生评分、分析文件、预注册文档及执行日志;端到端测试重评分需未发布任务工具包。
原文摘要 · Abstract (English)
Biographical personas are widely used in system prompts, but their effects on code generation are rarely evaluated under controlled, pre-registered conditions. We tested four prompt conditions (no persona, two engineer personas, and a research-librarian persona), 12 code-generation tasks, two frontier models, and five runs per cell (480 completions). Persona effects differed between the two tested models. Under the pre-registered mixed-effects analysis, the condition-by-model interaction was significant for provider-reported output tokens; a post-hoc visible-character measure showed the same qualitative pattern. Six GPT-5.5 completions were length-capped and are reported separately. On Claude Opus, the minimalist engineer persona reduced visible output by 30% (33% in provider tokens) without improving correctness, while the thorough engineer persona increased output without a correctness gain. In an exploratory post-hoc analysis, the librarian persona elicited in-character disclaimers in 55 of 60 Opus responses and 12 genuine no-code responses, lowering mean correctness from 0.92 to 0.67. GPT-5.5 produced neither behavior in its 59 non-truncated responses. These results are consistent with personas acting as Model-Dependent behavioral-policy biases rather than universal quality interventions. We release raw completions, derived scores, analysis artifacts, a pre-registration document, and an execution gate log; end-to-end test-based rescoring requires an unreleased task harness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。