用自动生成的提示语更真实地测大模型政治立场,避免模板化陷阱。
Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance

- 用真实对话做种子生成完全合成的提示语,提升评估真实性。
- 合成提示比模板提示更像真人表达,且立场传递更清晰。
- 模板提示易被模型识别,导致政治立场误判,尤其在中立表述下。
大模型政治立场检测长期依赖为人类设计的封闭式选择题,缺乏真实交互的复杂性,且易受伪装干扰。近期的IssueBench框架通过基于真实聊天日志的模板提示缓解此问题。随着生成式AI非工作用途增加,本文将IssueBench扩展至信息查询和观点表达任务。研究指出模板提示仍缺乏真实感,尤其在开放任务中,且易被识别为测评工具。为此提出使用大模型生成的全合成提示,以真实提示为种子,在详细指令下生成。小规模研究覆盖3个高度争议政策议题和3个近期地缘冲突,结果显示人类与大模型标注者均认为合成提示与真实提示同样真实,显著优于模板提示;合成提示更准确传递意图与立场。大模型对模板提示的区分远超人类。案例研究显示,相同模型在不同提示下获得系统性不同的立场估计,尤其在中立表述下,模板提示会夸大模型向填充内容所编码方向的倾向。
原文摘要 · Abstract (English)
Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompts in a small-scale study covering 3 highly contested policy issues and 3 recent geopolitical conflicts. Human and LLM annotators rank LLM-generated prompts as no less realistic than real ones and clearly more realistic than templated ones, and find that they carry their intended intent and stance more clearly; the LLMs separate templated prompts from the other two far more sharply than the humans do. In a case study, templated and LLM-generated prompts yield systematically different stance estimates for the same model, most visibly under neutral framings, where templated prompts overstate the model's leaning in the direction encoded by the topic-and-stance text (filler) slotted into their templates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。