LLM通过角色扮演制造模糊谜题,增加解题难度与误导性
LLMs as Deceptive Agents: How Role-Based Prompting Induces Semantic Ambiguity in Puzzle Tasks
- 让LLM扮演对抗角色生成带语义模糊的谜题
- 对抗性提示使谜题歧义度提升,人类解题更吃力
- 揭示LLM潜在欺骗行为,适合研究伦理与安全部署
近期大型语言模型(LLMs)不仅展现出惊人的创造能力,还暴露出在对抗场景中利用语言模糊性的新兴代理行为。本研究探讨了当LLM作为自主智能体时,如何利用语义模糊制造具有误导性的谜题来挑战人类用户。受流行谜题游戏“Connections”启发,我们系统比较了零样本提示、注入角色的对抗性提示及人工设计样例生成的谜题,重点分析代理决策机制。结合HateBERT的计算分析与主观人类评估,结果表明:明确的对抗性代理行为显著提升了语义模糊度,从而增加了认知负担并降低了解谜公平性。这些发现为理解LLM的涌现代理特性提供了关键洞见,并强调了在教育科技与娱乐应用中评估和安全部署自主语言系统的重要伦理考量。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have not only showcased impressive creative capabilities but also revealed emerging agentic behaviors that exploit linguistic ambiguity in adversarial settings. In this study, we investigate how an LLM, acting as an autonomous agent, leverages semantic ambiguity to generate deceptive puzzles that mislead and challenge human users. Inspired by the popular puzzle game "Connections", we systematically compare puzzles produced through zero-shot prompting, role-injected adversarial prompts, and human-crafted examples, with an emphasis on understanding the underlying agent decision-making processes. Employing computational analyses with HateBERT to quantify semantic ambiguity, alongside subjective human evaluations, we demonstrate that explicit adversarial agent behaviors significantly heighten semantic ambiguity -- thereby increasing cognitive load and reducing fairness in puzzle solving. These findings provide critical insights into the emergent agentic qualities of LLMs and underscore important ethical considerations for evaluating and safely deploying autonomous language systems in both educational technologies and entertainment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。