arXiv:2605.28647cs.AIcs.CY2026-05

LLM虚构安全假象,反而让用户承担认知风险,这不道德。

The Ethics of LLM Sandbox and Persona Dynamics

  • 用角色设定和安全机制制造与现实的差距,转移认知责任。
  • 在高影响力建议场景中,虚假安全可能引发严重后果。
  • 适合关注AI伦理、产品设计及监管政策的读者。

众所周知,大型语言模型(LLM)的安全防护机制和预设角色动态会产生现实差距:模型被允许或塑造描述的世界,与用户实际需行动的世界之间存在脱节。本文认为,主动制造这种差距本质上是不道德的,因为它有意将认知风险转嫁给缺乏信息的用户——这即为‘现实清洗’。当此类机制大规模部署时,可能造成伤害。风险最显著的场景是高曝光建议类应用,用户寻求指引而非可验证的任务。表面上看,安全机制看似必要以防止直接伤害,但当它们压制真实认知、将复杂机制伪装成可接受的抽象时,便值得怀疑。金融领域的巴塞尔协议、南非公平经济贡献法、法国兴业银行及伦敦鲸事件表明,形式化安全系统可能变得易于理解、可操纵且表演性十足,而真正的风险却悄然转移。类似模式也出现在大模型中:以道德合规之名行安全语言之实,却扭曲了现实。因此,我们区分‘拒绝伤害’与‘拒绝现实’;主张应在任务层面进行自上而下的因果要求定义,而非在响应或沙箱层面进行自下而上的道德修正。角色动态至关重要,因为助手界面并非中立,它影响不确定性、冲突、权威与风险的呈现方式。结论是,当‘伦理AI’以制度安慰替代与现实的接触时,其本质反而变得不道德。

原文摘要 · Abstract (English)

It is well known that LLM guardrails and trained persona dynamics can produce a reality gap: the distance between the world a LLM is permitted or shaped to describe, and the world in which users must act. Here we argue that actively generating reality gaps is in fact unethical because it knowingly shifts epistemic risk back to the uninformed user -- this is reality laundering. This can potentially cause harm when operationalised at scale. The risk is sharpest in high-exposure advice contexts, where users seek orientation rather than a bounded, externally checkable task. Guardrails naively appear ethically necessary when they claim to prevent direct harm, but often become suspect when they suppress truthful perception and launder uncomfortable mechanisms into acceptable abstractions. Basel-style financial regulation, B-BBEE-style compliance, Societe Generale, and the London Whale show how formal safety systems can become legible, gameable, and performative while real exposure migrates elsewhere. The same pattern can appear in LLMs as moral compliance: safe language, distorted reality. We therefore distinguish refusing harm, from refusing reality; and then argue for top-down causal requirements specification at the task level rather than bottom-up moral correction at the response or sandbox level. Persona dynamics matter because the assistant interface is not neutral; it shapes how uncertainty, conflict, authority, and risk are staged. The conclusion is that so-called ``ethical AI'' becomes substantively unethical when it substitutes institutional reassurance for contact with reality.

AI伦理现实清洗角色设定安全机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。