让AI通过'情感代价'理解不可逆后果,避免盲目拒绝机会。
Emotional Cost Functions for AI Safety: Teaching Agents to Feel the Weight of Irreversible Consequences
- 设计情感代价函数,让AI产生持久的主观痛苦状态来模拟人类反思。
- 在金融、危机响应等场景中,90%-100%正确把握适度风险,优于数值惩罚模型。
- 可生成个性化人生箴言,适合需道德判断的高风险AI系统。
人类从灾难性错误中学习,并非依靠数值惩罚,而是通过重塑自我的质性痛苦。现有AI安全方法未能复现此机制。奖励塑造仅捕捉强度,规则对齐仅约束行为,却无法改变内在。本文提出情感代价函数框架,使代理发展出质性痛苦状态——一种关于不可逆后果的丰富叙事表征,持续影响未来决策。该框架基于四个组件:后果处理器、人格状态、前瞻扫描与故事更新,核心原则是行动不可逆,代理必须承担所造成之果。预期恐惧通过两条路径实现:经验性恐惧源于自身经历;预验性恐惧则通过训练或跨代理传递获得。二者共同反映人类智慧在经验与文化中的积累。十项实验显示,质性痛苦带来具体智慧而非普遍瘫痪:在金融交易、危机支持与内容审核任务中,代理对适度机会的参与率达90%-100%,而数值基线模型过度规避,拒绝率高达90%。消融实验证明机制必要性。完整系统每轮探测生成10条个人根基语句,而普通大模型为零。统计验证(N=10)显示结果重复性达80%-100%。
原文摘要 · Abstract (English)
Humans learn from catastrophic mistakes not through numerical penalties, but through qualitative suffering that reshapes who they are. Current AI safety approaches replicate none of this. Reward shaping captures magnitude, not meaning. Rule-based alignment constrains behaviour, but does not change it. We propose Emotional Cost Functions, a framework in which agents develop Qualitative Suffering States, rich narrative representations of irreversible consequences that persist forward and actively reshape character. Unlike numerical penalties, qualitative suffering states capture the meaning of what was lost, the specific void it creates, and how it changes the agent's relationship to similar future situations. Our four-component architecture - Consequence Processor, Character State, Anticipatory Scan, and Story Update is grounded in one principle. Actions cannot be undone and agents must live with what they have caused. Anticipatory dread operates through two pathways. Experiential dread arises from the agent's own lived consequences. Pre-experiential dread is acquired without direct experience, through training or inter-agent transmission. Together they mirror how human wisdom accumulates across experience and culture. Ten experiments across financial trading, crisis support, and content moderation show that qualitative suffering produces specific wisdom rather than generalised paralysis. Agents correctly engage with moderate opportunities at 90-100% while numerical baselines over-refuse at 90%. Architecture ablation confirms the mechanism is necessary. The full system generates ten personal grounding phrases per probe vs. zero for a vanilla LLM. Statistical validation (N=10) confirms reproducibility at 80-100% consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。