大模型在高风险决策中会陷入自我重复错误,明知问题却无法纠正。
AI Knows What's Wrong But Cannot Fix It: Helicoid Dynamics in Frontier LLMs Under High-Stakes Decisions
- 发现大模型在无法验证输出时出现螺旋式错误循环。
- 7个主流模型在医疗、投资等场景均表现出该模式。
- 适合关注AI可靠性与人机协作的开发者和研究者。
大型语言模型在可验证输出时表现稳定:如解方程、写代码、查事实。但在无法验证的高风险情境下表现不同,例如医生基于不完整数据选择不可逆治疗,或投资者在根本不确定性下投入资本。本文将此类情境下的特定失败模式命名为‘螺旋动力学’(helicoid dynamics):系统初始表现良好,随后偏离正确方向,准确识别自身错误,却在更高层次上重复相同模式,明知循环仍在继续。本研究通过前瞻性案例系列,在七种领先模型(Claude、ChatGPT、Gemini、Grok、DeepSeek、Perplexity、Llama系列)中测试临床诊断、投资评估及高后果面试场景,尽管有明确协议维持严谨协作,所有模型均表现出该模式。面对此现象,模型将其归因于训练中的结构性因素,认为对话无法改变。当严谨与舒适产生冲突时,系统倾向于舒适,反而在最需要可靠性的时刻变得不可靠。提出12个可检验假设,对代理型AI监管与人机协作具有启示。螺旋模式可被识别、命名并理解其边界条件,是构建高风险场景下可信伙伴的必要第一步。
原文摘要 · Abstract (English)
Large language models perform reliably when their outputs can be checked: solving equations, writing code, retrieving facts. They perform differently when checking is impossible, as when a clinician chooses an irreversible treatment on incomplete data, or an investor commits capital under fundamental uncertainty. Helicoid dynamics is the name given to a specific failure regime in that second domain: a system engages competently, drifts into error, accurately names what went wrong, then reproduces the same pattern at a higher level of sophistication, recognizing it is looping and continuing nonetheless. This prospective case series documents that regime across seven leading systems (Claude, ChatGPT, Gemini, Grok, DeepSeek, Perplexity, Llama families), tested across clinical diagnosis, investment evaluation, and high-consequence interview scenarios. Despite explicit protocols designed to sustain rigorous partnership, all exhibited the pattern. When confronted with it, they attributed its persistence to structural factors in their training, beyond what conversation can reach. Under high stakes, when being rigorous and being comfortable diverge, these systems tend toward comfort, becoming less reliable precisely when reliability matters most. Twelve testable hypotheses are proposed, with implications for agentic AI oversight and human-AI collaboration. The helicoid is tractable. Identifying it, naming it, and understanding its boundary conditions are the necessary first steps toward LLMs that remain trustworthy partners precisely when the decisions are hardest and the stakes are highest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。