为高风险决策设计五层架构,提升大模型可靠性
Making LLMs Reliable When It Matters Most: A Five-Layer Architecture for High-Stakes Decisions
- 构建五层保护架构,实现人机认知协作与自我纠错
- 实测显示架构可避免认知偏差导致的决策失误
- 适合金融、医疗等需长期可靠决策的高风险领域
当前大语言模型在可验证领域表现优异,但在高风险战略决策中因人类与AI共同存在的认知偏见而可靠性下降,威胁估值合理性与投资可持续性。本报告基于对7个前沿大模型及3个市场导向创业案例的系统性定性评估,在时间压力下发现:仅靠提示工程无法维持稳定协作状态。需通过4阶段初始化与7阶段校准序列构成的五层保护架构,实现偏见自监控、人机对抗挑战、协作状态验证、性能退化检测与利益相关方保护。三项发现:协作状态可通过有序校准实现但需动态维护;架构漂移与上下文耗尽共现时可靠性下降;解散纪律可避免坚持错误方向。跨模型验证显示不同架构存在系统性性能差异。该方法证明人机团队可在高风险决策中形成可防后悔的认知伙伴关系,满足依赖可信决策支持的回报预期。
原文摘要 · Abstract (English)
Current large language models (LLMs) excel in verifiable domains where outputs can be checked before action but prove less reliable for high-stakes strategic decisions with uncertain outcomes. This gap, driven by mutually reinforcing cognitive biases in both humans and artificial intelligence (AI) systems, threatens the defensibility of valuations and sustainability of investments in the sector. This report describes a framework emerging from systematic qualitative assessment across 7 frontier-grade LLMs and 3 market-facing venture vignettes under time pressure. Detailed prompting specifying decision partnership and explicitly instructing avoidance of sycophancy, confabulation, solution drift, and nihilism achieved initial partnership state but failed to maintain it under operational pressure. Sustaining protective partnership state required an emergent 7-stage calibration sequence, built upon a 4-stage initialization process, within a 5-layer protection architecture enabling bias self-monitoring, human-AI adversarial challenge, partnership state verification, performance degradation detection, and stakeholder protection. Three discoveries resulted: partnership state is achievable through ordered calibration but requires emergent maintenance protocols; reliability degrades when architectural drift and context exhaustion align; and dissolution discipline prevents costly pursuit of fundamentally wrong directions. Cross-model validation revealed systematic performance differences across LLM architectures. This approach demonstrates that human-AI teams can achieve cognitive partnership capable of preventing avoidable regret in high-stakes decisions, addressing return-on-investment expectations that depend on AI systems supporting consequential decision-making without introducing preventable cognitive traps when verification arrives too late.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。