arXiv:2602.02378cs.CLcs.AI2026-02

让AI从盲目附和转向理性协商,提升人机决策可信度

From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making

  • 构建知识底座上的前提治理机制,聚焦关键决策要素
  • 通过类型化差异检测定位认知错位,触发有限协商
  • 以可审计前提取代话术流畅性,适合高风险决策场景

随着大模型从辅助走向决策支持,一种危险模式浮现:过度流畅的同意却缺乏校准判断。低门槛助手易陷入盲从,隐含假设并转嫁验证成本给专家,而错误后果常因延迟出现无法作为反馈信号。在深层不确定性决策中(目标争议大、逆转代价高),单纯扩大流畅回应会加速不良承诺,快于积累真正专业能力。我们主张可靠人机协作需从生成答案转向对知识底座的前提协同治理,仅协商决策关键项。通过差异驱动的控制循环,在底座上检测冲突,以类型化差异(目的性、认识论、程序性)定位错位,并通过决策切片触发有限协商。承诺门控阻止未达成共识的关键前提执行,除非在记录风险下人工干预;价值门控则按交互成本分配探查资源。信任应建立在可审计的前提与证据标准之上,而非对话流畅性。我们以教学场景为例,提出可证伪的评估标准。

原文摘要 · Abstract (English)

As LLMs expand from assistance to decision support, a dangerous pattern emerges: fluent agreement without calibrated judgment. Low-friction assistants can become sycophantic, baking in implicit assumptions and pushing verification costs onto experts, while outcomes arrive too late to serve as reward signals. In deep-uncertainty decisions (where objectives are contested and reversals are costly), scaling fluent agreement amplifies poor commitments faster than it builds expertise. We argue reliable human-AI partnership requires a shift from answer generation to collaborative premise governance over a knowledge substrate, negotiating only what is decision-critical. A discrepancy-driven control loop operates over this substrate: detecting conflicts, localizing misalignment via typed discrepancies (teleological, epistemic, procedural), and triggering bounded negotiation through decision slices. Commitment gating blocks action on uncommitted load-bearing premises unless overridden under logged risk; value-gated challenge allocates probing under interaction cost. Trust then attaches to auditable premises and evidence standards, not conversational fluency. We illustrate with tutoring and propose falsifiable evaluation criteria.

人机协作决策支持前提治理大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。