让人类与AI对话时能互相提醒,避免错误累积。
From Consumption to Reflection: Designing Human-AI Relations for Stable Reasoning
- 在对话中插入可审计的反思环节,不改模型本身
- 识别推理断裂点,关键时刻主动引入反思
- 适合需要高可靠性决策的场景,如医疗、法律
大语言模型(LLMs)提升了信息获取效率,却未改善人类的推理方式。它们的流畅性加速了信息消费,却跳过了形成可靠判断所需的缓慢反思过程。本文提出关系式反思智能(RRI),一种运行于模型之外的推理期治理层,通过可审计的推理循环实现反思。核心观点是:LLMs继承了人类认知的缺陷——依赖直觉捷径、混淆表征与现实、偏好一致性而非证伪。当人与模型共享这些倾向时,错误会相互放大,称为关系性漂移,其根源在于交互而非模型本身。解决此问题需从关注词语间关系转向构建模型输出与人类推理间的结构化关系。RRI通过三个组件实现:玫瑰框架(Rose-Frame)识别推理断裂点;建筑师之笔(Architect's Pen)在关键节点引入反思步骤;推理期工作流将这些步骤嵌入流程,无需重训练模型。三者共同将人机交互转化为具有显式检查点、冲突暴露和可审计假设链的联合推理系统。RRI不强求机器像人思考或让人像机器推理,而是建立一种互补的结构化互动,将AI安全重构为认知架构问题,即可靠的决策依赖于将反思直接嵌入交互过程。
原文摘要 · Abstract (English)
Large language models (LLMs) have transformed how humans access information, but not how we reason with it. Their fluency accelerates consumption while bypassing the slow, reflective processes that underpin sound judgment. This paper introduces Relational Reflective Intelligence (RRI), an inference-time governance layer that operationalizes reflection through auditable reasoning loops. RRI operates not inside the model but around it, providing a practical structure for stable, auditable reasoning between humans and LLMs. The core premise is that LLMs inherit cognitive vulnerabilities similar to those that shape human thought: reliance on intuitive shortcuts, confusion between representation and reality, and a preference for coherence over falsification. When humans and models share these tendencies, their errors compound. We refer to this as relational drift, a failure that arises from interaction rather than from the model alone. Addressing this requires a shift from modeling relations between words to structuring relations between model outputs and human reasoning. RRI provides this missing layer through three components: the Rose-Frame, which identifies likely breakdowns in reasoning; the Architect's Pen, which introduces targeted reflection steps at critical moments; and an inference-time workflow that embeds these steps without retraining the model. Together, these elements transform human-AI interaction into a joint reasoning system with explicit checkpoints, conflict surfacing, and an auditable trail of assumptions. Rather than making machines think like humans or forcing humans to reason like machines, RRI creates a structured interaction in which both compensate for each other's limitations. It reframes AI safety as a cognitive architecture problem, where reliable decisions depend on embedding reflection directly into the interaction process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。