arXiv:2604.25136cs.CLcs.AI2026-04被引 1

让大模型学会何时、如何干预以降低认知和规范风险。

Frictive Policy Optimization for LLMs: Epistemic Intervention, Risk-Sensitive Control, and Reflective Alignment

  • 将澄清、质疑等行为作为可控动作,动态调节信念演化
  • 通过干预决策提升下游认知质量,而非仅追求即时奖励
  • 适合追求安全可靠对话的场景,如医疗与法律咨询

我们提出摩擦策略优化(FPO),一种用于训练语言模型策略的框架,使其不仅能决定说什么,还能决定何时以及如何干预,以管理认知与规范风险。不同于传统对齐方法仅优化表面偏好或任务效用,FPO将澄清、验证、挑战、引导和拒绝视为显式控制动作,旨在塑造信念、承诺与不确定性随时间的演变。我们将对齐形式化为风险敏感的认知控制问题,干预决策基于其对下游认知质量的预期影响,而不仅依赖即时奖励。我们引入了摩擦干预的紧凑分类法、结构化摩擦函数以实现多种对齐失败模式,并构建涵盖奖励塑造、偏好配对、组间相对排序与风险条件信任区域的统一FPO方法族。此外,我们提出评估框架,通过澄清行为、校准度、矛盾修复、拒绝合理性与信息效率直接测量认知能力。这些成果共同为学习在结果与认知行为上均对齐的智能体提供了形式化与算法基础。

原文摘要 · Abstract (English)

We propose Frictive Policy Optimization (FPO), a framework for learning language model policies that regulate not only what to say, but when and how to intervene in order to manage epistemic and normative risk. Unlike standard alignment methods that optimize surface-level preference or task utility, FPO treats clarification, verification, challenge, redirection, and refusal as explicit control actions whose purpose is to shape the evolution of belief, commitment, and uncertainty over time. We formalize alignment as a risk-sensitive epistemic control problem in which intervention decisions are selected based on their expected effect on downstream epistemic quality rather than on immediate reward alone. We introduce a compact taxonomy of frictive interventions, a structured friction functional that operationalizes multiple alignment failure modes, and a unified family of FPO methods spanning reward shaping, preference pairing, group-relative ranking, and risk-conditioned trust regions. We further propose an evaluation framework that measures epistemic competence directly through clarification behavior, calibration, contradiction repair, refusal proportionality, and information efficiency. Together, these results provide a formal and algorithmic foundation for learning agents that are aligned not only in outcome, but in epistemic conduct.

大模型对齐认知控制风险敏感干预机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。