arXiv:2605.19151cs.AIcs.HC2026-05

用偏好学习方法自动调节智能体工具使用时的信任水平。

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

  • 将信任校准建模为偏好学习问题,通过高斯过程推断人类风险容忍度。
  • 仅在人类决策最不确定时才请求人工干预,减少不必要审批。
  • 适合需要安全与效率平衡的自动化系统设计者参考。

我们将智能体工具使用中的信任校准(决定自动化建议是否可自主执行或需人工批准)形式化为一个偏好学习问题。策略网关维护一个关于潜在人类风险容忍度函数的高斯过程后验分布,通过二元批准/拒绝反馈的probit似然进行观测,并在审批结果最不确定处触发人工介入。该方法在结构上属于偏好贝叶斯优化,继承其近似高斯过程分类推理机制和以不确定性为导向的高效采样优势,但目标不同:不是优化设计参数,而是对动作空间进行允许/禁止/询问区域的分类。

原文摘要 · Abstract (English)

We formalize trust calibration for agentic tool use (deciding when an automated agent's proposed action may execute autonomously versus require human approval) as a preference-learning problem. A policy gateway maintains a Gaussian-process posterior over a latent human risk-tolerance function, observed through a probit likelihood on binary approve/deny feedback, and escalates to the human exactly where the approval outcome is most uncertain. We show this is structurally an instance of Preferential Bayesian Optimization, inheriting its inference machinery (approximate Gaussian-process classification) and its sample-efficiency argument (uncertainty-targeted querying), while differing in objective: classifying an action space into allow/block/ask regions rather than optimizing a design.

信任校准偏好学习智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。