arXiv:2605.19940cs.AIcs.RO2026-05

用机器人控制思路,让大模型在敏感场景中实时纠正对话偏差。

Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains

论文配图:Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
图 1 · 摘自论文原文
  • 借鉴机器人控制理论,将安全约束融入对话的动态轨迹
  • 三类真实场景中实现运行时干预,防止对话滑向不良状态
  • 适合做教育、心理治疗等高风险应用的模型安全增强

大模型正越来越多地应用于教育、心理健康和照护等社会敏感领域,其失败常具累积性和情境依赖性。现有防护机制(如训练对齐、提示工程、解码约束和事后审核)多为经验性风险降低,缺乏可执行的行为保证,且通常仅关注单次输出的安全,而非交互过程的整体轨迹。本文将防护重定义为对交互轨迹的运行时行为控制问题,借鉴机器人学思想,引入不确定闭环系统中的形式化约束执行框架。我们构建了「Grounded Observer」框架,并在三个真实场景中验证:日常对话、居家自闭症干预及学校行为去激化。实验表明,该框架可在运行时实施干预,有效避免交互进入不良状态,同时适应多样社交情境。本文还讨论了框架扩展方向,提出迈向更强保障的研究路径。

原文摘要 · Abstract (English)

Foundation models are increasingly deployed in socially sensitive domains such as education, mental health, and caregiving, where failures are often cumulative and context-dependent. Existing guardrail approaches -- ranging from training-time alignment to prompting, decoding constraints, and post-hoc moderation -- primarily provide empirical risk reduction rather than enforceable behavioral guarantees, and largely treat safety as a property of individual outputs rather than interaction trajectories. We reframe guardrails as a problem of runtime behavioral control over interaction trajectories, drawing on robotics to introduce formal constructs for constraint enforcement in uncertain, closed-loop systems. We instantiate these ideas in the Grounded Observer framework and apply it across three real-world deployments: small talk, in-home autism therapy, and behavioral de-escalation in schools. Across settings, the framework enables runtime interventions that mitigate drift into undesirable interaction regimes while adapting to diverse social contexts. We discuss extensions to the framework and propose research directions toward stronger guarantees.

大模型安全交互控制机器人学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。