用实时行为评估约束大模型,让其在敏感领域更靠谱
A Grounded Observer Framework for Establishing Guardrails for Foundation Models in Socially Sensitive Domains
- 通过低维行为特征实时监控,动态调整模型输出
- 实现对话上下文一致的自然闲聊,支持无剧本人机互动
- 适合医疗、金融等高风险场景,可推广至多社交情境
随着大模型在医疗、金融、心理健康等敏感领域的广泛应用,确保其行为符合预期目标与社会规范变得至关重要。由于这些高维模型的复杂性,传统依赖低维离散状态与动作空间的行为约束方法难以直接适用。受机器人动作选择技术启发,本文提出一种基于实证观察的框架,可在保证行为可靠性的前提下实现实时动态调整。该方法通过实时评估低阶行为特征,动态调节模型行为并提供上下文反馈。为验证有效性,我们构建了一个能维持上下文恰当闲聊能力的系统,并将其应用于机器人,实现与人类的全新非脚本化交互。最后,讨论了该框架在其他社交场景中的潜在应用及未来研究方向。
原文摘要 · Abstract (English)
As foundation models increasingly permeate sensitive domains such as healthcare, finance, and mental health, ensuring their behavior meets desired outcomes and social expectations becomes critical. Given the complexities of these high-dimensional models, traditional techniques for constraining agent behavior, which typically rely on low-dimensional, discrete state and action spaces, cannot be directly applied. Drawing inspiration from robotic action selection techniques, we propose the grounded observer framework for constraining foundation model behavior that offers both behavioral guarantees and real-time variability. This method leverages real-time assessment of low-level behavioral characteristics to dynamically adjust model actions and provide contextual feedback. To demonstrate this, we develop a system capable of sustaining contextually appropriate, casual conversations ("small talk"), which we then apply to a robot for novel, unscripted interactions with humans. Finally, we discuss potential applications of the framework for other social contexts and areas for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。