arXiv:2608.13622cs.AIcs.CL2026-08

解决开放世界交互中奖励不公平问题,让模型行为比较更合理。

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

论文配图:ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
图 1 · 摘自论文原文
  • 用条件分组法重构交互轨迹比较,确保公平性
  • 在τ/τ²工具使用基准上显著提升性能
  • 适合研究交互式AI与强化学习公平性的学者

开放世界交互存在多种有效行为模式:代理可直接回答、请求澄清、提供进展或行动前确认。这种灵活性打破了基于群体的强化学习的核心假设——同一组内的轨迹不再具有行为可比性。结果,奖励模型对交互风格的偏好会扭曲相对优势,引导优化偏向奖励偏好行为而非情境适配行为。本文将此问题形式化为“奖励公平性问题”,提出ARC(通过条件控制实现优势正则化)训练方法,通过策略条件分组、混合奖励和熵正则化恢复更公平的相对比较。我们在新提出的inter框架下评估,该框架支持响应式、可引导、执行感知的用户-代理交互,解耦可见沟通与隐式推理及工具使用。inter还提供了用于构建inter-86K策略标注训练语料库的标注与蒸馏流程。实证显示,ARC显著提升了核心τ/τ²工具使用基准表现;相比思考型基线,inter将首次响应时间从4.91秒降低至1.27秒。这些结果表明,开放世界交互学习的关键瓶颈不仅是奖励设计,更是行为比较是否公平。ARC代码与inter-86K数据集将公开。

原文摘要 · Abstract (English)

Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a \textit{reward fairness problem} and propose \textbf{ARC} (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core $τ/τ^2$ tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.

强化学习交互系统公平性工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。