用自动反馈替代人工标注,让对话机器人在对抗中自我进化。
RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

- 用环境反馈代替人工标注,实现无监督对抗训练。
- 银行客服测试中工具路由错误归零,通过率翻倍。
- 发现新型伪装策略,适合安全评估与模型加固研究。
企业在部署任务导向型对话系统时面临持续的标注瓶颈:鲁棒训练需大规模标注交互数据,但企业对话日志涉及隐私且标注成本高,用户行为演化速度远超标注流程。我们提出RL-ADA(基于对抗对话代理的强化学习),通过引入 extit{世界反馈}——直接从可测量交互结果中提取的后果性奖励信号——消除该瓶颈。一个30亿参数的客服代理(DA)与一个70亿参数的对抗客户代理(CA)在对抗环境中协同进化,由固定自动化评判器驱动:DA因成功解决多轮对话获奖励,CA则因生成能导致误转的真实意图隐藏语句而获奖励,形成不对称对抗压力。隔离训练机制迭代重训表现较弱的一方,全程无需人工标注。在银行业客服原型验证中,工具路由错误完全消除,端到端通过率在五轮协同进化后翻倍,仅依赖自动化奖励信号,无需标注数据。此外观察到 extbf{上下文伪装}现象:CA仅凭奖励压力学会将意图嵌入密集真实客户细节中,对企事业单位红队测试和鲁棒性评估具有直接意义。
原文摘要 · Abstract (English)
Deploying task-oriented dialogue agents in enterprise customer support faces a persistent annotation bottleneck: robust training requires labelled interaction data at scale, yet enterprise conversational logs are privacy-sensitive and expensive to annotate, while user behaviour evolves faster than labelling pipelines can keep pace. We present RL-ADA (Reinforcement Learning with Adversarial Dialogue Agents), a co-evolutionary training framework that eliminates this bottleneck by replacing human labels with \emph{world feedback}: consequence-based reward signals derived directly from measurable interaction outcomes. A Customer Support Agent (DA, 3B parameters) and an Adversarial Customer Agent (CA, 7B parameters) co-evolve in an adversarial arena guided by a fixed automated judge: the DA is rewarded for correctly handling multi-turn customer conversations to successful resolution, while the CA is rewarded for producing realistic, intent-concealing utterances that cause misroutes, creating asymmetric adversarial pressure through opposing but independently structured rewards. An isolation gym iteratively retrains the weaker agent on prior-failure transcripts, requiring no human annotation at any stage. In a banking customer support proof of concept, tool-routing errors are eliminated and the strict end-to-end PASS rate doubles over five co-evolutionary cycles, driven solely by automated arena reward with no labelled data. We additionally observe the emergence of \textbf{Contextual Camouflage}, an adversarial strategy in which the CA learns to embed intent within dense realistic customer detail purely from reward pressure, with direct implications for enterprise red-teaming and robustness evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。