构建真实客服对话评估与优化闭环,提升模型落地表现
Benchmarking and Learning Real-World Customer Service Dialogue
- 设计OlaBench多场景评测基准,覆盖检索生成、流程系统与智能体设置
- OlaMind模型在评测中超越GPT-5.2与Gemini 3 Pro,线上解决率提升23.67%
- 通过专家对话提炼策略,结合分阶段强化学习实现可部署的客服能力提升
现有工业级智能客服(ICS)的评测与训练体系与真实对话需求脱节,过度关注可验证的任务成功率,而忽视主观服务质量与真实失败模式,导致离线指标提升难以转化为实际部署效果。为此,我们提出评测到优化的闭环:首先引入OlaBench,一个涵盖检索增强生成、流程化系统和智能体设置的ICS评测基准,评估服务能力、安全性与延迟敏感性;基于OlaBench结果发现顶尖大模型仍存在短板,我们提出OlaMind,从专家对话中提炼可复用的推理模式与服务策略,并采用实例级评分指导的分阶段探索-利用强化学习来提升模型能力。OlaMind在OlaBench上得分83.64,优于GPT-5.2(70.58)与Gemini 3 Pro(70.84),在线A/B测试中平均问题解决率提升23.67%,人工转接率下降6.6%,有效弥合了离线指标与实际部署之间的差距。OlaBench与OlaMind共同推动智能客服向更类人、专业、可靠的方向演进。项目主页与评测地址见https://olamind-olabench.github.io。
原文摘要 · Abstract (English)
Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements, overemphasizing verifiable task success while under-measuring subjective service quality and realistic failure modes, leaving a gap between offline gains and deployable dialogue behavior. We close this gap with a benchmark-to-optimization loop: we first introduce OlaBench, an ICS benchmark spanning retrieval-augmented generation, workflow-based systems, and agentic settings, which evaluates service capability, safety, and latency sensitivity; moreover, motivated by OlaBench results showing state-of-the-art LLMs still fall short, we propose OlaMind, which distills reusable reasoning patterns and service strategies from expert dialogues and applies staged exploration--exploitation reinforcement learning with instance-level rubric-aware guidance to improve model capability. OlaMind surpasses GPT-5.2 and Gemini 3 Pro on OlaBench (83.64 vs. 70.58/70.84) and, in online A/B tests, delivers an average +23.67% issue resolution and -6.6% human transfer rate versus the baseline, bridging offline gains to deployment. Together, OlaBench and OlaMind advance ICS systems toward more anthropomorphic, professional, and reliable deployment. The project page and evaluation are available at https://olamind-olabench.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。