用自适应采样提升大模型异常事件检测能力,显著增强跨场景泛化。
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
- 引入双循环动态课程学习,按困惑度自适应选择难样本
- 在真实外卖对话中实现17.19%的F1提升和9.59%的跨域性能增益
- 适合需要快速部署、跨业务场景通用的工业级异常检测应用
真实客户客服对话中的异常事件检测因业务数据复杂性和交互动态性而极具挑战。模型需具备强跨领域(OOD)泛化能力,以快速适应不同业务场景并最大化商业价值。本文提出一种新型自适应困惑度感知强化学习框架(APARL),利用大语言模型的推理能力进行异常事件检测。APARL采用双循环动态课程学习架构,使模型在能力提升过程中逐步聚焦更难样本,有效缓解性能瓶颈,显著增强跨域迁移能力。在外卖配送对话任务上的大量评估表明,该模型在适应性和鲁棒性上均有显著提升,平均F1得分提高17.19%,跨域测试平均提升9.59%。该方法为异常检测模型的工业部署提供了更优解决方案,有助于提升运营效率与商业效益。
原文摘要 · Abstract (English)
Detecting abnormal events in real-world customer service dialogues is highly challenging due to the complexity of business data and the dynamic nature of customer interactions. Moreover, models must demonstrate strong out-of-domain (OOD) generalization to enable rapid adaptation across different business scenarios and maximize commercial value. In this work, we propose a novel Adaptive Perplexity-Aware Reinforcement Learning (APARL) framework that leverages the advanced reasoning capabilities of large language models for abnormal event detection. APARL introduces a dual-loop dynamic curriculum learning architecture, enabling the model to progressively focus on more challenging samples as its proficiency increases. This design effectively addresses performance bottlenecks and significantly enhances OOD transferability. Extensive evaluations on food delivery dialogue tasks show that our model achieves significantly enhanced adaptability and robustness, attaining the highest F1 score with an average improvement of 17.19\%, and an average improvement of 9.59\% in OOD transfer tests. This method provides a superior solution for industrial deployment of anomaly detection models, contributing to improved operational efficiency and commercial benefits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。