用大模型打造高效语音系统,自动采集海量地点信息
DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition

- 用有限状态机生成多样化对话数据,解决真实场景中罕见问题
- 通过思维链机制与选择性生成,提升输出稳定性和准确性
- 支持持续优化且人工干预少,适合大规模工业级对话应用
精准获取兴趣点(POI)属性对位置服务至关重要,但传统模块化语音应答系统存在错误累积和维护成本高的问题。我们提出DuIVRS-2,一个基于大语言模型(LLM)的端到端框架,用于百度地图的大规模POI属性采集。为应对真实交互中的长尾分布,方法首先采用有限状态机(FSM)引导的数据增强策略,合成平衡多样化的训练数据集;随后通过选择性生成结合思维链(CoT)机制,简化对话管理,确保输出稳定性并有效消除幻觉。为实现低人工成本的持续策略优化,设计了双评估器投票的协作式迭代学习框架。系统上线两个月,日处理0.4万通电话,任务成功率(TSR)达83.9%,较前代提升4个百分点,响应时间仅130ms。本工作为构建高效、低成本的工业级大模型对话代理提供了可落地的参考。
原文摘要 · Abstract (English)
Accurate Point of Interest (POI) attribute acquisition is essential for location-based services, yet traditional modular Interactive Voice Response (IVR) systems suffer from error accumulation and high maintenance overhead. We present DuIVRS-2, a large language model (LLM)-based end-to-end framework designed for large-scale POI attribute acquisition at Baidu Maps. To address the long-tail distribution of real-world interactions, our methodology first employs a finite state machine (FSM)-guided data augmentation strategy to synthesize a balanced and diverse training dataset. We then streamline dialogue management via a selective generation scheme combined with a Chain-of-Thought (CoT) mechanism, which ensures output stability and effectively eliminates hallucinations in industrial settings. To facilitate continuous policy refinement with minimal manual effort, we design a cooperative iterative learning framework that leverages a dual-evaluator voting system. Deployed in production for two months, DuIVRS-2 processed 0.4 million calls daily and achieved a 83.9\% Task Success Rate (TSR), outperforming its predecessor by 4 percentage points while maintaining a low reaction time of 130ms. This work provides a production-proven reference for developing robust, cost-effective LLM agents for large-scale industrial dialogue applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。