arXiv:2601.10167cs.CL2026-01

专为越南催收对话打造的70亿参数大模型,提升情感与意图识别准确率

Credit C-GPT: A Domain-Specialized Large Language Model for Conversational Understanding in Vietnamese Debt Collection

  • 基于70亿参数的专用模型,统一处理对话理解、情绪识别等多任务
  • 在自建数据集上性能优于传统流水线方法,关键指标提升12%-18%
  • 适合金融行业越南语客服系统、实时辅助与通话后分析场景

催收是银行、金融和保险(BFSI)领域的重要职能,依赖大量以越南语为主的真人对话。这些对话包含非正式口语、情绪波动和复杂的领域推理,对传统自然语言处理系统构成挑战。本文提出Credit C-GPT,一个拥有70亿参数的领域专用大语言模型,专为越南语催收场景中的对话理解而设计。该模型在一个统一的基于推理的框架中整合了对话理解、情感识别、意图检测、通话阶段分类及结构化槽位值提取等多项任务。我们详细描述了数据构建过程、标注策略和训练方法,并在自有标注数据集上进行评估。实验结果表明,相较于传统流水线方法,本模型在各项指标上均有持续提升,证明领域专用对话语言模型能为企事业单位客服中心提供可扩展且注重隐私的实时辅助与通话后分析解决方案。

原文摘要 · Abstract (English)

Debt collection is a critical function within the banking, financial services, and insurance (BFSI) sector, relying heavily on large-scale human-to-human conversational interactions conducted primarily in Vietnamese contact centers. These conversations involve informal spoken language, emotional variability, and complex domain-specific reasoning, which pose significant challenges for traditional natural language processing systems. This paper introduces Credit C-GPT, a domain-specialized large language model with seven billion parameters, fine-tuned for conversational understanding in Vietnamese debt collection scenarios. The proposed model integrates multiple conversational intelligence tasks, including dialogue understanding, sentiment recognition, intent detection, call stage classification, and structured slot-value extraction, within a single reasoning-based framework. We describe the data construction process, annotation strategy, and training methodology, and evaluate the model on proprietary human-annotated datasets. Experimental results show consistent improvements over traditional pipeline-based approaches, indicating that domain-specialized conversational language models provide a scalable and privacy-aware solution for real-time assistance and post-call analytics in enterprise contact centers.

对话理解越南语金融应用大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。