构建首个越南语客服问答基准数据集,助力高效模型选型
A Benchmark Dataset and Evaluation Framework for Vietnamese Large Language Models in Customer Support
- 基于9000+真实客服对话构建专用数据集
- 11个轻量级越南语模型在该数据集上表现差异显著
- 适合研究者和企业评估越南语大模型实际应用能力
随着人工智能快速发展,大语言模型(LLMs)已成为提升客户服务效率的关键技术。越南语大模型(ViLLMs)凭借精度高、效率优和隐私保护等优势,成为轻量开源模型的实用选择。然而,领域特定评估仍不充分,缺乏反映真实客户交互的基准数据集,导致企业难以筛选适用模型。为此,本文构建了客户支持对话数据集(CSConDa),包含超过9,000对来自大型越南软件公司真实人机客服交互的问答对,涵盖定价、产品可得性、技术故障处理等多样化主题。我们进一步提出一套综合评估框架,对11个轻量级开源越南语模型在该数据集上进行自动指标与句法分析,揭示模型性能差异、优势弱点及语言模式。研究为模型行为提供洞察,识别改进方向,推动下一代越南语大模型发展。通过建立可靠基准与系统评估体系,本工作支持客户支持问答系统的模型选型,促进越南语大模型研究进展。数据集已公开:https://huggingface.co/datasets/ura-hcmut/Vietnamese-Customer-Support-QA。
原文摘要 · Abstract (English)
With the rapid growth of Artificial Intelligence, Large Language Models (LLMs) have become essential for Question Answering (QA) systems, improving efficiency and reducing human workload in customer service. The emergence of Vietnamese LLMs (ViLLMs) highlights lightweight open-source models as a practical choice for their accuracy, efficiency, and privacy benefits. However, domain-specific evaluations remain limited, and the absence of benchmark datasets reflecting real customer interactions makes it difficult for enterprises to select suitable models for support applications. To address this gap, we introduce the Customer Support Conversations Dataset (CSConDa), a curated benchmark of over 9,000 QA pairs drawn from real interactions with human advisors at a large Vietnamese software company. Covering diverse topics such as pricing, product availability, and technical troubleshooting, CSConDa provides a representative basis for evaluating ViLLMs in practical scenarios. We further present a comprehensive evaluation framework, benchmarking 11 lightweight open-source ViLLMs on CSConDa with both automatic metrics and syntactic analysis to reveal model strengths, weaknesses, and linguistic patterns. This study offers insights into model behavior, explains performance differences, and identifies key areas for improvement, supporting the development of next-generation ViLLMs. By establishing a robust benchmark and systematic evaluation, our work enables informed model selection for customer service QA and advances research on Vietnamese LLMs. The dataset is publicly available at https://huggingface.co/datasets/ura-hcmut/Vietnamese-Customer-Support-QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。