小模型通过摘要对话可实现接近大模型的客服问答效果
Can Small Language Models Handle Context-Summarized Multi-Turn Customer-Service QA? A Synthetic Data-Driven Comparative Evaluation
- 用对话摘要保留上下文,让小模型理解多轮客服对话
- 部分小模型表现接近商用大模型,但整体差异明显
- 适合资源有限场景下的客服系统研发与评估
客服问答系统越来越依赖对话语言理解。尽管大语言模型(LLMs)表现优异,但其高计算成本和部署限制使其在资源受限环境中难以应用。小语言模型(SLMs)提供更高效的替代方案,但在需要保持对话连续性和上下文理解的多轮客服问答任务中,其有效性仍待深入探索。本研究聚焦指令微调的小模型在经过对话摘要处理的多轮客服问答任务中的表现,采用历史摘要策略以保留关键对话状态。同时引入基于对话阶段的定性分析方法,评估模型在不同交互阶段的行为表现。我们对九个低参数量的指令微调小模型进行了评估,并与三个商用大模型进行对比,使用词汇和语义相似性指标,结合人工评估及大模型作为裁判的评价方式。结果显示,不同小模型表现差异显著,部分模型达到接近大模型的性能,而另一些则难以维持对话连贯性和上下文一致性。这些发现揭示了低参数模型在真实客服问答系统中的潜力与当前局限。
原文摘要 · Abstract (English)
Customer-service question answering (QA) systems increasingly rely on conversational language understanding. While Large Language Models (LLMs) achieve strong performance, their high computational cost and deployment constraints limit practical use in resource-constrained environments. Small Language Models (SLMs) provide a more efficient alternative, yet their effectiveness for multi-turn customer-service QA remains underexplored, particularly in scenarios requiring dialogue continuity and contextual understanding. This study investigates instruction-tuned SLMs for context-summarized multi-turn customer-service QA, using a history summarization strategy to preserve essential conversational state. We also introduce a conversation stage-based qualitative analysis to evaluate model behavior across different phases of customer-service interactions. Nine instruction-tuned low-parameterized SLMs are evaluated against three commercial LLMs using lexical and semantic similarity metrics alongside qualitative assessments, including human evaluation and LLM-as-a-judge methods. Results show notable variation across SLMs, with some models demonstrating near-LLM performance, while others struggle to maintain dialogue continuity and contextual alignment. These findings highlight both the potential and current limitations of low-parameterized language models for real-world customer-service QA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。