arXiv:2606.05176cs.CLcs.AI2026-06

用LoRA微调小模型,提升电信客服对话质量并分析能耗与性能关系。

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

论文配图:PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis
图 1 · 摘自论文原文
  • 通过组合式合成数据生成3万条电信场景对话样本。
  • 发现最低损失模型未必最符合人类评价,需结合定性评估。
  • 提供低耗能部署方案,适合对合规与效率有要求的行业应用。

尽管大语言模型在自然语言理解与生成方面表现强劲,但其在电信客户支持领域特定约束下的评估与适配仍有限。数据主权、监管要求及敏感客户与网络信息处理使外部托管基础模型的应用复杂化。本文系统研究将参数高效微调(PEFT)中的低秩适应(LoRA)应用于Qwen2.5-3B,构建领域专用对话助手。提出基于52个行业术语词典的组合式合成数据生成方法,借助Gemini 2.0 Flash生成约3万条训练样本,覆盖1,560种不同问题场景。评估16种LoRA配置,调整超参数与目标模块。评估不仅包含标准指标,还引入能耗分析及基于GPT-5.2和Claude 4.5 Sonnet的LLM-as-a-judge框架进行定性评估。结果表明量化与质化表现存在显著背离:最低验证损失(0.5024)仅在定性评估中排名第6-7位,而最高损失(0.6807)反而获得双模型最优排名。本工作贡献包括:(1) 一种组合式合成数据构建方法;(2) 对LoRA注入目标模块选择的影响洞察;(3) 证明验证损失不足以作为对话AI微调配置选择依据;(4) 提供可持续部署的能效-性能权衡分析。

原文摘要 · Abstract (English)

While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-specific constraints in telecommunications customer support remain limited. In addition, data sovereignty, regulatory constraints, and the handling of sensitive customer and network information complicate the use of externally hosted foundation models in this domain. We present a systematic study of parameter-efficient fine-tuning (PEFT) using Low-Rank Adaptation (LoRA) applied to Qwen2.5-3B to build a domain-specific conversational assistant. We introduce a combinatorial synthetic data generation approach based on a glossary of 52 industry-specific terms, producing approximately 30,000 training examples across 1,560 distinct problem scenarios via a generative pipeline powered by Gemini 2.0 Flash. We evaluate 16 LoRA configurations by varying hyperparameters and target modules. Our evaluation extends beyond standard metrics by incorporating energy consumption analysis and qualitative assessment using an LLM-as-a-judge framework with GPT-5.2 and Claude 4.5 Sonnet. Results show a clear divergence between quantitative and qualitative performance: models achieving the lowest validation loss do not necessarily obtain the best human-aligned rankings. The best validation loss (0.5024) ranks only 6th-7th in qualitative evaluation, while the worst loss (0.6807) ranks first according to both judges. This work contributes (1) a combinatorial method for synthetic dataset construction, (2) insights into the impact of target module selection for LoRA injection, (3) evidence that validation loss alone is insufficient for selecting fine-tuning configurations in conversational AI, and (4) an energy-performance trade-off analysis for sustainable LLM deployment.

LoRA对话系统能效分析电信客服

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。