arXiv:2607.27421cs.CLcs.AI2026-07

41个开源大模型零样本评估,帮开发者选适合对话系统的模型。

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

论文配图:Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
图 1 · 摘自论文原文
  • 系统性测试41个开源模型在8个数据集上的表现
  • 3B指令微调模型性能超过部分7B基础模型
  • 发现主流基准已饱和,需新评测标准

意图分类是任务导向对话系统的核心组件,但实践者在计算资源、延迟和鲁棒性约束下缺乏系统的开源大模型选择指导。本文对41个开源语言模型(涵盖15个模型家族,参数量135M至9B)进行了零样本系统评估,覆盖8个英文单标签意图分类数据集,其中ATIS数据集使用5个标注示例作为辅助五样本结果。评估包含标准基准、大规模语音助手语料库及生产环境电商数据集。除精确匹配准确率外,还分析了置信度校准、对真实输入扰动的鲁棒性、模型排名统计可靠性、部署效率及基准饱和度。结果显示:指令微调的3B模型可超越多个7B基础模型;在MASSIVE数据集上,领先模型间的差异在配对麦克内马尔检验中无统计显著性;如SNIPS等常用基准已趋于饱和,无法有效区分当前开源模型。指令微调对置信度校准的影响不一致,并非普遍有害。研究为意图分类场景下的开源模型选择与评估提供了实用指导。

原文摘要 · Abstract (English)

Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under compute, latency, and robustness constraints. We present a systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M--9B parameter range across eight English single-label intent-classification datasets. A ninth dataset, ATIS, uses five labeled demonstrations and is reported as an auxiliary five-shot result. The evaluation includes standard benchmarks, a large-scale voice-assistant corpus, and production-derived e-commerce datasets. Beyond exact-match accuracy, we analyze confidence calibration, robustness to realistic input perturbations, statistical reliability of model rankings, deployment efficiency, and benchmark saturation. Our results show that instruction-tuned 3B models can outperform several evaluated 7B base models, that differences among leading models on MASSIVE are statistically indistinguishable under pairwise McNemar tests, and that widely used benchmarks such as SNIPS have become saturated and no longer meaningfully discriminate among current open-weight models. Instruction tuning's effect on confidence calibration is inconsistent rather than uniformly harmful. These findings provide practical guidance for selecting and evaluating open-weight language models for intent classification.

意图分类大模型评估零样本开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。