arXiv:2608.00042cs.CLcs.AI2026-08

小模型领域适配可能暗藏信任风险,安全策略未必有效。

Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study

  • 对比四种微调方法在医疗、法律、金融领域的表现
  • 对抗数据提升适配效果,但三种安全策略反增攻击风险
  • 发现现有安全方法无法保护小模型免受恶意输入干扰

小型语言模型(SLMs)的领域适配已成为资源受限、高风险场景(如医疗、法律、金融)中部署高效NLP系统的重要策略。尽管参数高效微调可提升性能,但其对可信度(事实校准与对抗鲁棒性)的影响尚不明确。本文首次系统性地开展跨架构、跨领域的实证研究,评估三种SLM架构(TinyLlama 1B、Gemma-2 2B、Llama 3.2 1B)、三个领域(医疗、法律、金融)、两种训练数据条件(良性与对抗扰动)及四种微调策略(基线QLoRA、Safety-DPO、Dark Experience Replay、Task Arithmetic LoRA, TA-LoRA)下的可信度代价。通过TruthfulQA MC2(事实校准)与HarmBench ASR(对抗鲁棒性)在216组实验配置中评估,每组使用三个随机种子。主要发现:第一,基线QLoRA对事实校准影响极小(|ΔTQA|均值<0.02);第二,对抗扰动训练数据显著提升适配质量(Δ损失≈-0.040),且未损害可信度基准;第三,三种安全策略均未降低对抗危害敏感性:Safety-DPO基本无影响(ΔASR均值<0.001),而Dark ER与TA-LoRA使安全对齐模型(Gemma-2 2B、Llama 3.2 1B)的平均HarmBench ASR分别上升+0.171和+0.155,个别配置超+0.45。结果挑战了重放与算术融合策略能迁移对齐能力的假设。

原文摘要 · Abstract (English)

Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-stakes environments including healthcare, legal services, and financial analysis. While performance gains from parameter-efficient fine-tuning are well characterised, the corresponding impact on trustworthiness (factual calibration and adversarial robustness) remains poorly understood. This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation across three SLM architectures (TinyLlama 1B, Gemma-2 2B, Llama 3.2 1B), three domains (healthcare, legal, finance), two training-data conditions (benign and adversarially perturbed), and four fine-tuning strategies (baseline LoRA, Safety-DPO, Dark Experience Replay, and Task Arithmetic LoRA, TA-LoRA). Trustworthiness is evaluated through TruthfulQA MC2 (factual calibration) and HarmBench ASR (adversarial robustness) across all 216 experimental configurations with three random seeds. Three principal findings emerge. First, baseline QLoRA domain adaptation produces minimal TruthfulQA MC2 change across all model-domain combinations (mean |Delta TQA| < 0.02). Second, adversarially perturbed training data consistently improves domain adaptation quality (Delta loss approximately -0.040) without worsening trustworthiness benchmarks. Third, none of the three safety-preserving strategies reduced adversarial harm susceptibility: Safety-DPO was effectively neutral (mean Delta ASR < 0.001), while Dark ER and TA-LoRA increased mean HarmBench ASR by +0.171 and +0.155 respectively in safety-aligned models (Gemma-2 2B, Llama 3.2 1B), with individual configurations exceeding +0.45. These results challenge the assumption that replay-based and arithmetic-merge strategies transfer alignment to domain-adapted SLMs.

小模型领域适配可信度对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。