arXiv:2605.27835cs.LGcs.CL2026-05

无需标注理由,让大模型解释更可信且高效。

CAREF: Calibration-Aware Regularization for Explanation Faithfulness Without Rationale Supervision

论文配图:CAREF: Calibration-Aware Regularization for Explanation Faithfulness Without Rationale Supervision
图 1 · 摘自论文原文
  • 用统一损失函数同时优化预测准确率和解释一致性。
  • 仅用6.43%参数就达到89.04准确率和81.00 nBERT对齐度。
  • 适合追求轻量化可解释性微调的研究者或应用开发者。

我们提出CAREF,一种参数高效的微调框架,通过校准感知正则化联合优化预测准确率与解释忠实性。其核心是将基于熵的校准与令牌级稀疏性控制结合于单一损失函数——解释忠实性的校准感知正则化(LSCED),无需理由标注。在四个自然语言理解基准(COS-E、ECQA、ComVE、e-SNLI)上使用Flan-T5评估,轻量版CAREF-AQ仅用6.43%可训练参数,即取得平均准确率89.04和解释对齐度81.00 nBERT,优于LoRA与AdaLoRA。据我们所知,CAREF是首个将熵与稀疏性正则化统一于单一训练目标中用于可解释大模型微调的方法。

原文摘要 · Abstract (English)

We introduce CAREF, a parameter-efficient fine-tuning framework that jointly optimizes predictive accuracy and explanation faithfulness via calibration-aware regularization. At its core, CAREF couples entropy-based calibration with token-level sparsity control through a single unified loss, the Calibration-Aware Regularization for Explanation Faithfulness (LSCED), without requiring rationale supervision. Evaluated on four NLE benchmarks (COS-E, ECQA, ComVE, e-SNLI) with Flan-T5, our lightweight CAREF-AQ variant attains the best average accuracy (89.04) and explanation alignment (81.00 nBERT) using only 6.43% of trainable parameters, outperforming LoRA and AdaLoRA. To our knowledge, CAREF is the first method to unify entropy and sparsity regularization in a single training objective for interpretable LLM fine-tuning.

可解释性参数高效大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。