arXiv:2604.15589cs.CLcs.AI2026-04中稿 · ICCCBE 2026被引 1

对比不同微调方法与模型规模对代码合规生成解释性的影响

LLM attribution analysis across different fine-tuning strategies and model scales for automated code compliance

  • 用扰动法分析不同微调方式的解释行为差异
  • 7B以上模型性能趋稳,但解释策略更聚焦数值约束
  • 全量微调比高效微调更具可解释性,适合监管场景

现有针对大语言模型在自动化代码合规任务中的研究多关注性能表现,将模型视为黑箱,忽视训练策略对其解释行为的影响。本文采用基于扰动的归因分析,比较全量微调(FFT)、低秩适应(LoRA)及量化LoRA微调在不同模型规模下的解释行为。结果表明,FFT产生的归因模式在统计上更集中且不同于参数高效微调方法;随着模型规模扩大,大模型发展出优先关注数值约束和规则标识符的解释策略,但超过7B参数后,生成规则与参考规则在语义相似度上的性能提升趋于饱和。本研究为模型可解释性提供关键洞见,推动建筑、工程、施工领域中高监管要求任务的透明化大模型建设。

原文摘要 · Abstract (English)

Existing research on large language models (LLMs) for automated code compliance has primarily focused on performance, treating the models as black boxes and overlooking how training decisions affect their interpretive behavior. This paper addresses this gap by employing a perturbation-based attribution analysis to compare the interpretive behaviors of LLMs across different fine-tuning strategies such as full fine-tuning (FFT), low-rank adaptation (LoRA) and quantized LoRA fine-tuning, as well as the impact of model scales which include varying LLM parameter sizes. Our results show that FFT produces attribution patterns that are statistically different and more focused than those from parameter-efficient fine-tuning methods. Furthermore, we found that as model scale increases, LLMs develop specific interpretive strategies such as prioritizing numerical constraints and rule identifiers in the building text, albeit with performance gains in semantic similarity of the generated and reference computer-processable rules plateauing for models larger than 7B. This paper provides crucial insights into the explainability of these models, taking a step toward building more transparent LLMs for critical, regulation-based tasks in the Architecture, Engineering, and Construction industry.

大模型解释性代码合规微调对比模型规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。