arXiv:2604.15789cs.CL2026-04

系统评估无需训练的可信大模型方法,揭示其优劣与权衡。

A Systematic Study of Training-Free Methods for Trustworthy Large Language Models

论文配图:A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
图 1 · 摘自论文原文
  • 按输入、内部、输出三阶段分类,梳理干预位置
  • 发现现有方法在可信性与性能间存在显著权衡
  • 为无训练部署提供实用建议,适合安全研究者

随着大语言模型(LLMs)在各领域广泛应用,其生成有害内容、偏见、无依据声明及对抗攻击脆弱性等问题日益受关注。为实现快速低成本适配,无需训练的方法作为后训练对齐的替代方案兴起。然而,现有方法评估不一致,覆盖维度有限,且可能引发性能下降和鲁棒性减弱等副作用。本文系统重评现有无需训练方法在多种可信场景下的有效性,及其对模型效用、鲁棒性与计算开销的影响。基于推理过程中干预位置,将方法分为输入、内部、输出三类,对不同规模与架构的LLM家族中代表性方法进行综合分析。研究揭示当前方法存在的多方面权衡与未解挑战,总结文献核心发现与局限,并提出无需额外训练即可平衡可信性、效用与鲁棒性的实践建议。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) receive increasing attention and are being deployed across various domains, their potential risks, including generating harmful or biased content, producing unsupported claims, and exhibiting vulnerabilities to adversarial attacks, have drawn significant attention. To enable quick and low-cost adaptation, training-free methods have recently emerged as cost-effective alternatives to post-training alignment techniques. Despite their promising results, these methods are evaluated inconsistently across the literature, cover limited dimensions of trustworthiness, and can introduce undesirable side effects, such as utility degradation and increased brittleness. To fully assess the impacts of these training-free methods, we take a step back and systematically re-evaluate the effectiveness of existing training-free methods against various trustworthy settings and their influence on utility, robustness, and computational overhead. We also categorize these methods into three levels (input, internal, and output) based on where they intervene in the model's information flow during inference. Using this taxonomy, we conduct a comprehensive analysis of various representative and effective methods from each level across different LLM families and sizes. Our analysis highlights several trade-offs and unresolved challenges in current approaches. We summarize key findings and limitations in the existing literature, and propose practical recommendations for balancing trustworthiness, utility, and robustness in LLMs without the need for additional training.

大模型安全无需训练可信性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。