识别大模型预测中‘看似正确却脆弱’的隐性风险,提升鲁棒性。
FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness

- 提出边际感知误差流机制,捕捉正确但易错的预测
- 在多个基准上实现扰动后最差类准确率提升
- 适用于对模型可靠性要求高的场景
大模型鲁棒性评估常依赖平均准确率或一致性,但这些指标可能掩盖一种结构化失败:预测虽正确,但概率质量已从真实类别流向决策边界附近的系统性错误类别。本文将此现象形式化为边际感知误差流,并提出FragileFlow——一种可插拔正则化器,利用校准的边际缓冲区识别正确但脆弱的预测,并将其非本类概率质量组织为类别级脆弱风险矩阵。理论上,首次给出了该误差流对象的PAC-Bayes上界,证明了经验谱控制在稳定性条件下可保守地保证确定性最差类鲁棒性。在多项选择型LLM基准和少样本CLIP适配实验中,FragileFlow始终优于匹配基线,在多数设置下提升了扰动后的最差类准确率,同时保持干净数据上的准确率不变。
原文摘要 · Abstract (English)
Robust adaptation of LLMs and VLMs is often evaluated by average accuracy or average consistency under perturbations. However, these averages can hide a structured failure mode: a prediction may remain correct while probability mass already flows from particular true classes toward systematic wrong competitors near the decision boundary. In this paper, we formalize this phenomenon as margin-aware error flow and introduce FragileFlow, a plug-in regularizer that uses a calibrated margin buffer to identify correct-but-fragile predictions and organize their off-class probability mass into a class-wise vulnerable-risk matrix. Theoretically, we provide the first PAC-Bayes upper bound for this margin-aware error-flow object, showing how empirical spectral control yields a conservative route to deterministic worst-class robustness under a stability condition. Experiments on multiple-choice LLM benchmarks and few-shot CLIP adaptation show that FragileFlow consistently improves the proposed theory-facing risk measures over matched baselines, yields perturbed worst-class accuracy gains in most settings, and preserves clean accuracy across comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。