无需外部标注,通过干预内部神经元实现自引导去偏。
NeuronTune: Towards Self-Guided Spurious Bias Mitigation
- 在隐空间探测并调节导致误判的神经元
- 在多个模型和数据上显著降低虚假关联依赖
- 适合提升模型在分布外数据上的鲁棒性
深度神经网络常产生虚假偏差,即依赖非关键特征与类别的共现关系进行预测。例如,模型可能基于频繁出现的背景而非物体本身特征来识别对象,导致在缺乏这些共现关系的数据上性能下降。现有方法通常依赖外部提供的虚假关联标注,获取困难且不适用于模型自身产生的偏差。本文提出 NeuronTune,一种后处理的自引导去偏方法,直接干预模型内部决策过程。该方法在模型的隐空间中探测并调控引发虚假预测行为的神经元。我们从理论上证明该方法可使模型更接近无偏状态。与以往方法不同,NeuronTune 不需要虚假相关性标注,因此在实际应用中更具可行性与有效性。跨多种架构和数据模态的实验表明,该方法能以自引导方式显著缓解虚假偏差。
原文摘要 · Abstract (English)
Deep neural networks often develop spurious bias, reliance on correlations between non-essential features and classes for predictions. For example, a model may identify objects based on frequently co-occurring backgrounds rather than intrinsic features, resulting in degraded performance on data lacking these correlations. Existing mitigation approaches typically depend on external annotations of spurious correlations, which may be difficult to obtain and are not relevant to the spurious bias in a model. In this paper, we take a step towards self-guided mitigation of spurious bias by proposing NeuronTune, a post hoc method that directly intervenes in a model's internal decision process. Our method probes in a model's latent embedding space to identify and regulate neurons that lead to spurious prediction behaviors. We theoretically justify our approach and show that it brings the model closer to an unbiased one. Unlike previous methods, NeuronTune operates without requiring spurious correlation annotations, making it a practical and effective tool for improving model robustness. Experiments across different architectures and data modalities demonstrate that our method significantly mitigates spurious bias in a self-guided way.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。