arXiv:2504.02577cs.AIcs.CL2025-04

揭示深度模型推理不一致问题并提出检测与缓解方法

Reasoning Inconsistencies and How to Mitigate Them in Deep Learning

  • 提出新方法检测语言与图像模型内部推理不一致
  • 在低资源场景下通过合成数据提升公平性与性能
  • 适合关注模型可解释性与鲁棒性的研究者

深度学习模型在多任务、多模态上取得显著进展,但其内部推理过程仍缺乏理解,尤其在系统性不一致或逻辑错误方面。这些不一致表现为矛盾输出、泛化失败或特定情境下错误结论,其根源可能来自模型内部机制不透明、训练数据偏见或任务复杂性。本文提出两种检测与量化自然语言和图像模型预测不一致的方法;针对数据偏见,提出高效采样方法与低资源场景下的合成数据生成策略;此外还设计两种优化复杂推理任务的模型方法。这些技术提升了模型性能,增强了推理过程的可解释性与可信度。本文构建了一个涵盖鲁棒性、公平性与可解释性的综合框架,适用于多种任务与模态。

原文摘要 · Abstract (English)

The recent advancements in Deep Learning models and techniques have led to significant strides in performance across diverse tasks and modalities. However, while the overall capabilities of models show promising growth, our understanding of their internal reasoning processes remains limited, particularly concerning systematic inconsistencies or errors patterns of logical or inferential flaws. These inconsistencies may manifest as contradictory outputs, failure to generalize across similar tasks, or erroneous conclusions in specific contexts. Even detecting and measuring such reasoning discrepancies is challenging, as they may arise from opaque internal procedures, biases and imbalances in training data, or the inherent complexity of the task. Without effective methods to detect, measure, and mitigate these errors, there is a risk of deploying models that are biased, exploitable, or logically unreliable. This thesis aims to address these issues by producing novel methods for deep learning models that reason over knowledge graphs, natural language, and images. The thesis contributes two techniques for detecting and quantifying predictive inconsistencies originating from opaque internal procedures in natural language and image processing models. To mitigate inconsistencies from biases in training data, this thesis presents a data efficient sampling method to improve fairness and performance and a synthetic dataset generation approach in low resource scenarios. Finally, the thesis offers two techniques to optimize the models for complex reasoning tasks. These methods enhance model performance while allowing for more faithful and interpretable exploration and exploitation during inference. Critically, this thesis provides a comprehensive framework to improve the robustness, fairness, and interpretability of deep learning models across diverse tasks and modalities.

推理一致性模型可解释性公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。