arXiv:2510.13393cs.AI2025-10被引 2

提出新方法解决模型解释中生成内容单一问题,提升可解释性与准确性。

Learnable Game-theoretic Policy Optimization for Data-centric Self-explanation Rationalization

  • 从博弈论视角设计策略干预机制,动态优化生成与预测协同过程。
  • 在9个真实数据集上实现最高8.1%的性能提升,有效缓解模式坍缩。
  • 适用于需高可解释性的场景,如医疗、金融等关键决策领域。

理性化是一种以数据为中心的框架,旨在构建自解释模型,通过生成输入中人类可理解的部分(即推理依据)来解释预测结果。该过程涉及一个协作博弈模型:生成器负责生成最具可读性的输入片段(即推理依据),随后由预测器基于这些生成的推理依据做出预测。传统理性化方法通常通过正则化项施加约束以校准或惩罚不期望的生成行为,但这类方法存在模式坍缩问题——预测器虽能正确预测,生成器却持续输出重复的、缺乏多样性的推理依据。此外,现有研究多针对特定坍缩模式分别设计,缺乏统一处理思路。本文从新的博弈论视角系统重审协作理性化,识别出根本原因:生成器不再探索新策略以发现信息丰富的推理依据,最终导致系统收敛至次优博弈均衡(正确预测但推理依据单一)。为此,我们提出一种新方法——面向理性的博弈策略优化(PORAT),通过逐步引入策略干预,在协作博弈过程中调节博弈均衡,引导模型趋向更优解。我们理论上分析了次优均衡的成因,并证明所提方法的可行性。在九个广泛使用的现实世界数据集及两个合成设置上验证,PORAT相比现有最先进方法性能提升最高达8.1%。

原文摘要 · Abstract (English)

Rationalization, a data-centric framework, aims to build self-explanatory models to explain the prediction outcome by generating a subset of human-intelligible pieces of the input data. It involves a cooperative game model where a generator generates the most human-intelligible parts of the input (i.e., rationales), followed by a predictor that makes predictions based on these generated rationales. Conventional rationalization methods typically impose constraints via regularization terms to calibrate or penalize undesired generation. However, these methods are suffering from a problem called mode collapse, in which the predictor produces correct predictions yet the generator consistently outputs rationales with collapsed patterns. Moreover, existing studies are typically designed separately for specific collapsed patterns, lacking a unified consideration. In this paper, we systematically revisit cooperative rationalization from a novel game-theoretic perspective and identify the fundamental cause of this problem: the generator no longer tends to explore new strategies to uncover informative rationales, ultimately leading the system to converge to a suboptimal game equilibrium (correct predictions v.s collapsed rationales). To solve this problem, we then propose a novel approach, Game-theoretic Policy Optimization oriented RATionalization (PORAT), which progressively introduces policy interventions to address the game equilibrium in the cooperative game process, thereby guiding the model toward a more optimal solution state. We theoretically analyse the cause of such a suboptimal equilibrium and prove the feasibility of the proposed method. Furthermore, we validate our method on nine widely used real-world datasets and two synthetic settings, where PORAT achieves up to 8.1% performance improvements over existing state-of-the-art methods.

可解释性博弈论生成模型理性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。