arXiv:2606.21915cs.CV2026-06

用博弈论显式对齐影像与报告,提升肺部X光报告的临床一致性。

GTA-Net: Cooperative Game Theory for Vision-Language Alignment in Chest X-Ray Report Generation

论文配图:GTA-Net: Cooperative Game Theory for Vision-Language Alignment in Chest X-Ray Report Generation
图 1 · 摘自论文原文
  • 将图文对齐建模为合作博弈,通过收益矩阵和重要性权重显式计算区域与词语关联。
  • 在CheXpertPlus和IU-XRay上生成报告的临床一致性显著优于现有方法。
  • 适合医学影像生成、临床辅助诊断等需要高可靠性的应用场景。

自动化胸部X光报告生成需要精确的跨模态对齐以确保临床可靠性。然而,现有视觉语言模型依赖隐式注意力机制,无法强制实现区域-词语对应关系和疾病级别的一致性。本文提出博弈论对齐网络(GTA-Net),将报告生成建模为合作博弈对齐问题。模型引入二元博弈对齐器,基于相似性收益矩阵与类谢林值的重要性加权,建模图像区域与文本词元间的交互。为进一步强化临床语义,还设计了疾病感知三元对齐器,捕捉图像、报告与结构化疾病概念间的联合交互。GTA-Net结合Swin视觉编码器与LoRA微调的大语言模型,采用统一目标函数进行生成与对齐联合训练。在CheXpertPlus和IU-XRay数据集上的实验表明,该模型在标准生成指标上达到当前最优性能,并显著提升临床一致性,验证了显式博弈论对齐在医疗视觉语言生成中的有效性。

原文摘要 · Abstract (English)

Automated chest X-ray report generation requires precise cross-modal grounding to ensure clinically reliable descriptions. However, existing vision-language models rely on implicit attention mechanisms that fail to enforce explicit region-word correspondence and disease-level consistency. We propose Game-Theoretic Alignment Network (GTA-Net), a vision-language framework that formulates report generation as a cooperative game-theoretic alignment problem. The model introduces a BinaryGameAligner that models interactions between image regions and text tokens using similarity-based payoff matrices with Shapley-inspired importance weighting. To enforce clinical semantics, we further develop a Disease-Aware Ternary Aligner, which captures joint interactions among images, reports, and structured disease concepts. GTA-Net combines a Swin-based visual encoder with a LoRA-adapted large language model and is trained with a unified objective for generation and alignment. Experiments on CheXpertPlus and IU-XRay demonstrate state-of-the-art performance across standard generation metrics and improved clinical consistency, highlighting the effectiveness of explicit game-theoretic alignment for medical vision-language generation.

视觉语言医学影像博弈论报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。