用混合马尔可夫逻辑网络解释图像描述生成背后的推理过程
On Explaining Visual Captioning with Hybrid Markov Logic Networks
- 用混合马尔可夫逻辑网络建模训练样本对生成结果的影响
- 通过分布变化量化哪些训练样例影响了当前描述生成
- 适合关注模型可解释性的研究者和开发者
深度神经网络在多模态任务如图像描述生成中取得显著进展,但解释模型如何融合视觉、语言与知识信息生成合理描述仍具挑战。传统评估依赖生成结果与人工标注的对比,难以揭示内部整合机制。本文提出基于混合马尔可夫逻辑网络(HMLN)的新解释框架,该语言可结合符号规则与实值函数。我们假设训练数据中的相关样本可能影响了最终描述的生成,并学习训练实例上的HMLN分布。通过条件化于生成样本,推断分布变化,从而量化哪些训练样本提供了更丰富的信息。在多个先进模型上使用Amazon Mechanical Turk进行实验,验证了该方法在解释性上的有效性,并实现了模型间可解释性的比较。
原文摘要 · Abstract (English)
Deep Neural Networks (DNNs) have made tremendous progress in multimodal tasks such as image captioning. However, explaining/interpreting how these models integrate visual information, language information and knowledge representation to generate meaningful captions remains a challenging problem. Standard metrics to measure performance typically rely on comparing generated captions with human-written ones that may not provide a user with a deep insights into this integration. In this work, we develop a novel explanation framework that is easily interpretable based on Hybrid Markov Logic Networks (HMLNs) - a language that can combine symbolic rules with real-valued functions - where we hypothesize how relevant examples from the training data could have influenced the generation of the observed caption. To do this, we learn a HMLN distribution over the training instances and infer the shift in distributions over these instances when we condition on the generated sample which allows us to quantify which examples may have been a source of richer information to generate the observed caption. Our experiments on captions generated for several state-of-the-art captioning models using Amazon Mechanical Turk illustrate the interpretability of our explanations, and allow us to compare these models along the dimension of explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。