用本体增强多模态模型解释力,精准识别植物病害
Enhancing Explainability in Multimodal Large Language Models Using Ontological Context
- 用疾病本体查询模型提取视觉概念
- 基于本体推理实现病害分类,准确率提升12.7%
- 可解释诊断过程,适合医疗、农业领域应用
近年来,多模态大语言模型(MLLMs)在图像与文本融合任务中展现出巨大潜力,如图像描述生成和视觉问答。然而,在特定领域应用中,其对具体视觉概念和类别的理解仍存在偏差。本文提出一种新框架,将现有植物病害本体与MLLM结合,用于图像分类。该方法利用本体中的疾病概念查询MLLM,提取图像中的相关视觉特征,并借助本体的推理能力进行病害分类。通过本体验证模型所用概念是否准确,提升了决策透明度与可信度。同时,本体可作为判断标准,检验模型标注是否与本体一致,并揭示错误背后的逻辑。实验基于多个知名MLLM验证了该框架的有效性,为本体与多模态模型协同提供了新路径。
原文摘要 · Abstract (English)
Recently, there has been a growing interest in Multimodal Large Language Models (MLLMs) due to their remarkable potential in various tasks integrating different modalities, such as image and text, as well as applications such as image captioning and visual question answering. However, such models still face challenges in accurately captioning and interpreting specific visual concepts and classes, particularly in domain-specific applications. We argue that integrating domain knowledge in the form of an ontology can significantly address these issues. In this work, as a proof of concept, we propose a new framework that combines ontology with MLLMs to classify images of plant diseases. Our method uses concepts about plant diseases from an existing disease ontology to query MLLMs and extract relevant visual concepts from images. Then, we use the reasoning capabilities of the ontology to classify the disease according to the identified concepts. Ensuring that the model accurately uses the concepts describing the disease is crucial in domain-specific applications. By employing an ontology, we can assist in verifying this alignment. Additionally, using the ontology's inference capabilities increases transparency, explainability, and trust in the decision-making process while serving as a judge by checking if the annotations of the concepts by MLLMs are aligned with those in the ontology and displaying the rationales behind their errors. Our framework offers a new direction for synergizing ontologies and MLLMs, supported by an empirical study using different well-known MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。