用模糊推理增强多模态问答,提升准确性和可解释性。
Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework
- 引入模糊规则与多跳推理,实现跨领域知识融合
- 在多个数据集上优于现有方法,答案更准确一致
- 适合需要深度推理和不确定处理的多模态应用
在多模态问答(MQA)中,模型需联合编码并整合文本、图像、语音等异构信息,以进行复杂语义推理与决策。尽管已有进展,传统深度学习模型及大模型或提示框架仍面临三大挑战:不同模态间特征分布差异导致模态偏差;问题涉及多领域知识带来显著不确定性;现有方法多依赖浅层语义匹配,推理深度与可解释性不足。为此,受传统模糊系统(FS)启发,提出多模态生成模糊系统(MMGFS),通过多模态协同反思机制缓解模态偏差,并引入模糊规则与多跳推理机制,支持跨域知识融合与分层推理,强化不确定性建模与深层语义理解。在开放域数据集MultimodalQA、WebQA及领域特定基准BioMol-VQA、EHRxQA上进行评估,结果表明MMGFS在多个数据集上持续优于现有方法,有效缓解模态偏差与问题不确定性,显著提升答案准确率、一致性和泛化能力。
原文摘要 · Abstract (English)
In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalities, including text, images, and speech, to perform complex semantic reasoning and decision making. Despite recent advances, existing approaches, including traditional deep learning models and Large Models (LMs) or prompt-based frameworks, continue to face several critical challenges. First, modality bias arises from discrepancies in feature distributions across different modalities, which limits effective cross modal collaborative understanding. Second, many questions require knowledge drawn from multiple domains, introducing significant uncertainty. Third, current methods often rely on shallow semantic matching, resulting in limited reasoning depth an reduced interpretability. To address these issues, inspired by the traditional fuzzy system (FS) framework, we propose a fuzzy-inference-guided multimodal generative architecture termed the Multi-Modal Generative Fuzzy System (MMGFS). The main contributions of MMGFS are two folds. First, it alleviates modality bias through a multimodal collaborative rumination mechanism. Second, it introduces fuzzy rules and a multi-hop inference mechanism to support cross-domain knowledge fusion and hierarchical reasoning, thereby strengthening uncertainty modelling and deepening semantic understanding. We conduct comprehensive evaluations on open-domain question answering datasets, including MultimodalQA and WebQA, as well as domain-specific benchmarks, including BioMol-VQA and EHRxQA. Experimental results demonstrate that MMGFS consistently outperforms existing methods across multiple datasets. It effectively mitigates modality bias and question uncertainty while achieving superior performance in answer accuracy, consistency, and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。