用机器生成的图像增强语言模型推理,缓解常识知识偏差。
Zero-shot Commonsense Reasoning over Machine Imagination
- 将图像生成器融入语言模型推理,用视觉信号补充文本输入。
- 在多个基准上显著优于现有方法,提升常识推理准确率。
- 适合研究多模态推理、降低数据偏见的AI从业者。
零样本常识推理近期进展使预训练语言模型(PLMs)能在不针对特定场景的情况下学习广泛常识知识。然而,这些方法常受文本常识知识中的人类报告偏差影响,导致模型与人类理解存在差异。本文提出一种名为Imagine(基于机器想象的推理)的新框架,通过引入视觉信号来弥补这一差距。该框架将图像生成器嵌入推理过程,赋予PLMs想象能力,并构建合成预训练数据集以模拟视觉问答任务。大量实验和分析表明,Imagine在多个推理基准上显著超越现有方法,验证了机器想象在缓解报告偏差和提升泛化能力方面的有效性。
原文摘要 · Abstract (English)
Recent approaches to zero-shot commonsense reasoning have enabled Pre-trained Language Models (PLMs) to learn a broad range of commonsense knowledge without being tailored to specific situations. However, they often suffer from human reporting bias inherent in textual commonsense knowledge, leading to discrepancies in understanding between PLMs and humans. In this work, we aim to bridge this gap by introducing an additional information channel to PLMs. We propose Imagine (Machine Imagination-based Reasoning), a novel zero-shot commonsense reasoning framework designed to complement textual inputs with visual signals derived from machine-generated images. To achieve this, we enhance PLMs with imagination capabilities by incorporating an image generator into the reasoning process. To guide PLMs in effectively leveraging machine imagination, we create a synthetic pre-training dataset that simulates visual question-answering. Our extensive experiments on diverse reasoning benchmarks and analysis show that Imagine outperforms existing methods by a large margin, highlighting the strength of machine imagination in mitigating reporting bias and enhancing generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。