用机器生成图像增强语言模型的零样本常识推理能力
Enhancing Zero-shot Commonsense Reasoning by Integrating Visual Knowledge via Machine Imagination
- 将图像生成嵌入推理流程,让语言模型‘想象’视觉场景
- 在多个基准上超越现有零样本方法,接近大模型表现
- 适合研究多模态推理与降低文本偏见的学者
近期零样本常识推理进展使预训练语言模型(PLMs)无需特定任务微调即可获取广泛常识知识。然而,这些模型常受文本知识中人类报告偏差的影响,导致机器与人类理解存在差异。为弥合这一差距,我们引入视觉模态以增强PLMs的推理能力。提出Imagine(基于机器想象的推理)框架,通过在推理流程中集成图像生成器,使PLMs具备‘想象’能力,并补充文本输入的视觉信号。为有效利用生成的视觉上下文,构建了模拟视觉问答场景的合成数据集。在多个常识推理基准上的全面评估表明,Imagine显著优于现有零样本方法,甚至超越部分先进大语言模型。结果证明,机器想象能有效缓解报告偏差,显著提升常识推理模型的泛化能力。
原文摘要 · Abstract (English)
Recent advancements in zero-shot commonsense reasoning have empowered Pre-trained Language Models (PLMs) to acquire extensive commonsense knowledge without requiring task-specific fine-tuning. Despite this progress, these models frequently suffer from limitations caused by human reporting biases inherent in textual knowledge, leading to understanding discrepancies between machines and humans. To bridge this gap, we introduce an additional modality to enrich the reasoning capabilities of PLMs. We propose Imagine (Machine Imagination-based Reasoning), a novel zero-shot commonsense reasoning framework that supplements textual inputs with visual signals from machine-generated images. Specifically, we enhance PLMs with the ability to imagine by embedding an image generator directly into the reasoning pipeline. To facilitate effective utilization of this imagined visual context, we construct synthetic datasets designed to emulate visual question-answering scenarios. Through comprehensive evaluations on multiple commonsense reasoning benchmarks, we demonstrate that Imagine substantially outperforms existing zero-shot approaches and even surpasses advanced large language models. These results underscore the capability of machine imagination to mitigate reporting bias and significantly enhance the generalization ability of commonsense reasoning models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。