用图像生成语言含义,实现零样本自然语言推理。
Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- 用文生图模型将前提转为视觉表征,再与假设对比
- 无需微调即达高准确率,对文本偏见有强鲁棒性
- 适合追求可靠语义理解的AI研究者
我们提出一种零样本自然语言推理方法,通过将语言在视觉上下文中进行定位来利用多模态表征。该方法使用文生图模型生成前提的视觉表征,并通过比较这些表征与文本假设来进行推理。我们评估了两种推理技术:余弦相似度和视觉问答。该方法在无需任务特定微调的情况下实现了高准确率,表现出对文本偏差和表面启发式方法的鲁棒性。此外,我们设计了一个受控对抗数据集以验证方法的鲁棒性。研究结果表明,将视觉模态作为语义表征具有潜力,可推动自然语言理解的可靠性发展。
原文摘要 · Abstract (English)
We propose a zero-shot method for Natural Language Inference (NLI) that leverages multimodal representations by grounding language in visual contexts. Our approach generates visual representations of premises using text-to-image models and performs inference by comparing these representations with textual hypotheses. We evaluate two inference techniques: cosine similarity and visual question answering. Our method achieves high accuracy without task-specific fine-tuning, demonstrating robustness against textual biases and surface heuristics. Additionally, we design a controlled adversarial dataset to validate the robustness of our approach. Our findings suggest that leveraging visual modality as a meaning representation provides a promising direction for robust natural language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。