让大模型学会读懂话中之意,识别隐含的逻辑关系。
Entailed Between the Lines: Incorporating Implication into NLI
- 将隐含蕴含引入自然语言推理任务,扩展传统NLI
- 构建INLI数据集,提升模型对隐含推理的识别能力
- 适合研究语义理解与大模型推理能力的学者
人类交流很大程度依赖于隐含意义,即通过字面之外的信息传递思想、意图和情感。为使模型更好理解并促进人际沟通,必须响应文本的隐含含义。本文聚焦自然语言推理(NLI)这一核心语言任务,发现当前最先进的NLI模型与数据集难以识别多种隐含蕴含的情形。为此,我们正式将隐含蕴含定义为NLI任务的拓展,并提出Implied NLI数据集(INLI),以帮助当前大语言模型识别更广泛的隐含蕴含,同时区分隐含与显式蕴含。实验表明,经INLI微调的模型能有效理解隐含蕴含,并在不同数据集与领域间实现泛化。
原文摘要 · Abstract (English)
Much of human communication depends on implication, conveying meaning beyond literal words to express a wider range of thoughts, intentions, and feelings. For models to better understand and facilitate human communication, they must be responsive to the text's implicit meaning. We focus on Natural Language Inference (NLI), a core tool for many language tasks, and find that state-of-the-art NLI models and datasets struggle to recognize a range of cases where entailment is implied, rather than explicit from the text. We formalize implied entailment as an extension of the NLI task and introduce the Implied NLI dataset (INLI) to help today's LLMs both recognize a broader variety of implied entailments and to distinguish between implicit and explicit entailment. We show how LLMs fine-tuned on INLI understand implied entailment and can generalize this understanding across datasets and domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。