arXiv:2410.10449cs.CL2024-10EMNLP被引 7

构建真实贝叶斯推理数据集,评估模型对语言化不确定性的量化能力。

QUITE: Quantifying Uncertainty in Natural Language Text in Bayesian Reasoning Scenarios

  • 提出自然语言形式的不确定性前提与证据,要求输出概率估计。
  • 逻辑模型在因果、证据和解释性推理上全面优于大模型。
  • 适合研究复杂推理与神经符号系统的研究者使用。

推理是许多决策过程的核心,需整合带有不同程度不确定性的规则式前提与观察结果以得出结论。现有概率推理数据集常简化任务,如仅要求排序文本选项、仅含二元随机变量或使用有限模板导致文本多样性不足。本文提出QUITE,一个包含类别型随机变量与复杂关系的真实世界贝叶斯推理问答数据集。QUITE提供高质量的自然语言前提描述与证据陈述,并要求以概率值回答问题。实验表明,基于逻辑的模型在所有推理类型(因果、证据、解释性)上均优于现成的大语言模型。结果支持神经符号模型是提升复杂推理能力的有前途方向。数据集与代码已开源至GitHub。

原文摘要 · Abstract (English)

Reasoning is key to many decision making processes. It requires consolidating a set of rule-like premises that are often associated with degrees of uncertainty and observations to draw conclusions. In this work, we address both the case where premises are specified as numeric probabilistic rules and situations in which humans state their estimates using words expressing degrees of certainty. Existing probabilistic reasoning datasets simplify the task, e.g., by requiring the model to only rank textual alternatives, by including only binary random variables, or by making use of a limited set of templates that result in less varied text. In this work, we present QUITE, a question answering dataset of real-world Bayesian reasoning scenarios with categorical random variables and complex relationships. QUITE provides high-quality natural language verbalizations of premises together with evidence statements and expects the answer to a question in the form of an estimated probability. We conduct an extensive set of experiments, finding that logic-based models outperform out-of-the-box large language models on all reasoning types (causal, evidential, and explaining-away). Our results provide evidence that neuro-symbolic models are a promising direction for improving complex reasoning. We release QUITE and code for training and experiments on Github.

贝叶斯推理不确定性量化自然语言推理神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。