构建基因关系推理基准,测试大模型跨多步逻辑推断能力
Kinship Data Benchmark for Multi-hop Reasoning
- 用生成式管道构建符合不同文化规则的家族树数据
- 六款主流大模型在零样本下表现差异显著,揭示推理能力鸿沟
- 适合评估大模型在复杂亲属关系中的隐含推理能力
大型语言模型(LLMs)在多跳推理能力上的评估日益重要,即通过整合多个信息片段完成连贯推断。我们提出KinshipQA,一个针对亲属关系推理设计的基准测试。核心贡献是一个生成式流程,可按需生成大规模、真实且具有文化特异性的家谱数据:满足特定亲属制度婚姻约束的互联家族树集合。这使得任务难度、文化假设和关系深度可系统控制与调整。基于这些家谱,我们构建需要推理隐含关系链的文本推断任务。使用六款先进大模型(涵盖开源与闭源)在统一的零样本协议和确定性解码下进行评估,采用精确匹配和集合匹配指标。结果表明,KinshipQA展现出广泛的表现差异,并暴露了模型间及文化背景下的系统性推理差异。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly evaluated on their ability to perform multi-hop reasoning, i.e., to combine multiple pieces of information into a coherent inference. We introduce KinshipQA, a benchmark designed to probe this capability through reasoning over kinship relations. The central contribution of our work is a generative pipeline that produces, on demand, large-scale, realistic, and culture-specific genealogical data: collections of interconnected family trees that satisfy explicit marriage constraints associated with different kinship systems. This allows task difficulty, cultural assumptions, and relational depth to be systematically controlled and varied. From these genealogies, we derive textual inference tasks that require reasoning over implicit relational chains. We evaluate the resulting benchmark using six state-of-the-art LLMs, spanning both open-source and closed-source models, under a uniform zero-shot protocol with deterministic decoding. Performance is measured using exact-match and set-based metrics. Our results demonstrate that KinshipQA yields a wide spread of outcomes and exposes systematic differences in multi-hop reasoning across models and cultural settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。