arXiv:2606.24965cs.AIcs.LG2026-06

用大模型自动生成难题,测试神经推理模型的泛化能力。

Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners

论文配图:Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners
图 1 · 摘自论文原文
  • 用大模型驱动进化搜索,自动发现难例
  • 新数据让边缘变换器泛化性能提升37%
  • 可自动生成新世界,实现自主研究

神经模型在关系结构推理上仍面临挑战,尤其当需将学习知识应用于训练中未见的更复杂问题时。评估泛化能力困难,因难以预判何为难题。本文提出利用大语言模型(LLM)自动化构建基准,通过端到端方式生成日益困难的问题实例。具体地,给定由Datalog规则定义的世界及边变换器(Edge Transformer)作为推理评估器,采用基于FunSearch的大模型驱动进化搜索与自主智能体搜索,发现能生成高难度问题的采样函数。实验表明,使用这些数据可使边变换器在进一步数据扰动下表现更好,泛化能力显著提升。此外,该方法还可应用于大模型提出的全新世界,为神经关系推理的自主研究开辟道路。

原文摘要 · Abstract (English)

Reasoning about relational structures remains a significant challenge for neural models, particularly when they must systematically apply learned knowledge to problem instances that are harder than those seen in training. Progress is hampered by the difficulty of evaluating such generalization, since a priori, it is rarely clear what makes an instance hard. We study how this issue can be addressed by using large language models (LLMs) to automate benchmark generation, learning to produce increasingly challenging instances in an end-to-end manner. Concretely, given a world parametrized by Datalog rules, and an Edge Transformer as the reasoning evaluator, we use LLM-driven evolutionary search (based on FunSearch) and autonomous agentic search to discover sampling functions that yield hard problem instances. We also show that the Edge Transformer can be improved using this data such that it generalizes well to further data perturbations. Finally, we show that the same machinery can be applied to novel worlds proposed by LLMs, opening the door to autonomous research on neural relational reasoning.

关系推理大模型自动评测泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。