构建多分辨率空间推理评估框架,揭示神经网络在几何拓扑上的系统性缺陷
A Multi-Resolution Benchmark Framework for Spatial Reasoning Assessment in Neural Networks
- 用VoxLogicA生成两类合成数据:迷宫连通性(拓扑)与距离计算(几何)
- 在多分辨率下测试,发现神经网络在基础空间任务上表现不佳,指标偏低
- 提供可复现流程,适合研究医疗图像中神经网络空间理解的改进方向
本文提出一个全面的基准评估框架,用于系统性评测神经网络的空间推理能力,重点关注连通性与距离关系等形态属性。该框架利用空间模型检测器VoxLogicA生成两类合成数据集:用于拓扑分析的迷宫连通性问题和用于几何理解的距离计算任务。每类任务在多个分辨率下进行评估,以检验模型的可扩展性与泛化能力。自动化流程涵盖合成数据生成、标准化交叉验证训练、推理执行及基于Dice系数和IoU(交并比)的综合评估。初步实验表明,神经网络在基本几何与拓扑理解任务中存在系统性失败。该框架提供可复现的实验协议,帮助研究人员识别具体局限。未来可通过结合神经网络与符号推理的混合方法,在临床应用中提升空间理解能力,为持续研究神经网络空间推理缺陷及其解决方案奠定基础。
原文摘要 · Abstract (English)
This paper presents preliminary results in the definition of a comprehensive benchmark framework designed to systematically evaluate spatial reasoning capabilities in neural networks, with a particular focus on morphological properties such as connectivity and distance relationships. The framework is currently being used to study the capabilities of nnU-Net, exploiting the spatial model checker VoxLogicA to generate two distinct categories of synthetic datasets: maze connectivity problems for topological analysis and spatial distance computation tasks for geometric understanding. Each category is evaluated across multiple resolutions to assess scalability and generalization properties. The automated pipeline encompasses a complete machine learning workflow including: synthetic dataset generation, standardized training with cross-validation, inference execution, and comprehensive evaluation using Dice coefficient and IoU (Intersection over Union) metrics. Preliminary experimental results demonstrate significant challenges in neural network spatial reasoning capabilities, revealing systematic failures in basic geometric and topological understanding tasks. The framework provides a reproducible experimental protocol, enabling researchers to identify specific limitations. Such limitations could be addressed through hybrid approaches combining neural networks with symbolic reasoning methods for improved spatial understanding in clinical applications, establishing a foundation for ongoing research into neural network spatial reasoning limitations and potential solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。