用符号推理动态分配专家,让AI更懂复杂表格的结构与语义。
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
- 基于符号推理的路由机制,按语义角色分配表格元素到专用专家
- 在4个真实场景挑战集上超越现有模型,最高提升17.3%准确率
- 适合需要理解金融、医疗等复杂表格的开发者与研究者
真实场景中的多模态表格理解因结构复杂、符号密集和视觉退化(模糊、倾斜、水印、不完整结构或字体、多跨度或嵌套布局)而极具挑战。现有多模态大模型在野化结构(WildStruct)条件下表现受限,泛化能力差。为此,我们提出TableMoE,一种专为多模态表格结构化推理设计的神经符号混合专家(MoCE)架构。其创新的神经符号路由机制可预测隐含语义标记角色(如表头、数据单元格、坐标轴、公式),并基于符号推理图的置信度感知门控策略,将表格元素动态路由至专用专家(表转HTML、JSON、代码)。为支持对齐预训练,我们构建了大规模的TableMoE-Align数据集,包含120万条跨金融、科学、生物医学和工业领域的表格-HTML-JSON-代码四元组。评估方面,我们精心构建并发布了四个挑战性野化结构基准:WMMFinQA、WMMTatQA、WMMTabDialog和WMMFinanceMath,专门测试模型在真实多模态退化和结构复杂性下的表现。实验表明,TableMoE显著优于现有最先进模型。大量消融实验验证了核心组件的有效性,凸显神经符号路由与结构化专家对齐的关键作用。定性分析进一步展示了TableMoE的可解释性与增强鲁棒性,证实神经符号推理在多模态表格理解中的有效性。
原文摘要 · Abstract (English)
Multimodal understanding of tables in real-world contexts is challenging due to the complexity of structure, symbolic density, and visual degradation (blur, skew, watermarking, incomplete structures or fonts, multi-span or hierarchically nested layouts). Existing multimodal large language models (MLLMs) struggle with such WildStruct conditions, resulting in limited performance and poor generalization. To address these challenges, we propose TableMoE, a neuro-symbolic Mixture-of-Connector-Experts (MoCE) architecture specifically designed for robust, structured reasoning over multimodal table data. TableMoE features an innovative Neuro-Symbolic Routing mechanism, which predicts latent semantic token roles (e.g., header, data cell, axis, formula) and dynamically routes table elements to specialized experts (Table-to-HTML, Table-to-JSON, Table-to-Code) using a confidence-aware gating strategy informed by symbolic reasoning graphs. To facilitate effective alignment-driven pretraining, we introduce the large-scale TableMoE-Align dataset, consisting of 1.2M table-HTML-JSON-code quadruples across finance, science, biomedicine and industry, utilized exclusively for model pretraining. For evaluation, we curate and release four challenging WildStruct benchmarks: WMMFinQA, WMMTatQA, WMMTabDialog, and WMMFinanceMath, designed specifically to stress-test models under real-world multimodal degradation and structural complexity. Experimental results demonstrate that TableMoE significantly surpasses existing state-of-the-art models. Extensive ablation studies validate each core component, emphasizing the critical role of Neuro-Symbolic Routing and structured expert alignment. Through qualitative analyses, we further showcase TableMoE's interpretability and enhanced robustness, underscoring the effectiveness of integrating neuro-symbolic reasoning for multimodal table understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。