arXiv:2604.12133cs.AI2026-04

提出表格表征的柏拉图假设,让模型不被排版顺序干扰。

Towards Platonic Representation for Table Reasoning: A Foundation for Permutation-Invariant Retrieval

  • 用几何不变性原则重构表格编码,避免线性化导致的结构失真。
  • 发现大模型对表格布局微小变化敏感,嵌入向量剧烈漂移。
  • 新模型通过单元格头对齐机制,实现更稳定的语义检索。

传统表格表示学习多沿用自然语言处理的序列范式,但这种线性化会丢失表格固有的几何与关系结构,使表征对布局排列极为脆弱。本文提出表格的柏拉图表征假设(PRH),主张语义稳健的表征空间必须具备内在排列不变性(PI)。我们通过回顾分析表推理任务,揭示了普遍存在的序列化偏差。为此提出基于中心核对齐(CKA)的两个度量:(i) PI,衡量完全结构打乱下的嵌入漂移;(ii) rho,基于斯皮尔曼相关性的指标,追踪潜空间随结构逐步恢复向标准形式的收敛过程。实证表明,现代大语言模型即使在微小布局扰动下,表征嵌入也产生显著且不成比例的语义偏移,暴露了检索增强生成(RAG)系统对布局噪声的脆弱性。为此,我们设计一种新型结构感知的表征编码器,显式建模单元格头对齐的认知机制,在几何稳定性上优于现有方法,逼近理想的排列不变性。本工作既批判了线性化表征编码器,也为语义稳定、排列不变的检索提供了理论基础,开辟了表推理的新方向。

原文摘要 · Abstract (English)

Historical approaches to Table Representation Learning (TRL) have largely adopted the sequential paradigms of Natural Language Processing (NLP). We argue that this linearization of tables discards their essential geometric and relational structure, creating representations that are brittle to layout permutations. This paper introduces the Platonic Representation Hypothesis (PRH) for tables, positing that a semantically robust latent space for table reasoning must be intrinsically Permutation Invariant (PI). To ground this hypothesis, we first conduct a retrospective analysis of table-reasoning tasks, highlighting the pervasive serialization bias that compromises structural integrity. We then propose a formal framework to diagnose this bias, introducing two principled metrics based on Centered Kernel Alignment (CKA): (i) PI, which measures embedding drift under complete structural derangement, and (ii) rho, a Spearman-based metric that tracks the convergence of latent structures toward a canonical form as structural information is incrementally restored. Our empirical analysis quantifies an expected flaw in modern Large Language Models (LLMs): even minor layout permutations induce significant, disproportionate semantic shifts in their table embeddings. This exposes a fundamental vulnerability in RAG systems, in which table retrieval becomes fragile to layout-dependent noise rather than to semantic content. In response, we present a novel, structure-aware TRL encoder architecture that explicitly enforces the cognitive principle of cell header alignment. This model demonstrates superior geometric stability and moves towards the PI ideal. Our work provides both a foundational critique of linearized table encoders and the theoretical scaffolding for semantically stable, permutation invariant retrieval, charting a new direction for table reasoning in information systems.

表格推理排列不变语义稳定大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。