通过段落级预训练,提升异构表格的语义一致性与结构理解。
Segment-driven Structural Induction and Semantic Alignment for Heterogeneous Tabular Representation

- 以标题-值对为单元,融合模式与分布双重证据
- 在多场景下实现更优的重建与下游任务表现
- 适合处理跨领域异构表格的表示学习
现实世界中的异构表格往往具有变化的表头,但其底层属性语义共享,仅依靠表内局部证据难以推导领域专属语义。现有编码器虽部分解决此问题,但常忽视列级取值分布,并对不同语义角色的属性采用统一目标。我们提出NAVI,一种以段落为中心的预训练框架,将每个标题-值对视为聚合模式级结构证据与列级分布证据的单元。通过掩码段建模与熵驱动段对齐,联合强化标题-值间的结构耦合与跨稳定与实例特异性属性的语义对齐。在异构域内表格上的实验表明,该方法在多种评估设置下均提升了重构精度、语义一致性和下游实用性。
原文摘要 · Abstract (English)
Real-world domains often contain heterogeneous tables whose headers vary while their underlying attribute semantics are shared, making it difficult to induce domain-specialized semantics from table-local evidence alone. Existing encoders model parts of this problem, but often underuse column-level value distributions and apply uniform objectives across attributes with different semantic roles. We propose NAVI, a segment-centric pretraining framework that treats each header-value pair as the unit for aggregating schema-level structural evidence and column-level distributional evidence. We realize this design through Masked Segment Modeling and Entropy-driven Segment Alignment, which jointly enforce structured header-value coupling and semantic alignment across stable and instance-specific attributes. Experiments on heterogeneous in-domain tables show improved reconstruction, semantic consistency, and downstream utility across evaluation settings overall.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。