让表格模型理解数据背后的业务逻辑,提升智能系统可靠性
Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding
- 用代码与数据对训练模型,学习业务规则运作机制
- 在私有数据上实现零样本推理,准确率显著优于传统方法
- 适合构建能自主理解数据的智能决策系统
本文指出,孤立的表格基础模型存在根本性局限——缺乏对数据生成与管理过程中的操作逻辑、声明规则和领域知识的感知。当前方法仅关注单表泛化或模式级关系,未能捕捉赋予数据意义的操作背景。我们提出语义关联表(SLT)与面向SLT的基础模型(FMSLT),通过双阶段训练:先在开源代码-数据对及合成系统上预训练,学习业务逻辑机制;再在专有数据上进行零样本推理。引入‘操作图灵测试’基准,论证操作性接地对复杂数据环境中自主代理的重要性。
原文摘要 · Abstract (English)
This position paper argues that foundation models for tabular data face inherent limitations when isolated from operational context - the procedural logic, declarative rules, and domain knowledge that define how data is created and governed. Current approaches focus on single-table generalization or schema-level relationships, fundamentally missing the operational knowledge that gives data meaning. We introduce Semantically Linked Tables (SLT) and Foundation Models for SLT (FMSLT) as a new model class that grounds tabular data in its operational context. We propose dual-phase training: pre-training on open-source code-data pairs and synthetic systems to learn business logic mechanics, followed by zero-shot inference on proprietary data. We introduce the ``Operational Turing Test'' benchmark and argue that operational grounding is essential for autonomous agents in complex data environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。