通过合成数据扩展任务空间,提升模型在开放世界中的泛化能力
Task Expansion and Cross Refinement for Open-World Conditional Modeling
- 用大模型生成多样化数据结构并弱标注,扩大任务覆盖范围
- 跨模型训练与反向修正合成数据,降低偏差、提升伪数据质量
- 适用于需要泛化到未知任务的条件建模场景
开放世界条件建模(OCM)要求单一模型在异构数据集上回答任意条件查询,其中观测变量和目标不断变化,来自一个无限的任务空间。由于真实数据集仅覆盖该空间的一小部分,本文提出任务扩展与交叉精炼(TEXR),一种半监督框架,通过结构化合成与精炼语义数据上下文来扩展有效任务覆盖。TEXR 首先利用大语言模型引导的结构化概率生成器,生成多样化的未实例化数据模式并弱实例化;随后在互不重叠的数据分区上进行跨模型训练,并在不同分区间迭代修正合成值,以减少确认偏差并提升伪值质量。最终将精炼后的合成数据与真实数据合并,训练统一的条件模型。在多个异构表格基准上,TEXR 持续提升多种 OCM 主干模型在零样本、少样本及多样本下的表现,证明了结构化任务扩展与交叉精炼能有效增强开放世界条件建模能力。
原文摘要 · Abstract (English)
Open-world conditional modeling (OCM), requires a single model to answer arbitrary conditional queries across heterogeneous datasets, where observed variables and targets vary and arise from a vast open-ended task universe. Because any finite collection of real-world datasets covers only a small fraction of this space, we propose Task Expansion and Cross Refinement (TEXR), a semi-supervised framework that enlarges effective task coverage through structured synthesis and refinement of semantic data contexts. TEXR first generates diverse uninstantiated dataset schemas and weakly instantiates them via structured probabilistic generators guided by large language models. It then performs cross-model refinement by training on disjoint data partitions and revising synthetic values across splits to reduce confirmation bias and improve pseudo-value quality. The refined synthetic datasets are aggregated with real data to train a unified conditional model. Across heterogeneous tabular benchmarks, TEXR consistently improves zero-, few-, and many-shot performance for multiple OCM backbones, demonstrating that structured task expansion and cross refinement enhance open-world conditional modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。