arXiv:2608.04926cs.LGcs.AI2026-08

让图表、表格和代码三者自动对齐,提升跨模态理解能力。

Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning

论文配图:Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
图 1 · 摘自论文原文
  • 通过显式一对一对应关系,用表示间一致性驱动模型联合优化。
  • 在四个基准上,训练与推理时均显著提升跨模态任务性能。
  • 适合需要多模态对齐的可视化分析、自动化数据理解场景。

随着图表图像、表格数据和可视化代码在各领域的重要性不断提升,跨模态理解面临根本性挑战:模态间关系本质上为一对多,标注模糊且成本高昂,模型优化缺乏方向自适应且泛化性强的信号。本文提出 CoCoEvolve,旨在增强图表、表格与代码表示间的一致性。不同于将跨模态映射视为一对多问题,我们定义明确的一对一对应关系,并通过表示间的一致性进行模型优化,无需额外标注。训练阶段,CoCoEvolve@Train 在图表-表格-代码循环中执行协同进化;推理阶段,CoCoEvolve@Test 利用相同一致性目标实现测试时协同优化。此外,我们还构建了覆盖全部六类跨模态任务的评估套件 CoCoEvolve@Eval。在四个基准上的实验表明,CoCoEvolve 在训练与测试阶段均显著提升性能。

原文摘要 · Abstract (English)

As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross-representation understanding across these modalities poses fundamental challenges for AI systems: the relationships across representations are inherently \textit{one-to-many}, supervision is ambiguous and costly, and model optimization lacks a principled signal that is both direction-adaptive and representation-generalizable beyond task-specific objectives. We introduce CoCoEvolve to improve consistency across chart, table, and code representations. Instead of treating cross-representation mapping as a one-to-many problem, we define explicit one-to-one correspondences and optimize models using agreement between representations, without additional annotations. During training, CoCoEvolve@Train performs co-evolution across the chart-table-code cycle, while CoCoEvolve@Test applies the same consistency objective at inference time for test-time co-optimization. We also present CoCoEvolve@Eval, an evaluation suite covering all six cross-representation tasks. Across four benchmarks, CoCoEvolve improves performance in both training-time and test-time settings. Our project page: https://xhguo7.github.io/CoCoEvolve/.

跨模态学习自监督一致性优化可视化理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。