arXiv:2604.08324cs.NEcs.AI2026-04中稿 · PPSN 2026

探究多模态优化在符号回归中的对齐效果,发现现有方法实际对齐不足。

Multi-Modal Learning meets Genetic Programming: Analyzing Alignment in Latent Space Optimization

  • 用多模态编码器在潜在空间对齐符号与数值表达式,尝试提升搜索效率。
  • 实验表明优化过程中跨模态对齐未改善,且对齐过于粗糙。
  • 适合关注符号回归、潜在空间优化与模型对齐的研究者阅读。

符号回归(SR)旨在从数据中发现数学表达式,传统上通过遗传编程(GP)在符号结构中进行组合搜索。潜在空间优化(LSO)方法利用神经编码器将符号表达式映射到连续空间,将组合搜索转化为连续优化。SNIP(Meidani等,2024)受CLIP启发,提出一种多模态方法:在共享潜在空间中对齐符号与数值编码器,学习表型-基因型映射,使数值空间优化能隐式引导符号搜索。然而,该方法依赖细粒度的跨模态对齐,而类似模型(如CLIP)的文献表明此类对齐通常较粗粒度。本文研究发现:(1)优化过程中跨模态对齐并未提升,即使适应度增加;(2)SNIP学到的对齐过于粗糙,无法在符号空间中实现有原则的搜索。结果表明,尽管多模态LSO在符号回归中潜力巨大,但有效对齐引导的优化尚未实现,凸显细粒度对齐是未来关键方向。

原文摘要 · Abstract (English)

Symbolic regression (SR) aims to discover mathematical expressions from data, a task traditionally tackled using Genetic Programming (GP) through combinatorial search over symbolic structures. Latent Space Optimization (LSO) methods use neural encoders to map symbolic expressions into continuous spaces, transforming the combinatorial search into continuous optimization. SNIP (Meidani et al., 2024), a contrastive pre-training model inspired by CLIP, advances LSO by introducing a multi-modal approach: aligning symbolic and numeric encoders in a shared latent space to learn the phenotype-genotype mapping, enabling optimization in the numeric space to implicitly guide symbolic search. However, this relies on fine-grained cross-modal alignment, whereas literature on similar models like CLIP reveals that such an alignment is typically coarse-grained. In this paper, we investigate whether SNIP delivers on its promise of effective bi-modal optimization for SR. Our experiments show that: (1) cross-modal alignment does not improve during optimization, even as fitness increases, and (2) the alignment learned by SNIP is too coarse to efficiently conduct principled search in the symbolic space. These findings reveal that while multi-modal LSO holds significant potential for SR, effective alignment-guided optimization remains unrealized in practice, highlighting fine-grained alignment as a critical direction for future work.

符号回归多模态学习潜在空间优化对齐机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。