发现并尝试解决无监督组合优化中训练与测试的不一致问题
On Training-Test (Mis)alignment in Unsupervised Combinatorial Optimization: Observation, Empirical Exploration, and Analysis
- 将可微化的去随机化过程引入训练阶段,增强训练与测试的一致性
- 实验证明该方法能提升训练-测试对齐度,但带来新的训练挑战
- 适合关注模型泛化性能与训练设计的研究者阅读
在无监督组合优化(UCO)中,训练阶段追求对每个训练实例生成具有概率优势的连续决策,以实现对初始离散且不可导问题的端到端训练。测试阶段则从连续决策出发,通过去随机化获得最终确定性解。现有方法不断改进测试时的去随机化策略以提升性能与理论保障。然而我们观察到,当前方法在训练与测试之间存在错位:更低的训练损失并不一定带来更好的去随机化后表现,即使在无数据分布偏移的训练实例上也是如此。我们实证发现了此类不良现象,并初步探索将可微去随机化纳入训练以改善对齐。实验表明该思路确实提升了对齐程度,但也引入了显著的训练难题。
原文摘要 · Abstract (English)
In unsupervised combinatorial optimization (UCO), during training, one aims to have continuous decisions that are promising in a probabilistic sense for each training instance, which enables end-to-end training on initially discrete and non-differentiable problems. At the test time, for each test instance, starting from continuous decisions, derandomization is typically applied to obtain the final deterministic decisions. Researchers have developed more and more powerful test-time derandomization schemes to enhance the empirical performance and the theoretical guarantee of UCO methods. However, we notice a misalignment between training and testing in the existing UCO methods. Consequently, lower training losses do not necessarily entail better post-derandomization performance, even for the training instances without any data distribution shift. Empirically, we indeed observe such undesirable cases. We explore a preliminary idea to better align training and testing in UCO by including a differentiable version of derandomization into training. Our empirical exploration shows that such an idea indeed improves training-test alignment, but also introduces nontrivial challenges into training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。