用人类反馈指导生成数据,提升模型在未知环境下的泛化能力。
Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm
- 通过条件神经微分方程生成数据,结合强化学习优化物理参数
- 专家反馈越可靠,部署损失越低,生成数据与目标分布差距越小
- 适用于动态与非动态系统,适合需要人机协同的鲁棒学习场景
将机器学习模型推广到与训练分布不同的环境仍是关键挑战,尤其当目标域数据完全或部分不可得时。我们提出生成式元学习结合人类反馈(GMHF)框架,通过专家直觉引导数据合成来弥合领域差异。基于泛化误差的理论分析,我们推导出边界,表明使生成数据分布与人类对目标物理规律的信念一致可显著降低风险。GMHF利用条件神经微分方程(cNODE)作为生成式数字孪生,并结合强化学习(RL)代理,根据反馈迭代优化生成轨迹的潜在物理参数,有效引导元学习器逼近未观测的目标分布。在非线性Duffing振子上的实证验证表明,随着专家可靠性提高,部署损失显著下降,且生成数据与目标数据的分歧随之减小,直接验证了理论预测的分歧最小化机制。在非动力学概率模型上的进一步实验确认该框架可扩展至非微分方程驱动系统,确立人机协作作为分布偏移下鲁棒泛化的严谨催化剂。
原文摘要 · Abstract (English)
Generalizing machine learning models to environments that differ from their training distribution remains a critical hurdle, particularly when data from the target domain is entirely or partially unavailable. We propose Generative Meta-Learning with Human Feedback (GMHF), a novel framework that bridges this domain gap by leveraging expert intuition to guide data synthesis. Grounded in a theoretical analysis of generalization error, we derive bounds demonstrating that aligning the distribution of generated data with human beliefs regarding the target physics significantly mitigates risk. GMHF operationalizes this insight by employing a Conditional Neural ODE (cNODE) as a generative digital twin, coupled with a Reinforcement Learning (RL) agent. The agent iteratively refines the latent physical parameters of the generated trajectories based on feedback, effectively steering the meta-learner toward the unobserved target distribution. Empirical validation on a nonlinear Duffing oscillator shows that GMHF substantially reduces deployment loss as expert reliability increases, and that the divergence between generated and target data falls under reliable feedback, directly corroborating the divergence-minimisation mechanism predicted by our theory. Further experiments on a non-dynamical probabilistic model confirm that the framework extends beyond ODE-governed systems, establishing human-AI collaboration as a rigorous catalyst for robust generalisation under distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。