arXiv:2606.16870cs.CVcs.GR2026-06中稿 · IEEE/CVF Conferenc…

用强化学习从果皮断裂行为反推食材物理参数,10毫秒内完成精准估计。

Latent Space Reinforcement Learning for Inverse Material Estimation in Food Fracture Simulation

论文配图:Latent Space Reinforcement Learning for Inverse Material Estimation in Food Fracture Simulation
图 1 · 摘自论文原文
  • 在隐空间中训练强化学习策略,实现从断裂描述到材料参数的端到端逆向映射。
  • 在9维原空间上达到0.642的参数恢复率,比传统方法提升23%。
  • 支持任意目标无需重训,适合视频驱动的食材材质识别应用。

真实模拟食物操作需要精确的材料参数,但这些参数难以直接测量,且在同一食物内部存在异质性。本文针对非可微分连续损伤力学模拟器中的逆问题,基于橙子剥皮场景进行研究。利用2000次前向仿真训练神经代理模型,并在原始9维参数空间及两种学习得到的4维隐空间中,对比协方差矩阵自适应进化策略(CMA-ES)与近端策略优化(PPO)算法。由于不同橙子材料属性各异,实际系统需在不重新训练的前提下处理任意目标。为此,我们训练了一个目标条件化的PPO策略,仅需一次前向传播(8次代理评估,约10毫秒)即可生成对应材料参数估计。在归一化流隐空间中使用共享代理评估器,该策略在模拟器验证下实现0.642的实际恢复率,比原空间提升23%。通过冷启动扩展,将CMA-ES初始化于策略输出,进一步提升恢复率达0.828(540次评估)。研究为食品物理逆问题提供了实用框架,也为从食物操作视频中实现视觉驱动的材质识别奠定基础。

原文摘要 · Abstract (English)

Realistic visual simulation of food manipulation requires accurate material parameters, yet these are difficult to measure directly and vary across the heterogeneous regions of a single food item. We address the inverse problem of estimating material parameters from a target description of fracture behavior in a non-differentiable continuum damage mechanics simulator. Using orange peeling as a test case, we train a neural surrogate on 2,000 forward simulations and compare Covariance Matrix Adaptation Evolution Strategy (CMA-ES, a gradient-free evolutionary optimizer) with Proximal Policy Optimization (PPO, a reinforcement learning algorithm) across the original 9-dimensional parameter space and two learned 4-dimensional latent representations. Since different oranges have different material properties, a practical inverse system must handle arbitrary targets without retraining. We train a goal-conditioned PPO policy that learns a general inverse mapping: given any target description of peeling behavior, the policy produces a material parameter estimate in a single forward pass (8 surrogate evaluations, approximately 10ms). Operating in a normalizing flow latent space with a shared surrogate evaluator, the goal-conditioned policy achieves 0.642 actual recovery when validated through the simulator, outperforming the original parameter space by 23%. A warm-start extension that initializes CMA-ES refinement from the policy's output further improves recovery to 0.828 with 540 evaluations. These findings provide a practical framework for inverse food physics and lay groundwork for vision-driven material identification from video observations of food manipulation.

逆问题强化学习食品模拟隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。