专为物理智能设计的多模态评判模型,提升感知与因果推理能力。
PhyCritic: Multimodal Critic Models for Physical AI
- 分两阶段训练:先强化物理感知,再自参照评判优化稳定性
- 在物理与通用多模态评测中超越开源基线模型
- 适合需要可靠判断的物理场景生成与决策任务
随着大型多模态模型的快速发展,可靠的判别与评判模型已成为开放式评估和偏好对齐的关键,可提供成对偏好、数值评分及解释性理由来评估模型生成结果。然而,现有评判模型主要在图像描述或图像问答等通用视觉领域训练,对涉及感知、因果推理与规划的物理人工智能任务仍缺乏深入探索。我们提出 PhyCritic,一种通过双阶段强化学习与验证(RLVR)流程优化的多模态评判模型:第一阶段为物理技能预热,增强物理导向的感知与推理能力;第二阶段为自参照评判微调,评判模型在判断候选响应前先生成自身预测作为内部参考,从而提升判断稳定性和物理正确性。在物理与通用多模态评判基准测试中,PhyCritic 均显著优于开源基线模型;当用作策略模型时,进一步提升了物理情境下任务的感知与推理表现。
原文摘要 · Abstract (English)
With the rapid development of large multimodal models, reliable judge and critic models have become essential for open-ended evaluation and preference alignment, providing pairwise preferences, numerical scores, and explanatory justifications for assessing model-generated responses. However, existing critics are primarily trained in general visual domains such as captioning or image question answering, leaving physical AI tasks involving perception, causal reasoning, and planning largely underexplored. We introduce PhyCritic, a multimodal critic model optimized for physical AI through a two-stage RLVR pipeline: a physical skill warmup stage that enhances physically oriented perception and reasoning, followed by self-referential critic finetuning, where the critic generates its own prediction as an internal reference before judging candidate responses, improving judgment stability and physical correctness. Across both physical and general-purpose multimodal judge benchmarks, PhyCritic achieves strong performance gains over open-source baselines and, when applied as a policy model, further improves perception and reasoning in physically grounded tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。