arXiv:2511.12365cs.CV2025-11被引 2

用强化学习构建数字孪生表示,统一解决多模态视觉推理问题

Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning

  • 通过强化学习训练大模型构建视觉输入的数字孪生表示
  • 在六项基准测试中超越现有专用模型,提升显著
  • 适合需要跨任务、跨模态通用推理的研究者

视觉推理需模型理解图像视频并响应隐含文本查询,输出形式涵盖像素级分割图到自然语言描述。现有方法依赖特定任务的监督微调与架构设计,如推理分割、定位、摘要和视觉问答均需独立建模与训练,难以实现统一解决方案,限制跨任务与跨模态泛化。为此,我们提出DT-R1,一种基于强化学习的框架,让大语言模型构建复杂多模态视觉输入的数字孪生表示,并以此高阶表征为统一基础进行视觉推理。具体而言,采用GRPO训练并设计新奖励函数,同时验证结构完整性和输出准确性。在覆盖两种模态和四种任务类型的六个视觉推理基准上评估,DT-R1持续优于当前最优的专用模型。该工作开启了一条新路径:视觉推理可从数字孪生表示的强化学习中涌现。

原文摘要 · Abstract (English)

Visual reasoning may require models to interpret images and videos and respond to implicit text queries across diverse output formats, from pixel-level segmentation masks to natural language descriptions. Existing approaches rely on supervised fine-tuning with task-specific architectures. For example, reasoning segmentation, grounding, summarization, and visual question answering each demand distinct model designs and training, preventing unified solutions and limiting cross-task and cross-modality generalization. Hence, we propose DT-R1, a reinforcement learning framework that trains large language models to construct digital twin representations of complex multi-modal visual inputs and then reason over these high-level representations as a unified approach to visual reasoning. Specifically, we train DT-R1 using GRPO with a novel reward that validates both structural integrity and output accuracy. Evaluations in six visual reasoning benchmarks, covering two modalities and four task types, demonstrate that DT-R1 consistently achieves improvements over state-of-the-art task-specific models. DT-R1 opens a new direction where visual reasoning emerges from reinforcement learning with digital twin representations.

视觉推理数字孪生强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。