用物理约束提升多模态图像评估准确性
Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
- 结合大模型推理与视觉语言模型,分三阶段评估图像结构与语义
- 通过物理规则验证物体位置、对齐与一致性,避免生成错误场景
- 适合需要高精度图像评估的科研或工业场景
当前主流评估指标如BLEU、CIDEr、VQA分数、SigLIP-2和CLIPScore难以捕捉特定领域或上下文相关的语义与结构准确性。为此,本文提出物理约束多模态数据评估(PCMDE)指标,融合大语言模型推理、基于知识的映射与视觉语言模型,克服上述局限。该架构包含三个主要阶段:(1) 通过目标检测与视觉语言模型提取空间与语义特征;(2) 采用置信度加权组件融合,实现自适应的组件级验证;(3) 利用大语言模型进行物理引导推理,强制执行结构与关系约束(如对齐、位置、一致性)。该方法能有效识别不符合物理规律的合成图像。
原文摘要 · Abstract (English)
Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accuracy, especially for domain-specific or context-dependent scenarios. For this, this paper proposes a Physics-Constrained Multimodal Data Evaluation (PCMDE) metric combining large language models with reasoning, knowledge based mapping and vision-language models to overcome these limitations. The architecture is comprised of three main stages: (1) feature extraction of spatial and semantic information with multimodal features through object detection and VLMs; (2) Confidence-Weighted Component Fusion for adaptive component-level validation; and (3) physics-guided reasoning using large language models for structural and relational constraints (e.g., alignment, position, consistency) enforcement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。