arXiv:2512.12287cs.CV2025-12被引 1

首个含真实目标图的拖拽编辑评测基准,解决模型评估无标准问题。

RealDrag: The First Dragging Benchmark with Real Target Image

  • 构建包含400+标注样本的拖拽编辑数据集,含源图/目标图及操作点
  • 提出4项新指标,量化像素匹配、非编辑区域保留与语义对齐程度
  • 首次系统评测17个顶级模型,揭示当前方法的性能权衡

基于点的图像编辑模型评估因缺乏标准化基准和度量而不可靠。问题根源在于评价协议不一致,尤其缺少包含真实目标图像的数据集,导致不同方法难以客观比较。为此,我们提出首个综合性点拖拽图像编辑基准——RealDrag,包含成对的真实目标图像。数据集涵盖400多个来自多样化视频源的人工标注样本,提供源图/目标图、操作点/目标点、可编辑区域掩码以及图像与编辑动作的描述性标题。我们还提出四项任务专用新指标:语义距离(SeD)、外掩码保持分数(OMPS)、内块保持分数(IPPS)和方向相似性(DiS),分别用于量化像素级匹配精度、检查非编辑区域保持情况、衡量与目标任务的语义一致性。利用该基准,我们首次开展大规模系统性分析,评估了17个当前最优模型。结果揭示了现有方法间的明确权衡,并建立了稳健可复现的基准以指导未来研究。数据集与评估工具包将公开发布。

原文摘要 · Abstract (English)

The evaluation of drag based image editing models is unreliable due to a lack of standardized benchmarks and metrics. This ambiguity stems from inconsistent evaluation protocols and, critically, the absence of datasets containing ground truth target images, making objective comparisons between competing methods difficult. To address this, we introduce \textbf{RealDrag}, the first comprehensive benchmark for point based image editing that includes paired ground truth target images. Our dataset contains over 400 human annotated samples from diverse video sources, providing source/target images, handle/target points, editable region masks, and descriptive captions for both the image and the editing action. We also propose four novel, task specific metrics: Semantical Distance (SeD), Outer Mask Preserving Score (OMPS), Inner Patch Preserving Score (IPPS), and Directional Similarity (DiS). These metrics are designed to quantify pixel level matching fidelity, check preservation of non edited (out of mask) regions, and measure semantic alignment with the desired task. Using this benchmark, we conduct the first large scale systematic analysis of the field, evaluating 17 SOTA models. Our results reveal clear trade offs among current approaches and establish a robust, reproducible baseline to guide future research. Our dataset and evaluation toolkit will be made publicly available.

图像编辑评测基准视觉生成人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。