arXiv:2603.29697cs.CV2026-03被引 1

提出细粒度人脸表情编辑评估基准,解决评测偏差与指令遵循难题。

FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing

  • 构建747组高质量三元组数据,包含原图、指令和真实目标图。
  • 设计跨粒度评分体系,精准衡量指令遵循、保真度与表情变化幅度。
  • 揭示当前模型在保真度与精确控表间的瓶颈,适合表情生成研究者使用。

人脸表情图像编辑需精细控制,在严格保持身份与背景的前提下精准调整表情。然而现有编辑基准多聚焦通用场景,缺乏高质量人脸图像及对应编辑指令,且评估指标存在系统性偏差,常偏好简单或过拟合的编辑方式。为此,我们提出FED-Bench,一个全面的基准测试平台,具备严谨的测试流程与精准的评估体系。首先,通过级联可扩展的流水线构建了包含747个三元组的数据集,每个三元组由原始图像、编辑指令和真实目标图像组成,支持精准评估。其次,提出FED-Score,一种跨粒度评估协议,将评估拆分为三个维度:对齐性(验证指令遵循)、保真度(测试图像质量与身份保留)和相对表情增益(量化表情变化幅度),有效缓解上述评估偏差。第三,对18种图像编辑模型进行基准测试,发现当前方法难以同时实现高保真度与精确表情操控,细粒度指令遵循是主要瓶颈。最后,利用所提基准引擎的可扩展特性,构建了一个20,000+的野外人脸训练数据集,并通过微调基线模型验证其有效性。相关代码与基准将公开发布。

原文摘要 · Abstract (English)

Facial expression image editing requires fine-grained control to strictly preserve human identity and background while precisely manipulating expression. However, existing editing benchmarks primarily focus on general scenarios, lacking high-quality facial images and corresponding editing instructions. Furthermore, current evaluation metrics exhibit systemic biases in this task, often favoring lazy editing or overfit editing. To bridge these gaps, we propose FED-Bench, a comprehensive benchmark featuring rigorous testing and an accurate evaluation suite. First, we carefully construct a benchmark of 747 triplets through a cascaded and scalable pipeline, each comprising an original image, an editing instruction, and a ground-truth image for precise evaluation. Second, we introduce FED-Score, a cross-granularity evaluation protocol that disentangles assessment into three dimensions: Alignment for verifying instruction following, Fidelity for testing image quality and identity preservation, and Relative Expression Gain for quantifying the magnitude of expression changes, effectively mitigating the aforementioned evaluation biases. Third, we benchmark 18 image editing models, revealing that current approaches struggle to simultaneously achieve high fidelity and accurate expression manipulation, with fine-grained instruction following identified as the primary bottleneck. Finally, leveraging the scalable characteristic of introduced benchmark engine, we provide a 20k+ in-the-wild facial training set and demonstrate its effectiveness by fine-tuning a baseline model that achieves significant performance gains. Our benchmark and related code will be made publicly open soon.

表情编辑评估基准细粒度控制图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。