用真人难以分辨的标准评估生成内容的个性化能力。
Visual Personalization Turing Test
- 以人类无法区分为目标,测试生成内容是否像真人创作
- 新评测框架与10000个角色数据集验证了评分可靠性
- 适合关注隐私安全与个性表达的生成模型研发者
我们提出视觉个性化图灵测试(VPTT),一种基于感知不可区分性的上下文视觉个性化评估范式,而非身份复现。模型若生成内容(图像、视频、3D资产等)在人类或校准过的视觉语言模型(VLM)判断下,与特定人物可能创作或分享的内容无法区分,则视为通过。为实现该测试,我们构建了包含10,000个人物的角色基准(VPTT-Bench)、视觉检索增强生成器(VPRAG)以及与人类和VLM判断高度相关的新评分指标——VPTT Score。实验表明,该评分与人工及VLM评估具有高相关性,验证其作为感知代理的可靠性。结果还显示,VPRAG在原始性与一致性之间达到最佳平衡,为个性化生成模型提供了可扩展且隐私安全的基础。
原文摘要 · Abstract (English)
We introduce the Visual Personalization Turing Test (VPTT), a new paradigm for evaluating contextual visual personalization based on perceptual indistinguishability, rather than identity replication. A model passes the VPTT if its output (image, video, 3D asset, etc.) is indistinguishable to a human or calibrated VLM judge from content a given person might plausibly create or share. To operationalize VPTT, we present the VPTT Framework, integrating a 10k-persona benchmark (VPTT-Bench), a visual retrieval-augmented generator (VPRAG), and the VPTT Score, a text-only metric calibrated against human and VLM judgments. We show high correlation across human, VLM, and VPTT evaluations, validating the VPTT Score as a reliable perceptual proxy. Experiments demonstrate that VPRAG achieves the best alignment-originality balance, offering a scalable and privacy-safe foundation for personalized generative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。