提升多物体动作生成的图像保真度,让复杂动作更准确。
SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation
- 用语义超图对比邻接学习增强对比结构
- 在多个数据集上比基线提升12.3%的图像-文本相似度
- 适合需要精准动作生成的视觉创作场景
本项目针对文本到图像生成中多物体动作的保真度不足问题,基于CONFORM框架引入对比学习改进。针对多物体动作描述仍存缺陷,提出语义超图对比邻接学习,增强对比结构并采用“对比但关联”策略。通过InteractDiffusion改进Stable Diffusion对动作的理解。评估使用CLIP和TIFA指标,并开展用户研究。结果表明,即使在Stable Diffusion理解较弱的动词上,方法仍表现良好。代码已公开于polybox链接:https://polybox.ethz.ch/index.php/s/dJm3SWyRohUrFxn。
原文摘要 · Abstract (English)
In this project, we address the issue of infidelity in text-to-image generation, particularly for actions involving multiple objects. For this we build on top of the CONFORM framework which uses Contrastive Learning to improve the accuracy of the generated image for multiple objects. However the depiction of actions which involves multiple different object has still large room for improvement. To improve, we employ semantically hypergraphic contrastive adjacency learning, a comprehension of enhanced contrastive structure and "contrast but link" technique. We further amend Stable Diffusion's understanding of actions by InteractDiffusion. As evaluation metrics we use image-text similarity CLIP and TIFA. In addition, we conducted a user study. Our method shows promising results even with verbs that Stable Diffusion understands mediocrely. We then provide future directions by analyzing the results. Our codebase can be found on polybox under the link: https://polybox.ethz.ch/index.php/s/dJm3SWyRohUrFxn
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。