arXiv:2606.17619cs.CV2026-06

用检索增强实现跨主体视角对齐,提升图像生成一致性。

RAVA: Retrieval-Augmented Viewpoint Alignment for Subject-Driven Image Generation

论文配图:RAVA: Retrieval-Augmented Viewpoint Alignment for Subject-Driven Image Generation
图 1 · 摘自论文原文
  • 通过检索与锚点视角一致的参考图像,提供几何线索
  • 采用LogDet策略选出视图一致且结构互补的参考集
  • 适合需要精准视角控制的主体驱动图像生成任务

参考图像驱动的图像生成在身份保持方面进展迅速,但不同主体间的可靠视角控制仍不明确。难点在于:模型需仅凭图像级证据,推断目标主体的隐含视角并将其转移至另一主体,无需相机位姿、深度或光线信息。现有基于多参考图像的方法常依赖虚假语义关联,导致视角漂移、局部结构错位及缺失目标特有内容。本文将此问题建模为跨主体视角对齐,提出RAVA框架,在生成前引入显式几何证据。RAVA首先学习跨实例视角嵌入,检索与锚点视角对齐的目标主体图像;再使用LogDet-based子集选择策略,保留视图一致且结构互补的紧凑参考集;最后由微调后的多参考图像生成器处理。实验表明,通用语义嵌入在此任务中近乎随机,而所提检索器显著提升视角检索质量。在跨主体生成中,RAVA持续优于零样本基线和更强的检索替代方案,相同生成主干下表现更优。结果表明,跨主体视角对齐得益于检索增强的几何定位,而非仅依赖端到端生成。

原文摘要 · Abstract (English)

Reference-driven image generation has made rapid progress on identity preservation, but reliable viewpoint control across different subjects remains poorly understood. The difficulty is not merely generating a new image of the target subject: the model must infer the implicit viewpoint of one subject and transfer it to another subject using only image-level evidence, without camera poses, depth, or ray-based conditions. In this setting, existing generators conditioned on multiple image references often rely on spurious semantic correlations, which lead to viewpoint drift, part-level structural mismatches, and missing or unsupported target-specific content. We formulate this challenge as cross-subject viewpoint alignment and propose RAVA, a retrieval-augmented framework that supplies explicit geometric evidence before generation. RAVA first learns a cross-instance viewpoint embedding that retrieves target-subject images aligned with the anchor viewpoint, then applies a LogDet-based subset selection strategy to retain a compact reference set that is both view-consistent and structurally complementary. The selected references are finally consumed by a fine-tuned multi-reference image generator. Experiments show that generic semantic embeddings are nearly random for this task, while the proposed retriever substantially improves viewpoint retrieval quality. On cross-subject generation, RAVA consistently outperforms zero-shot baselines and stronger retrieval alternatives under the same generation backbone. These results indicate that cross-subject viewpoint alignment benefits from retrieval-augmented geometric grounding rather than relying on end-to-end generation alone.

图像生成视角对齐检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。