arXiv:2601.04614cs.CV2026-01

用双曲几何提升图文对齐评估,自适应调整评分更准

HyperAlign: Hyperbolic Entailment Cones for Adaptive Text-to-Image Alignment Assessment

  • 将图文特征映射到双曲空间,建模语义层次结构
  • 动态监督机制将逻辑蕴含转为连续几何约束
  • 自适应调制回归器提升跨数据集泛化能力

随着文本到图像生成技术的快速发展,准确评估生成图像与文本提示之间的对齐程度已成为关键挑战。现有方法依赖欧几里得空间度量,忽视了语义对齐的结构特性,且缺乏对不同样本的自适应能力。为此,我们提出HyperAlign,一种基于双曲蕴含几何的自适应图文对齐评估框架。首先,利用CLIP提取欧氏特征并映射至双曲空间;其次,设计动态监督蕴含建模机制,将离散蕴含逻辑转化为连续几何结构监督;最后,提出自适应调制回归器,利用双曲几何特征生成样本级调制参数,自适应校准欧氏余弦相似度以预测最终得分。HyperAlign在单数据库评估与跨数据库泛化任务中均取得极具竞争力的性能,充分验证了双曲几何建模在图文对齐评估中的有效性。

原文摘要 · Abstract (English)

With the rapid development of text-to-image generation technology, accurately assessing the alignment between generated images and text prompts has become a critical challenge. Existing methods rely on Euclidean space metrics, neglecting the structured nature of semantic alignment, while lacking adaptive capabilities for different samples. To address these limitations, we propose HyperAlign, an adaptive text-to-image alignment assessment framework based on hyperbolic entailment geometry. First, we extract Euclidean features using CLIP and map them to hyperbolic space. Second, we design a dynamic-supervision entailment modeling mechanism that transforms discrete entailment logic into continuous geometric structure supervision. Finally, we propose an adaptive modulation regressor that utilizes hyperbolic geometric features to generate sample-level modulation parameters, adaptively calibrating Euclidean cosine similarity to predict the final score. HyperAlign achieves highly competitive performance on both single database evaluation and cross-database generalization tasks, fully validating the effectiveness of hyperbolic geometric modeling for image-text alignment assessment.

图文对齐双曲几何自适应评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。