arXiv:2606.04689quant-phcs.LG2026-06

用量子模型提升长尾关系识别,参数少效果好。

QPredSGG: Hybrid Quantum Predicate Learning for Long-Tailed Scene Graph Generation

论文配图:QPredSGG: Hybrid Quantum Predicate Learning for Long-Tailed Scene Graph Generation
图 1 · 摘自论文原文
  • 用量子电路替代传统分类头,压缩特征并减少参数
  • 4比特量子模型达57.25%的平均召回率,是经典模型的1.4倍
  • 适合关注高效视觉推理与量子机器学习的科研人员

场景图生成(SGG)需对物体及其交互进行关系推理,但严重长尾的谓词分布限制了性能。传统模型依赖数据统计,偏向高频关系而非细粒度语义。现有去偏策略虽提升均值召回率,但分类模块仍需大量参数。本文提出在因果特征增强网络(CFEN)中引入量子谓词头(QP-Head),采用加权交叉熵训练。这是首次在Visual Genome 150上评估混合量子架构用于谓词分类。研究了量子比特数、编码策略、纠缠结构与电路深度的影响。最优4量子比特QP-Head使用振幅编码与强纠缠层,将4096维配对特征压缩至16维,实现256倍降维;在mR@100上达57.25%,较经典CFEN的41.1%显著提升,仅需96个可训练量子参数。8量子比特方案保持良好长尾表现,达55.38%的mR@100,参数为384,深度分析显示表达能力与运行开销存在权衡。结果表明,紧凑的混合量子谓词头可在复杂视觉推理中实现参数高效的长尾关系分类。

原文摘要 · Abstract (English)

Scene Graph Generation (SGG) requires relational reasoning over objects and their interactions, but performance is often limited by severe long-tail predicate imbalance. Classical SGG models frequently rely on dataset statistics, leading to biased predictions toward frequent relations rather than fine-grained semantic predicates. Although existing debiasing strategies improve mean recall, predicate classification in current frameworks still often depends on large classical decision modules with high parameter cost. This work introduces a hybrid quantum predicate classifier for SGG by replacing the classical predicate head in Causal Feature Enhancement Network (CFEN) with a Quantum Predicate Head (QP-Head) trained using weighted cross-entropy. To the best of our knowledge, this is among the first studies to evaluate a hybrid quantum architecture for scene graph predicate classification on Visual Genome 150. We study the effect of qubit count, encoding strategy, entangling structure, and circuit depth on relational prediction. The best 4-qubit QP-Head uses Amplitude Embedding and Strongly Entangling Layers to compress 4096-dimensional pair features into a 16-dimensional quantum-compatible representation, corresponding to a 256$\times$ reduction. It achieves an mR@100 of 57.25%, compared with 41.1% for the classical CFEN reference, while using only 96 trainable quantum parameters. Scaling to 8 qubits maintains strong long-tail performance, reaching an mR@100 of 55.38% with 384 quantum parameters, while the depth analysis shows a trade-off between expressibility and runtime overhead. These results suggest that compact hybrid quantum predicate heads can support parameter-efficient long-tail relational classification in complex visual reasoning tasks.

量子机器学习场景图生成长尾问题参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。