用量子模型解决视觉词义歧义,降低语义偏差。
Quantum Visual Word Sense Disambiguation: Unraveling Ambiguities Through Quantum Inference Model
- 将多个词义描述编码为量子叠加态,缓解语义偏见。
- 在标准数据集上超越现有方法,尤其善用大模型生成的非专业词义。
- 设计可运行于经典计算机的启发式版本,适合当前硬件条件。
视觉词义消歧关注多义词问题,候选图像常因语义模糊而混淆。传统方法使用经典概率计算图像与每个词义匹配的可能性,并求和得到后验概率。然而,由于语义不确定性,不同来源的词义描述不可避免携带语义偏见,导致消歧结果偏差。受量子叠加态建模不确定性的启发,本文提出无监督视觉词义消歧的量子推理模型(Q-VWSD),将目标词的多个词义描述编码为叠加态,以缓解语义偏见。通过执行量子电路并观测结果实现推理。形式化分析表明,Q-VWSD是经典概率方法的量子推广。在此基础上,进一步设计了可在经典计算机上高效运行的启发式版本。实验显示,该方法优于现有先进经典方法,尤其能有效利用大语言模型生成的非专业化词义描述,进一步提升性能。本研究展示了量子机器学习在实际应用中的潜力,并证明了在量子硬件尚未成熟时,利用量子建模优势在经典设备上实现高效推理的可行性。
原文摘要 · Abstract (English)
Visual word sense disambiguation focuses on polysemous words, where candidate images can be easily confused. Traditional methods use classical probability to calculate the likelihood of an image matching each gloss of the target word, summing these to form a posterior probability. However, due to the challenge of semantic uncertainty, glosses from different sources inevitably carry semantic biases, which can lead to biased disambiguation results. Inspired by quantum superposition in modeling uncertainty, this paper proposes a Quantum Inference Model for Unsupervised Visual Word Sense Disambiguation (Q-VWSD). It encodes multiple glosses of the target word into a superposition state to mitigate semantic biases. Then, the quantum circuit is executed, and the results are observed. By formalizing our method, we find that Q-VWSD is a quantum generalization of the method based on classical probability. Building on this, we further designed a heuristic version of Q-VWSD that can run more efficiently on classical computing. The experiments demonstrate that our method outperforms state-of-the-art classical methods, particularly by effectively leveraging non-specialized glosses from large language models, which further enhances performance. Our approach showcases the potential of quantum machine learning in practical applications and provides a case for leveraging quantum modeling advantages on classical computers while quantum hardware remains immature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。