arXiv:2512.08873cs.CVcs.AI2025-12被引 3

轻量级模型提升低分辨率图像生成描述的准确率

Siamese-Driven Optimization for Low-Resolution Image Latent Embedding in Image Captioning

  • 用双路结构的孪生网络优化图像嵌入表示
  • 在资源受限环境下保持高准确率,计算开销小
  • 适合移动端或边缘设备上的图像描述任务

图像描述在辅助视障人士、提升内容管理系统和增强人机交互方面至关重要。然而,低分辨率图像(LRI)的处理成为当前挑战。虽然使用更大模型如Transformer可提升性能,但这类模型通常体积庞大,需要大量计算资源和内存,导致重训练困难。为此,本文提出SOLI(Siamese-Driven Optimization for Low-Resolution Image Latent Embedding in Image Captioning),一种专为轻量级低分辨率图像描述设计的方法。该方法采用孪生网络架构优化图像潜在嵌入,提升图像到文本转换的效率与准确性。通过双路径神经网络结构,SOLI在不牺牲性能的前提下显著降低计算开销,使其成为资源受限场景下训练的理想选择。

原文摘要 · Abstract (English)

Image captioning is essential in many fields including assisting visually impaired individuals, improving content management systems, and enhancing human-computer interaction. However, a recent challenge in this domain is dealing with low-resolution image (LRI). While performance can be improved by using larger models like transformers for encoding, these models are typically heavyweight, demanding significant computational resources and memory, leading to challenges in retraining. To address this, the proposed SOLI (Siamese-Driven Optimization for Low-Resolution Image Latent Embedding in Image Captioning) approach presents a solution specifically designed for lightweight, low-resolution images captioning. It employs a Siamese network architecture to optimize latent embeddings, enhancing the efficiency and accuracy of the image-to-text translation process. By focusing on a dual-pathway neural network structure, SOLI minimizes computational overhead without sacrificing performance, making it an ideal choice for training on resource-constrained scenarios.

图像描述低分辨率轻量模型孪生网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。