用神经网络自动识别珠宝并生成多层级描述,帮翻译快速理解
Automatic Identification and Description of Jewelry Through Computer Vision and Neural Networks for Translators and Interpreters
- 分三层生成描述,结合计算机视觉与图像标题技术
- 模型准确率超90%,能识别多样珠宝样式
- 适合翻译、口译员快速获取珠宝专业信息
珠宝识别因款式多样而困难,精准描述通常仅限行业专家。本文提出一种基于神经网络的自动识别与描述方法,帮助翻译和口译人员快速获取准确信息。模型通过计算机视觉与图像标题技术,在三个层次上生成自然语言描述,模拟专家对配饰的分析。采用不同图像标题架构检测图像中的珠宝并生成细节各异的描述。为验证方法有效性,构建了包含多种饰品的图像数据库,并对比多种架构,重点评估编码器-解码器模型。最终模型在生成描述任务中准确率超过90%。
原文摘要 · Abstract (English)
Identifying jewelry pieces presents a significant challenge due to the wide range of styles and designs. Currently, precise descriptions are typically limited to industry experts. However, translators and interpreters often require a comprehensive understanding of these items. In this study, we introduce an innovative approach to automatically identify and describe jewelry using neural networks. This method enables translators and interpreters to quickly access accurate information, aiding in resolving queries and gaining essential knowledge about jewelry. Our model operates at three distinct levels of description, employing computer vision techniques and image captioning to emulate expert analysis of accessories. The key innovation involves generating natural language descriptions of jewelry across three hierarchical levels, capturing nuanced details of each piece. Different image captioning architectures are utilized to detect jewels in images and generate descriptions with varying levels of detail. To demonstrate the effectiveness of our approach in recognizing diverse types of jewelry, we assembled a comprehensive database of accessory images. The evaluation process involved comparing various image captioning architectures, focusing particularly on the encoder decoder model, crucial for generating descriptive captions. After thorough evaluation, our final model achieved a captioning accuracy exceeding 90 per cent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。