用情感和语义分析评估图像描述质量,发现人类标注有6%带强烈情绪。
Evaluating authenticity and quality of image captions via sentiment and semantic analyses
- 用预训练模型提取描述的情感分数和语义多样性
- 约6%的人类描述带有强烈情绪,受物体类别影响
- 自动生成描述情绪弱且不依赖物体类别,适合数据质检
深度学习的发展高度依赖大规模标注数据,如自然语言处理与计算机视觉任务。在图像到文本或图像到图像的流程中,模型可能无意中学到人类描述中的情感(情绪)信息,且学习效果受描述多样性的干扰。尽管大规模数据标注主要依赖众包或数据工人,但评估此类训练数据的质量至关重要。本研究提出一种基于情感与语义丰富度的评估方法,应用于包含约15万张图像及分割对象的COCCO-MS数据集。采用预训练模型(Twitter-RoBERTa-base 和 BERT-base)提取描述的情感分数与语义嵌入的变异性。通过多元线性回归分析情感分数与语义变异性在不同物体类别间的关联。结果表明,多数描述为中性,约6%的描述表现出受特定物体类别影响的强烈情绪;同一图像内的描述语义变异性较低,且与物体类别无显著相关性。模型生成的描述中,强烈情绪占比不足1.5%,且不受物体类别影响,也不与对应的人类描述情绪相关。该研究展示了一种基于图像内容评估众包标注质量的方法。
原文摘要 · Abstract (English)
The growth of deep learning (DL) relies heavily on huge amounts of labelled data for tasks such as natural language processing and computer vision. Specifically, in image-to-text or image-to-image pipelines, opinion (sentiment) may be inadvertently learned by a model from human-generated image captions. Additionally, learning may be affected by the variety and diversity of the provided captions. While labelling large datasets has largely relied on crowd-sourcing or data-worker pools, evaluating the quality of such training data is crucial. This study proposes an evaluation method focused on sentiment and semantic richness. That method was applied to the COCO-MS dataset, comprising approximately 150K images with segmented objects and corresponding crowd-sourced captions. We employed pre-trained models (Twitter-RoBERTa-base and BERT-base) to extract sentiment scores and variability of semantic embeddings from captions. The relation of the sentiment score and semantic variability with object categories was examined using multiple linear regression. Results indicate that while most captions were neutral, about 6% of the captions exhibited strong sentiment influenced by specific object categories. Semantic variability of within-image captions remained low and uncorrelated with object categories. Model-generated captions showed less than 1.5% of strong sentiment which was not influenced by object categories and did not correlate with the sentiment of the respective human-generated captions. This research demonstrates an approach to assess the quality of crowd- or worker-sourced captions informed by image content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。