arXiv:2410.15532cs.SDeess.AS2024-10

构建首个含人类听觉印象的环境音视频描述数据集。

Construction and Analysis of Impression Caption Dataset for Environmental Sounds

  • 用ChatGPT生成+人工筛选,构建3600条环境音印象描述。
  • 主观与客观评估验证了描述的恰当性。
  • 适合研究声音情感、人机交互与音频生成的学者。

现有环境音与文本转换数据集多关注声音内容和出现顺序,但极少包含人类听觉感受,如“尖锐”“惊艳”等印象描述。本研究构建了一个包含环境音印象描述的数据集,利用ChatGPT生成印象描述,并由人工筛选出最合适的条目。最终数据集包含3,600条印象描述。通过主观与客观评估验证其合理性,结果表明可有效生成符合人类感知的环境音印象描述。

原文摘要 · Abstract (English)

Some datasets with the described content and order of occurrence of sounds have been released for conversion between environmental sound and text. However, there are very few texts that include information on the impressions humans feel, such as "sharp" and "gorgeous," when they hear environmental sounds. In this study, we constructed a dataset with impression captions for environmental sounds that describe the impressions humans have when hearing these sounds. We used ChatGPT to generate impression captions and selected the most appropriate captions for sound by humans. Our dataset consists of 3,600 impression captions for environmental sounds. To evaluate the appropriateness of impression captions for environmental sounds, we conducted subjective and objective evaluations. From our evaluation results, we indicate that appropriate impression captions for environmental sounds can be generated.

环境音印象描述数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。