arXiv:2410.08324cs.CLcs.HC2024-10中稿 · DCASE 2024被引 3

分析用户如何用文字搜声音,揭示真实搜索习惯。

The language of sound search: Examining User Queries in Audio Search Engines

  • 对比问卷与网站日志,发现用户在无系统限制时会写更长查询。
  • 900万条真实搜索中,绝大多数是关键词而非完整句子。
  • 关键影响因素包括声源类型、使用场景和声音数量。

本研究考察了音效搜索引擎中的文本型用户查询,涵盖拟音、音效及通用音频检索等应用。现有研究对真实用户需求与行为关注不足。为弥补这一缺口,我们分析了两个来源的查询数据:一项自定义调查和Freesound网站的查询日志。调查旨在收集不受现有系统限制的假设性声音搜索请求,生成一个反映用户意图的数据集,该数据集已公开共享。相比之下,Freesound查询日志包含约900万条搜索记录,全面呈现真实使用模式。研究发现,调查中的查询普遍比Freesound中的更长,表明用户在无系统约束时倾向于提供详细描述。两个数据集均以关键词查询为主,极少用户使用完整句子。影响调查查询的关键因素包括主要声源、预期用途、感知位置及声源数量。这些发现对构建以用户为中心的文本音频检索系统具有重要意义,深化了对声音搜索行为的理解。

原文摘要 · Abstract (English)

This study examines textual, user-written search queries within the context of sound search engines, encompassing various applications such as foley, sound effects, and general audio retrieval. Current research inadequately addresses real-world user needs and behaviours in designing text-based audio retrieval systems. To bridge this gap, we analysed search queries from two sources: a custom survey and Freesound website query logs. The survey was designed to collect queries for an unrestricted, hypothetical sound search engine, resulting in a dataset that captures user intentions without the constraints of existing systems. This dataset is also made available for sharing with the research community. In contrast, the Freesound query logs encompass approximately 9 million search requests, providing a comprehensive view of real-world usage patterns. Our findings indicate that survey queries are generally longer than Freesound queries, suggesting users prefer detailed queries when not limited by system constraints. Both datasets predominantly feature keyword-based queries, with few survey participants using full sentences. Key factors influencing survey queries include the primary sound source, intended usage, perceived location, and the number of sound sources. These insights are crucial for developing user-centred, effective text-based audio retrieval systems, enhancing our understanding of user behaviour in sound search contexts.

声音搜索用户行为文本检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。