用新数据集验证大模型在音乐实体识别中表现优于小模型。
A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection
- 用用户生成元数据构建新数据集,测试大模型上下文学习能力。
- 大模型在上下文学习下准确率高于小模型,最高达87.3%。
- 实体曝光程度显著影响大模型性能,是关键影响因素。
检测歌曲标题、艺术家名称等音乐实体有助于处理音乐搜索查询或分析网络上的音乐消费行为。以往方法使用BERT等小型语言模型(SLMs)取得了良好效果,但研究发现模型预训练阶段的实体暴露程度对其性能影响极大。随着大型语言模型(LLMs)的出现,其在多种下游任务中表现超越SLMs,但在文本实体检测任务中仍存在争议,主要因幻觉问题。本文构建了一个新的用户生成元数据数据集,对近期大模型在上下文学习(ICL)设置下的表现进行基准测试与鲁棒性评估。结果表明,在上下文学习设置下,大模型性能优于小型模型;同时,我们发现实体暴露程度对表现最佳的大模型有显著影响。
原文摘要 · Abstract (English)
Detecting music entities such as song titles or artist names is a useful application to help use cases like processing music search queries or analyzing music consumption on the web. Recent approaches incorporate smaller language models (SLMs) like BERT and achieve high results. However, further research indicates a high influence of entity exposure during pre-training on the performance of the models. With the advent of large language models (LLMs), these outperform SLMs in a variety of downstream tasks. However, researchers are still divided if this is applicable to tasks like entity detection in texts due to issues like hallucination. In this paper, we provide a novel dataset of user-generated metadata and conduct a benchmark and a robustness study using recent LLMs with in-context-learning (ICL). Our results indicate that LLMs in the ICL setting yield higher performance than SLMs. We further uncover the large impact of entity exposure on the best performing LLM in our study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。