为汽车广告文本设计新命名实体识别方案,提升行业数据分析能力
Shifting NER into High Gear: The Auto-AdvER Approach
- 构建三类标签的命名实体识别框架:车况、历史记录、销售选项
- 标注一致率达92% F1分,大模型表现优于传统编码器模型
- 适合汽车领域数据挖掘与市场分析,尤其适用于企业级广告洞察
本文提出Auto-AdvER,一个面向汽车广告文本的专用命名实体识别(NER)体系与数据集。该体系包含三个标签:'Condition'(车况)、'Historic'(历史记录)和'Sales Options'(销售选项),旨在满足产业界对文本挖掘的需求。我们阐述了标注指导原则,描述了标注体系开发方法,并通过标注研究展示了92%的F1得分,表明高一致性。对比了编码器类模型(BERT、DeBERTaV3)与解码器类大语言模型(LLMs)如Llama、Qwen、GPT-4和Gemini的表现,结果显示大语言模型整体性能更优,但成本高且存在局限。本工作为更细粒度的广告分析奠定基础,可应用于市场动态分析与数据驱动的预测性维护。该标注体系及发现对私有或公共机构在汽车领域及其他专业领域的命名实体识别具有参考价值。
原文摘要 · Abstract (English)
This paper presents a case study on the development of Auto-AdvER, a specialised named entity recognition schema and dataset for text in the car advertisement genre. Developed with industry needs in mind, Auto-AdvER is designed to enhance text mining analytics in this domain and contributes a linguistically unique NER dataset. We present a schema consisting of three labels: "Condition", "Historic" and "Sales Options". We outline the guiding principles for annotation, describe the methodology for schema development, and show the results of an annotation study demonstrating inter-annotator agreement of 92% F1-Score. Furthermore, we compare the performance by using encoder-only models: BERT, DeBERTaV3 and decoder-only open and closed source Large Language Models (LLMs): Llama, Qwen, GPT-4 and Gemini. Our results show that the class of LLMs outperforms the smaller encoder-only models. However, the LLMs are costly and far from perfect for this task. We present this work as a stepping stone toward more fine-grained analysis and discuss Auto-AdvER's potential impact on advertisement analytics and customer insights, including applications such as the analysis of market dynamics and data-driven predictive maintenance. Our schema, as well as our associated findings, are suitable for both private and public entities considering named entity recognition in the automotive domain, or other specialist domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。