构建汽车参数抽取的细粒度命名实体数据集,助力智能车市信息挖掘。
AutoSpecNER: A Fine-Grained Named Entity Recognition Dataset for Vehicle Specification Extraction

- 专家标注659篇车况广告,涵盖15类实体,共超1万条标注。
- DeBERTa模型表现最优,微平均F1达90%,远超规则基线(43%)。
- 适合从事汽车知识图谱、智能售车系统研究者使用。
车辆广告包含丰富的规格信息,但汽车行业命名实体识别资源仍有限。本文提出AutoSpecNER,一个面向车辆列表的细粒度命名实体识别专家标注数据集。该数据集收录来自知名二手车网站的659篇广告,涵盖15个类别(如MODEL、ENGINE_SPEC、BATTERY_CAPACITY),共标注超过10,000个实体。通过标注者间一致性验证,平均评分达到91.5%。我们对基于规则的抽取方法、微调的Transformer编码器及大语言模型进行了基准测试。DeBERTa在该任务上表现最佳,微平均F1得分为90%,显著优于规则基线(43%)和最强的大语言模型(77.8%)。
原文摘要 · Abstract (English)
Vehicle advertisements contain rich specification information, but automotive NER resources remain limited. We introduce AutoSpecNER, an expert-annotated dataset for fine-grained entity recognition in vehicle listings. The dataset includes 659 advertisements from a popular car-selling website, with over 10,000 entities annotated across 15 categories, including MODEL, ENGINE_SPEC, and BATTERY_CAPACITY. Annotation quality was validated through inter-annotator agreement, achieving an average score of 91.5%. We benchmark rule-based extraction, fine-tuned transformer encoders, and large language models. DeBERTa achieves the best performance with a 90% micro-F1 score, outperforming the rule-based baseline (43%) and the strongest large language model (77.8%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。