用轻量模型自动提取意大利语垃圾管理术语,效果稳定可解释。
Peacemaker at ATE-IT: Automatic term extraction from Italian text for waste management data using encoder model
- 基于微调编码器的低资源术语抽取方法
- 在类型级和微观级指标上表现均衡,F1分数领先
- 适合资源有限但需可解释性的工业文本处理场景
自动术语抽取在现代技术中日益重要,广泛应用于各类搜索引擎。尽管近期进展显著,但受限于标注文档数量少、多词表达复杂及领域偏移等因素,精准标注仍具挑战。本文针对ATE共享任务中的任务A,提出一种低成本且可解释的自动术语抽取方法。该方法采用微调策略,可在少量计算资源下运行。通过类型级与微观级的精确率、召回率和F1分数评估系统性能。实验结果表明,所提方法在各项指标上表现一致且均衡,优于其他参赛团队。尽管技术本身较为简单,但为低资源模型提供了良好起点。研究结果表明,未来在模型扩展后仍可保持高表现力与可解释性。
原文摘要 · Abstract (English)
The development of automatic term extraction has become increasingly important in modern technology. Automatic term extraction can be found in virtually every search engine that is currently available to users. Recent advancements have provided promising results for the extraction of automatic terms; however, accurate labeling is difficult because of several factors, such as the limited number of annotated documents available for training and the complexity of extracting multi-word expressions due to shifts in the domain. In this paper, we will present a low-cost and interpretable method of automatic term extraction, developed specifically for Task A of the ATE Shared Task. This new method utilizes fine-tuning extraction strategies that can run on a small amount of computational resources. We evaluated our automated system using both type-level and micro-level measures of precision, recall, and F1-score to measure both complementary aspects of the extraction performance. According to the experimental results, our proposed approach achieves consistent and balanced performance compared to other teams. Even though the technique itself is relatively straightforward, it serves as a good starting point for low-resource models. Overall, the findings point toward the possibility of significant future advancements (in model expansion) with higher-level performance still able to retain their ability to be interpreted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。