arXiv:2606.21895cs.CL2026-06

受嗅觉启发的稀疏组合编码提升低资源命名实体识别效果

Olfactory-Inspired Sparse Combinatorial Coding for Low-Resource Named Entity Recognition

论文配图:Olfactory-Inspired Sparse Combinatorial Coding for Low-Resource Named Entity Recognition
图 1 · 摘自论文原文
  • 在嵌入层与序列模型间引入类嗅觉受体-球状体瓶颈结构
  • 1000句数据下多语言平均F1提升,孟加拉语增6.23%
  • 适合数据稀缺场景,尤其对低资源语言有显著优势

低资源语言的命名实体识别受限于监督信号不足和高质量预训练嵌入缺失。生物嗅觉通过受体与球状体组织实现稀疏组合编码,在不确定性下仍能学习鲁棒表征。本文提出一种受嗅觉启发的新架构——受体-球状体瓶颈,置于标准词嵌入与双向LSTM-CRF序列模型之间。我们在六个跨语言数据集上完全从零训练(无预训练嵌入),在不同数据规模条件下评估,包括严格的1000句低资源控制。结果表明,引入表征瓶颈在严重数据稀缺时显著提升F1,主要作为强正则化器。在1000句限制下,至少一种嗅觉配置在所有六组数据中取得最高平均F1。多数语言中性能接近通用瓶颈控制,但在孟加拉语等语言中,该架构表现更优(较标准基线+6.23%,较最佳控制基线+8.47%)。全量数据下泰卢固语也获+4.43%提升,且受体层自然出现稀疏专业化。研究显示,受嗅觉网络启发的结构化稀疏编码可作为有效归纳偏置与正则化器,在有限或噪声监督下学习表征。

原文摘要 · Abstract (English)

Named Entity Recognition (NER) in low-resource languages suffers from limited supervision and a lack of high-quality pretrained embeddings. Biological olfaction, which relies on sparse combinatorial coding through receptor and glomerular organization, offers a compelling paradigm for learning robust representations under uncertainty. In this paper, we introduce a receptor-glomerular bottleneck - a novel, biologically-inspired olfactory architecture - between standard token embeddings and a BiLSTM-CRF sequence model. We evaluate our architecture across six multilingual datasets trained entirely from scratch (without pre-trained embeddings) under varied data-scale conditions, including a strict 1k-sentence low-resource control. Our results demonstrate that introducing a representation bottleneck yields F1 score improvements under severe data scarcity, primarily by acting as a powerful regularizer. Under the 1k capped training condition, at least one olfactory-inspired configuration achieves the highest mean F1 score across all six datasets. While these improvements represent near-ties with generic bottleneck controls for most languages, the olfactory architecture provides a significant advantage in languages like Bangla (+6.23% F1 over standard baseline and +8.47% F1 over the best control baseline) where generic bottlenecks degrade performance. We also observe improvements in the ultra-low-resource Telugu setting (+4.43% F1) at full-scale, and find that sparse specialization naturally emerges within the receptor layer. Our findings suggest that structured sparse coding inspired by olfactory networks serves as an effective inductive bias and regularizer when representations must be learned from limited or noisy supervision.

命名实体识别低资源稀疏编码生物启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。