轻量级注意力机制提升多语言论文主题推荐效果
NBF at SemEval-2025 Task 5: Light-Burst Attention Enhanced System for Multilingual Subject Recommendation
- 用小维度自注意力编码句子嵌入,降低计算开销
- 跨语言任务下平均召回率达32.24%,部分场景超43%
- 适合资源受限场景的多语言主题推荐应用
我们提交的系统参与了SemEval 2025 Task 5,聚焦英文与德文学术领域的跨语言主题分类。方法在训练中利用双语数据,结合负采样与基于间隔的检索目标。通过设计内部维度显著降低的维度-令牌自注意力机制,有效编码句子嵌入以支持主题检索。定量评估显示,在通用量化设置下(涵盖所有主题),系统平均召回率为32.24%;在通用定性评估中,分别达到43.16%和31.53%的召回率,且仅需极少GPU资源,表现具有竞争力。结果表明该方法在资源受限条件下仍能有效捕捉相关主题信息,但仍有优化空间。
原文摘要 · Abstract (English)
We present our system submission for SemEval 2025 Task 5, which focuses on cross-lingual subject classification in the English and German academic domains. Our approach leverages bilingual data during training, employing negative sampling and a margin-based retrieval objective. We demonstrate that a dimension-as-token self-attention mechanism designed with significantly reduced internal dimensions can effectively encode sentence embeddings for subject retrieval. In quantitative evaluation, our system achieved an average recall rate of 32.24% in the general quantitative setting (all subjects), 43.16% and 31.53% of the general qualitative evaluation methods with minimal GPU usage, highlighting their competitive performance. Our results demonstrate that our approach is effective in capturing relevant subject information under resource constraints, although there is still room for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。