arXiv:2411.12901cs.CLcs.CV2024-11被引 4

Signformer让手语翻译轻量高效,可在边缘设备实时运行。

Signformer is all you need: Towards Edge AI for Sign Language

  • 从零开始设计轻量架构,不用预训练模型或NLP技巧。
  • 参数量比顶尖方法少467到1807倍,仅0.57万参数仍超多数模型。
  • 专为边缘设备优化,适合真实场景的低资源部署。

手语翻译在无词元范式下面临资源消耗过大、难以持续应用的困境。当前主流方法严重依赖大型语言模型(LLMs)、嵌入源或大规模数据集,导致参数和计算开销巨大,难以为现实场景提供可持续支持。尽管性能优异,但这一路径背离了本领域服务听障人群的核心使命。本文提出不依赖预训练模型、知识迁移或任何现成NLP策略的全新架构——Signformer。该模型基于从头设计的轻量级-巨量变换器结构,通过卷积与注意力机制创新,在无需外部辅助的前提下实现性能与效率的极致平衡。我们对手语本质进行分析以指导算法设计,并构建可扩展的Transformer流水线。实验表明,Signformer在基准测试中取得第二名成绩,相比2024年最优模型参数量减少467至1807倍,且在仅0.57万参数的轻量配置下超越绝大多数现有方法。

原文摘要 · Abstract (English)

Sign language translation, especially in gloss-free paradigm, is confronting a dilemma of impracticality and unsustainability due to growing resource-intensive methodologies. Contemporary state-of-the-arts (SOTAs) have significantly hinged on pretrained sophiscated backbones such as Large Language Models (LLMs), embedding sources, or extensive datasets, inducing considerable parametric and computational inefficiency for sustainable use in real-world scenario. Despite their success, following this research direction undermines the overarching mission of this domain to create substantial value to bridge hard-hearing and common populations. Committing to the prevailing trend of LLM and Natural Language Processing (NLP) studies, we pursue a profound essential change in architecture to achieve ground-up improvements without external aid from pretrained models, prior knowledge transfer, or any NLP strategies considered not-from-scratch. Introducing Signformer, a from-scratch Feather-Giant transforming the area towards Edge AI that redefines extremities of performance and efficiency with LLM-competence and edgy-deployable compactness. In this paper, we present nature analysis of sign languages to inform our algorithmic design and deliver a scalable transformer pipeline with convolution and attention novelty. We achieve new 2nd place on leaderboard with a parametric reduction of 467-1807x against the finests as of 2024 and outcompete almost every other methods in a lighter configuration of 0.57 million parameters.

手语识别边缘计算轻量化模型Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。