为手语处理打造适配Hugging Face的多模态框架
MultimodalHugs: Enabling Sign Language Processing in Hugging Face
- 基于Hugging Face构建,支持手势、姿态等多模态数据
- 实测可处理手语姿态与文本像素数据,提升实验可复现性
- 适合手语研究者及需要灵活数据接口的多模态团队
近年来,手语处理(SLP)在自然语言处理领域日益重要。然而,相较于语音语言研究,手语研究因复杂且非标准的代码实现,导致可复现性差、比较不公平。现有工具如Hugging Face虽支持快速实验,但难以无缝集成手语任务。我们对SLP研究者进行调查后确认此问题。为此,提出MultimodalHugs,一个基于Hugging Face构建的框架,支持多样数据模态与任务,继承Hugging Face生态优势。尽管以手语为重点,其抽象层也适用于其他不匹配标准模板的场景。定量实验表明,该框架能有效处理手语姿态数据或文本像素数据。
原文摘要 · Abstract (English)
In recent years, sign language processing (SLP) has gained importance in the general field of Natural Language Processing. However, compared to research on spoken languages, SLP research is hindered by complex ad-hoc code, inadvertently leading to low reproducibility and unfair comparisons. Existing tools that are built for fast and reproducible experimentation, such as Hugging Face, are not flexible enough to seamlessly integrate sign language experiments. This view is confirmed by a survey we conducted among SLP researchers. To address these challenges, we introduce MultimodalHugs, a framework built on top of Hugging Face that enables more diverse data modalities and tasks, while inheriting the well-known advantages of the Hugging Face ecosystem. Even though sign languages are our primary focus, MultimodalHugs adds a layer of abstraction that makes it more widely applicable to other use cases that do not fit one of the standard templates of Hugging Face. We provide quantitative experiments to illustrate how MultimodalHugs can accommodate diverse modalities such as pose estimation data for sign languages, or pixel data for text characters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。