用智能体框架让大模型自动标注手语,更准更快。
SignAgent: Agentic LLMs for Linguistically-Grounded Sign Language Annotation and Dataset Curation
- 设计智能体协调多个语言工具,实现手语的自动标注。
- 在伪词标注和词素聚类任务中表现优异,准确率高。
- 适合需要大规模手语数据集的研究者使用。
本文提出 SignAgent,一种基于大语言模型(LLM)的智能体框架,用于可扩展、语言学基础的手语(SL)标注与数据集构建。传统计算方法常仅在词素层面操作,忽略重要语言细节;而人工语言标注成本高、效率低,难以支撑大规模音位感知数据集的创建。SignAgent 通过 SignAgent Orchestrator(推理型 LLM)协调一系列语言工具,并结合 SignGraph(知识驱动型 LLM)提供词汇与语言学基础支持。我们在两个下游任务上评估该框架:一是伪词标注,代理利用多模态证据进行约束性分配,提取并排序合适的词素标签;二是识别词素标注,代理通过推理视觉相似性与音位重叠,检测并精炼视觉聚类,正确识别和分组词素变体。结果表明,该智能体方法在大规模、语言学感知的数据标注与整理中表现强劲。
原文摘要 · Abstract (English)
This paper introduces SignAgent, a novel agentic framework that utilises Large Language Models (LLMs) for scalable, linguistically-grounded Sign Language (SL) annotation and dataset curation. Traditional computational methods for SLs often operate at the gloss level, overlooking crucial linguistic nuances, while manual linguistic annotation remains a significant bottleneck, proving too slow and expensive for the creation of large-scale, phonologically-aware datasets. SignAgent addresses these challenges through SignAgent Orchestrator, a reasoning LLM that coordinates a suite of linguistic tools, and SignGraph, a knowledge-grounded LLM that provides lexical and linguistic grounding. We evaluate our framework on two downstream annotation tasks. First, on Pseudo-gloss Annotation, where the agent performs constrained assignment, using multi-modal evidence to extract and order suitable gloss labels for signed sequences. Second, on ID Glossing, where the agent detects and refines visual clusters by reasoning over both visual similarity and phonological overlap to correctly identify and group lexical sign variants. Our results demonstrate that our agentic approach achieves strong performance for large-scale, linguistically-aware data annotation and curation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。