无需重训练,用少量例子识别新文字和符号。
Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
- 利用上下文学习,通过极少样本识别未知脚本模式。
- 在合成数据上成功识别中、日、俄等多语言字符。
- 自适应分词器支持无限类别的开放词汇识别,适合新语言研究。
我们提出 Rosetta 模型,采用多模态上下文学习(MICL)方法,在文档中仅需少量示例即可分类全新脚本序列,避免显式重训练。为增强上下文学习能力,设计了具有不同上下文信息量的数据生成流程,提升模型在多种场景下的适应性。关键创新在于使用上下文感知分词器(CAT),实现开放词汇分类,使模型能识别训练时未见过的文本与符号模式,突破训练字母表限制。实验表明,Rosetta 在合成数据集上可成功识别分布外的视觉模式及多种语言脚本,包括中文、希腊文、俄文、法文、西班牙文和日文等。
原文摘要 · Abstract (English)
We introduce Rosetta, a multimodal model that leverages Multimodal In-Context Learning (MICL) to classify sequences of novel script patterns in documents by leveraging minimal examples, thus eliminating the need for explicit retraining. To enhance contextual learning, we designed a dataset generation process that ensures varying degrees of contextual informativeness, improving the model's adaptability in leveraging context across different scenarios. A key strength of our method is the use of a Context-Aware Tokenizer (CAT), which enables open-vocabulary classification. This allows the model to classify text and symbol patterns across an unlimited range of classes, extending its classification capabilities beyond the scope of its training alphabet of patterns. As a result, it unlocks applications such as the recognition of new alphabets and languages. Experiments on synthetic datasets demonstrate the potential of Rosetta to successfully classify Out-Of-Distribution visual patterns and diverse sets of alphabets and scripts, including but not limited to Chinese, Greek, Russian, French, Spanish, and Japanese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。