arXiv:2507.03949cs.CL2025-07

无需标注数据,从事故报告中自动提取人物属性。

A Modular Unsupervised Framework for Attribute Recognition from Unstructured Text

  • 分模块设计,按需调用,轻量高效。
  • 在缺失人员案例中准确提取属性,无需监督训练。
  • 适合跨领域文本属性抽取,尤其适合数据稀缺场景。

我们提出POSID,一种模块化、轻量且按需使用的框架,可在无需任务特定微调的情况下,从非结构化文本中提取基于属性的结构化信息。该方法设计为可跨领域适应,本文在事故报告中的人类属性识别任务上进行评估。POSID结合词法与语义相似性技术,识别相关句子并提取属性。我们在缺失人员案例上使用IinciText数据集验证了其有效性,实现了无需监督训练的属性提取。

原文摘要 · Abstract (English)

We propose POSID, a modular, lightweight and on-demand framework for extracting structured attribute-based properties from unstructured text without task-specific fine-tuning. While the method is designed to be adaptable across domains, in this work, we evaluate it on human attribute recognition in incident reports. POSID combines lexical and semantic similarity techniques to identify relevant sentences and extract attributes. We demonstrate its effectiveness on a missing person use case using the InciText dataset, achieving effective attribute extraction without supervised training.

属性抽取无监督学习文本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。