arXiv:2603.11679cs.AI2026-03被引 2

用大模型自动生成数据提取规范,提升医疗数据学习效率

LLMs can construct powerful representations and streamline sample-efficient supervised learning

  • 让大模型分析少量文本样例,生成可编程的提取规则
  • 在15个临床任务中超越传统特征模型和普通大模型
  • 规则可审计、易扩展,适合医疗等高可靠性场景

随着真实世界数据日益复杂多样,监督学习常受限于输入表示设计。处理时间序列、自由文本和结构化记录等多模态数据通常需要大量领域知识。本文提出一种代理式流水线:首先,大模型基于少量多样化文本序列输入,在上下文中归纳出全局规约(rubric),作为提取与组织证据的程序化规范;该规约随后将原始文本序列转化为下游模型更易处理的标准化格式。此外,还引入任务相关的局部规约,由大模型生成解释性摘要。在EHRSHOT基准的15个临床任务中,本方法显著优于计数特征模型、原始大模型基线及预训练数据量大得多的临床基础模型。除性能提升外,规约还具备可审计、规模化成本低、便于生成表格表示等优势。

原文摘要 · Abstract (English)

As real-world datasets become more complex and heterogeneous, supervised learning is often bottlenecked by input representation design. Modeling multimodal data, such as time-series, free text, and structured records, often requires non-trivial domain expertise. We propose an agentic pipeline to streamline this process. First, an LLM analyzes a small but diverse subset of text-serialized input examples in-context to synthesize a global rubric, which acts as a programmatic specification for extracting and organizing evidence. This rubric is then used to transform naive text-serializations of inputs into a more standardized format for downstream models. We also describe local rubrics, which are task-conditioned interpretive summaries generated by an LLM. Across 15 clinical tasks from the EHRSHOT benchmark, our rubric approaches significantly outperform count-feature models, naive LLM baselines, and a clinical foundation model pretrained on orders of magnitude more data. Beyond performance, rubrics offer operational advantages such as being easy to audit, cost-effectiveness at scale, and facilitating tabular representations.

大模型医疗AI数据规约高效学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。