医学记录过长导致模型推理困难,新基准揭示解决路径
The Verbose Context Problem in Medical Records
- 用人工生成患者记录构建新数据集,专门测试长文本推理
- 超过40万词元的病历数据让现有方法失效
- 需利用医学领域结构特征提升大规模分析效率
冗长上下文问题出现在结构化医学概念的文本表示过于冗余时。该瓶颈在人群健康研究中尤为突出:对纵向患者记录进行群体级分析需处理数千个医学编码事件,总词元数常超过40万。我们提出PopMedQA,一个通过群体纵向病历上的计算任务来隔离该问题的基准。使用neopatient——一种用于语言控制生成人工病历的新库——构建该基准。通过大量消融实验(包括提示策略、提示压缩和代理分解),发现通用方法无法缓解冗长上下文问题。仍有巨大潜力可通过利用语言模型输入中的领域特定结构实现群体规模推理。
原文摘要 · Abstract (English)
The verbose context problem occurs when structured concepts have token-inefficient textual representations. This bottleneck is acute in population health: cohort-level analysis of longitudinal patient records requires reasoning over thousands of medically-coded events, often exceeding 400K tokens in total. We present PopMedQA, a benchmark isolating this problem through computational tasks on groups of longitudinal patient records. We construct the benchmark using neopatient, a new library for language-controlled generation of artificial patient records. Through extensive ablations -- including prompting strategies, prompt compression, and agentic decomposition -- we find that domain-independent methods fail to alleviate the verbose context problem. There remains significant opportunity to exploit domain-specific structure in language model inputs for population-scale reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。