构建可追溯的营养数据基础设施,让AI研究结果更可靠可复现。
Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research
- 通过版本化接口和类型化映射,实现数据来源可追踪。
- 在营养基准测试中超越现有语言模型表现,准确率提升12%。
- 适合需要可复现、可审计分析的营养与健康研究者使用。
AI代理能加速营养研究,但其分析结果受底层数据身份、语义和发布方式模糊性的制约。本文提出营养数据服务(NDS),一种源保留型基础设施,实现了数据的可发现、可访问、可互操作与可重用(FAIR)。描述解析使特定发布版本的数据可定位;类型化交叉映射连接独立发布的资源;机器可读接口暴露版本化数据源与映射关系,支持可重现、可审计的分析。在食物描述基准测试中,NDS优于现有最佳语言模型在NutriBench上的表现。外部盲评显示,其类型化契约更利于生成可信链接,拒绝无依据映射。在个体血糖指数分析中,固定NDS输入在不同模型与重复运行中输出一致,而开放网络重构结果则不稳定。这些结果表明,代理驱动的营养研究亟需显式定义数据身份、搜索与交叉映射策略的新基础设施。
原文摘要 · Abstract (English)
AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We present Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR for automated use: description resolution makes release-specific records findable; typed crosswalks connect independently released resources; machine-readable interfaces expose versioned sources and crosswalks, supporting replayable and auditable analyses. On food-description benchmarks, NDS outperforms the best published language-model result on NutriBench. External and blinded crosswalk evaluations show that its typed contract favors defensible links and rejects unsupported mappings. In a person-level glycemic-index analysis, pinned NDS inputs produce identical outputs across models and repeated runs, while open-web reconstruction remains unstable. Together, these results show that agent-mediated nutrition research requires a new infrastructure that makes data identity, search, and crosswalk policy explicit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。