分析Claude AI健康问答的引用来源可信度,发现97.8%引用来自权威机构。
Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework
- 用权威信号框架分析3075个健康问题的10038条引用来源
- 97.8%引用来自医疗、政府和专业组织,商业信息仅占2.2%
- 顶尖机构如梅奥诊所贡献超四分之一引用,适合医疗AI可信度研究者参考
本研究旨在评估Anthropic公司开发的Claude AI在回答消费者健康问题时引用来源所呈现的权威信号。尽管关于大模型生成健康引文质量的讨论众多,但对其引用来源本身的可靠性及是否符合健康专业人士标准的研究仍有限。本研究采用基于Google Research的HealthSearchQA数据集(含3,172个消费者健康问题),经筛选后得到3,075个问题对应的10,038条引用进行分析。应用权威信号框架(Jacques et al., 2026)对4个领域中的10项权威信号进行评估,选取542个代表性来源作分层抽样。结果显示,机构类来源占所有引用的97.8%(n=9,818),其中医疗机构占比最高(36.5%),其次是政府资源(31.6%)和专业协会(28.4%);商业健康信息占比2.2%(n=220)。前十大机构贡献了57.8%的引用,梅奥诊所单独占24.7%。在商业来源中,86.4%标注医学审核声明,82.5%使用结构化数据标记,71.8%内容完整;而传统机构来源即使无此类标记也被引用。鉴于Anthropic将Claude定位为符合HIPAA要求的医疗应用,本研究为评估其引文行为提供了基准,并验证了该框架在跨平台评估AI健康信息中的实用性。
原文摘要 · Abstract (English)
This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questions. While there exists a great deal of discourse around the quality of health citations that LLMs produce, there is limited information on the integrity of the sources the citations originate from, and to what extent the sources are, from what health professionals would consider, credible sources. This descriptive cross-sectional study used data from HealthSearchQA, which contains 3,172 consumer health questions curated by Google Research. After exclusions, a final dataset of 3,075 questions yielding 10,038 citations was analyzed. The Authority Signals Framework (Jacques et al., 2026) was applied to examine 10 authority signals across four domains for a disproportionate stratified sample of 542 sources. Established institutional sources accounted for 97.8% of all citations (n = 9,818). Medical Institutions were the most frequently cited organization type (36.5%), followed by Government Resources (31.6%) and Professional Associations (28.4%). Commercial Health Information comprised 2.2% (n = 220). The top 10 organizations accounted for 57.8% of all citations, with Mayo Clinic alone representing 24.7%. Among commercial sources in the focused sample, 86.4% displayed medical review statements, 82.5% used schema markup, and 71.8% had comprehensive content, while traditional institutional sources appeared in Claude's citations with or without these same markers. As Anthropic positions Claude for HIPAA-ready healthcare applications, these findings establish a baseline for Claude's citation behavior and demonstrate the utility of the Authority Signals Framework as a tool for ongoing, cross-platform evaluation of AI-mediated health information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。