用智能体评估伊斯兰医学问答,提升准确与安全。
From RAG to Agentic: Validating Islamic-Medicine Responses with LLM Agents
- 设计统一评测流程,融合检索与自我批判提示。
- 检索使事实准确率提升13%,智能体再增10%。
- 适合关注文化敏感医疗AI的研究者与开发者。
千年历史的伊斯兰医学典籍如《医典》和先知疗法蕴含丰富的预防医学、营养学与整体疗法,但难以普及且未被现代AI系统充分应用。现有语言模型评测侧重事实记忆或用户偏好,缺乏对文化根基医疗建议的大规模验证。本文提出统一评测框架Tibbe-AG,将30个精心筛选的先知医学问题与人工验证的疗法对应,并对比三种LLM(LLaMA-3、Mistral-7B、Qwen2-7B)在直接生成、检索增强生成及科学自检过滤三种配置下的表现。每个回答由二级LLM作为智能体裁判评估,生成单一3C3H质量评分。检索使事实准确性提升13%,智能体提示进一步带来10%的改进,体现在更深入的机制理解与安全考量。结果表明,结合古典伊斯兰文本、检索与自我评估,可实现可靠且文化敏感的医疗问答。
原文摘要 · Abstract (English)
Centuries-old Islamic medical texts like Avicenna's Canon of Medicine and the Prophetic Tibb-e-Nabawi encode a wealth of preventive care, nutrition, and holistic therapies, yet remain inaccessible to many and underutilized in modern AI systems. Existing language-model benchmarks focus narrowly on factual recall or user preference, leaving a gap in validating culturally grounded medical guidance at scale. We propose a unified evaluation pipeline, Tibbe-AG, that aligns 30 carefully curated Prophetic-medicine questions with human-verified remedies and compares three LLMs (LLaMA-3, Mistral-7B, Qwen2-7B) under three configurations: direct generation, retrieval-augmented generation, and a scientific self-critique filter. Each answer is then assessed by a secondary LLM serving as an agentic judge, yielding a single 3C3H quality score. Retrieval improves factual accuracy by 13%, while the agentic prompt adds another 10% improvement through deeper mechanistic insight and safety considerations. Our results demonstrate that blending classical Islamic texts with retrieval and self-evaluation enables reliable, culturally sensitive medical question-answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。