arXiv:2607.11891cs.CLcs.AI2026-07

专精领域问答新基准,测试大模型的上下文对齐能力。

CANDI: Contextual Alignment for Niche Domains Question Answering

论文配图:CANDI: Contextual Alignment for Niche Domains Question Answering
图 1 · 摘自论文原文
  • 分两类问题:直接查证与多跳推理,评估上下文敏感性。
  • 现有大模型在专业场景下表现差,需增强上下文与符号推理。
  • 适合研究高风险领域可信AI的团队使用。

将大语言模型(LLMs)部署于医疗诊断、金融咨询等专精领域时,需超越通用知识的评估。传统问答基准难以捕捉此类场景所需的语境依存、用户意识与领域理解。为此,我们提出CANDI-QA(Contextual Alignment for Niche Domains Question Answering),一个新型数据集,用于评估模型在专精领域中生成准确、上下文敏感且与用户对齐的答案能力。该数据集包含专家标注的问答对,分为两类:(1) 信息协助类问题,为直接事实查询,要求精确提取;(2) 应用推理类问题,为多跳推理任务,需情境推断以生成可操作洞察。我们评估了十余种不同规模的语言模型,涵盖开源小型模型到最新闭源系统。作为强基线,我们提出MTSS-Net,一种轻量级神经符号框架,结合神经检索与规则推理。结果表明,在专精领域实现上下文对齐面临巨大挑战,当前模型缺乏上下文或符号集成即显不足。最终,CANDI-QA成为推动上下文感知语言模型研究的关键基准,促进高风险领域可信AI的发展。

原文摘要 · Abstract (English)

The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge. Traditional question-answering benchmarks often fail to capture the nuanced contextual grounding, user awareness, and domain understanding these fields require. To address this, we introduce CANDI-QA (Contextual Alignment for Niche Domains Question Answering), a novel dataset evaluating LLMs on delivering accurate, context-sensitive, and user-aligned answers in specialized settings. CANDI-QA features expert-curated question-answer pairs structured into two categories: (1) Information Assistance Questions, which are direct, factual queries requiring precise extraction, and (2) Applied Inference Questions, which are multi-hop reasoning tasks needing situational inference to generate actionable insights. We evaluate over ten diverse language models, from compact open-source to state-of-the-art proprietary systems. As a robust baseline, we present MTSS-Net, a lightweight neuro-symbolic framework combining neural retrieval with rule-based reasoning. Our findings highlight the profound challenges of achieving contextual alignment in niche domains, revealing the limitations of current LLMs without enhanced contextual or symbolic integration. Ultimately, CANDI-QA serves as a critical benchmark for advancing research in context-aware language models, stimulating the development of robust, trustworthy AI for high-stakes domains.

专精领域问答系统上下文对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。