arXiv:2601.04201cs.CLcs.AI2026-01中稿 · -papers被引 2

用社区故事构建本地AI知识库,解决大模型对地方问题答不准的问题。

Collective Narrative Grounding: Community-Coordinated Data Contributions to Improve Local AI Systems

  • 将社区叙事转为可提取实体时间地点的结构化数据,由社区自己管理
  • 审计1.4万条本地问答发现76.7%错误源于地理文化时间错配
  • 适合关注公平性、本地化AI与参与式设计的研究者和实践者

大型语言模型在回答社区特定问题时常出现失败,形成‘知识盲区’,边缘化地方声音并加剧认知不公。我们提出集体叙事定位(Collective Narrative Grounding)——一种参与式协议,将社区故事转化为结构化叙事单元,并在社区治理下融入AI系统。通过三次参与式地图工作坊(共24名成员),我们设计了能保留叙事丰富性的同时支持实体、时间、地点提取、验证与溯源控制的方法与模板。为界定问题,我们审计了一个县级基准数据集,包含14,782条本地信息问答对,发现事实缺失、文化误解、地理混淆和时间错位占错误总量的76.7%。在基于工作坊生成的参与式问答集上,最先进的LLM在无额外上下文时正确率低于21%,凸显本地化语境的重要性。缺失事实多出现在收集的叙事中,表明可通过补充叙事直接缓解主要错误类型。除协议与试点外,我们还提出代表性与权力、治理与控制、隐私与同意等关键设计张力,给出以检索优先、溯源可见、本地治理为核心的问答系统具体要求。整体而言,我们的分类体系、协议与参与式评估为构建真正扎根社区的AI提供了严谨基础。

原文摘要 · Abstract (English)

Large language model (LLM) question-answering systems often fail on community-specific queries, creating "knowledge blind spots" that marginalize local voices and reinforce epistemic injustice. We present Collective Narrative Grounding, a participatory protocol that transforms community stories into structured narrative units and integrates them into AI systems under community governance. Learning from three participatory mapping workshops with N=24 community members, we designed elicitation methods and a schema that retain narrative richness while enabling entity, time, and place extraction, validation, and provenance control. To scope the problem, we audit a county-level benchmark of 14,782 local information QA pairs, where factual gaps, cultural misunderstandings, geographic confusions, and temporal misalignments account for 76.7% of errors. On a participatory QA set derived from our workshops, a state-of-the-art LLM answered fewer than 21% of questions correctly without added context, underscoring the need for local grounding. The missing facts often appear in the collected narratives, suggesting a direct path to closing the dominant error modes for narrative items. Beyond the protocol and pilot, we articulate key design tensions, such as representation and power, governance and control, and privacy and consent, providing concrete requirements for retrieval-first, provenance-visible, locally governed QA systems. Together, our taxonomy, protocol, and participatory evaluation offer a rigorous foundation for building community-grounded AI that better answers local questions.

社区智能本地化AI参与式设计叙事数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。