arXiv:2606.20255cs.CLcs.AI2026-06

提出九维框架,让AI读懂尼日利亚话语背后的真正意图。

The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse

  • 构建九维标注体系,区分表层情绪与真实沟通意图。
  • 零样本下语域识别准确率仅33.3%,引入框架后升至73.3%。
  • 大模型能力与文化理解可分离,小模型反而更受框架帮助。

我们提出意义智能框架(MIF),一种用于尼日利亚公共话语的九维标注与评估体系,旨在区分表面情绪与真实交际意图。现有尼日利亚语言基准如 NaijaSenti 和 AfriSenti 将情感分类视为三极极性任务。我们认为,当前AI在尼日利亚话语中的主要失败模式并非翻译错误,而是语境缺失:同一语句在不同说话人、受众和情境下可能具有相反的语用效果。MIF通过九个可评分维度实现这一洞察:语域、表面情绪、真实意图、反讽、隐含信息、风险等级、标注者信心、说话人情绪及建议沟通行动。我们构建了一个包含30个样本的校准数据集,涵盖标准英语、尼日利亚英语、尼日利亚皮钦语及混合语码形式。评估了三个前沿语言模型(Gemini 2.5 Flash、GPT-5 和 Gemini 2.5 Pro)在零样本与框架引导提示下的表现。两大核心发现:第一,语域差距——零样本语域分类准确率为33.3%,引入MIF框架后提升至73.3%(+40个百分点);第二,模型能力与文化理解解耦:GPT-5(MIS 67.8)与 Gemini 2.5 Pro(MIS 65.4)得分低于 Flash(MIS 78.6),且均未从框架提示中获益。我们已公开框架规范、标注指南与校准集以支持可复现性。

原文摘要 · Abstract (English)

We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent. Existing benchmarks for Nigerian languages, including NaijaSenti and AfriSenti, treat sentiment classification as a three-way polarity task. We argue that the dominant failure mode of AI systems on Nigerian discourse is not translation failure but context failure: the same utterance carries opposite pragmatic force depending on speaker, audience, and situation. The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent, irony, coded subtext, risk tier, annotator confidence, speaker emotion, and recommended communications action. We construct a 30-item calibration dataset spanning Standard English, Nigerian English, Nigerian Pidgin, and code-mixed registers, and evaluate three frontier language models (Gemini 2.5 Flash, GPT-5, and Gemini 2.5 Pro) under zero-shot and schema-informed prompting conditions. Two headline findings emerge. First, the Register Gap: zero-shot register classification accuracy is 33.3%, rising to 73.3% (+40 points) when the model receives the MIF schema in-context. Second, model capability and cultural competence are decoupled: GPT-5 (MIS 67.8) and Gemini 2.5 Pro (MIS 65.4) score lower than Flash (MIS 78.6), and neither benefits from schema-informed prompting. We release the framework specification, annotation guidelines, and calibration set to support reproducibility.

语义理解文化适配多语种AI框架设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。