提出九维框架,让AI读懂尼日利亚话语背后的真正意图。
The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse
- 构建九维标注体系,区分表层情绪与真实沟通意图。
- 零样本下语域识别准确率仅33.3%,引入框架后升至73.3%。
- 大模型能力与文化理解可分离,小模型反而更受框架帮助。
我们提出意义智能框架(MIF),一种用于尼日利亚公共话语的九维标注与评估体系,旨在区分表面情绪与真实交际意图。现有尼日利亚语言基准如 NaijaSenti 和 AfriSenti 将情感分类视为三极极性任务。我们认为,当前AI在尼日利亚话语中的主要失败模式并非翻译错误,而是语境缺失:同一语句在不同说话人、受众和情境下可能具有相反的语用效果。MIF通过九个可评分维度实现这一洞察:语域、表面情绪、真实意图、反讽、隐含信息、风险等级、标注者信心、说话人情绪及建议沟通行动。我们构建了一个包含30个样本的校准数据集,涵盖标准英语、尼日利亚英语、尼日利亚皮钦语及混合语码形式。评估了三个前沿语言模型(Gemini 2.5 Flash、GPT-5 和 Gemini 2.5 Pro)在零样本与框架引导提示下的表现。两大核心发现:第一,语域差距——零样本语域分类准确率为33.3%,引入MIF框架后提升至73.3%(+40个百分点);第二,模型能力与文化理解解耦:GPT-5(MIS 67.8)与 Gemini 2.5 Pro(MIS 65.4)得分低于 Flash(MIS 78.6),且均未从框架提示中获益。我们已公开框架规范、标注指南与校准集以支持可复现性。
原文摘要 · Abstract (English)
We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent. Existing benchmarks for Nigerian languages, including NaijaSenti and AfriSenti, treat sentiment classification as a three-way polarity task. We argue that the dominant failure mode of AI systems on Nigerian discourse is not translation failure but context failure: the same utterance carries opposite pragmatic force depending on speaker, audience, and situation. The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent, irony, coded subtext, risk tier, annotator confidence, speaker emotion, and recommended communications action. We construct a 30-item calibration dataset spanning Standard English, Nigerian English, Nigerian Pidgin, and code-mixed registers, and evaluate three frontier language models (Gemini 2.5 Flash, GPT-5, and Gemini 2.5 Pro) under zero-shot and schema-informed prompting conditions. Two headline findings emerge. First, the Register Gap: zero-shot register classification accuracy is 33.3%, rising to 73.3% (+40 points) when the model receives the MIF schema in-context. Second, model capability and cultural competence are decoupled: GPT-5 (MIS 67.8) and Gemini 2.5 Pro (MIS 65.4) score lower than Flash (MIS 78.6), and neither benefits from schema-informed prompting. We release the framework specification, annotation guidelines, and calibration set to support reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。