arXiv:2603.04416cs.CL2026-03

用可靠性筛选弱信号,提升阿拉伯语情感分析的鲁棒性

Optimizing What We Trust: Reliability-Guided QUBO Selection of Multi-Agent Weak Framing Signals for Arabic Sentiment Prediction

  • 多智能体系统通过分歧与推理质量评估数据可信度
  • 选中的子集在跨域测试中表现更稳定,且保留可迁移结构
  • 适合需要低资源下可靠弱监督的自然语言处理任务

阿拉伯语社交媒体中的框架检测因解释模糊、文化依赖及标注稀缺而困难。现有基于大模型的弱监督方法多依赖标签聚合,在标注少且社会性依赖强时易失效。本文提出一种可靠性感知的弱监督框架,将重点从标签融合转向数据筛选。一个包含两个框架者、一个评论者和一个判别器的小型多智能体流水线,将分歧程度与推理质量作为认知信号,生成实例级可靠性评分。这些评分指导基于QUBO的子集选择过程,实现框架平衡并减少冗余。内在诊断与跨领域阿拉伯语情感迁移测试表明,所选子集更具可靠性,编码了非随机且可迁移的结构,同时不降低纯文本强基线性能。

原文摘要 · Abstract (English)

Framing detection in Arabic social media is difficult due to interpretive ambiguity, cultural grounding, and limited reliable supervision. Existing LLM-based weak supervision methods typically rely on label aggregation, which is brittle when annotations are few and socially dependent. We propose a reliability-aware weak supervision framework that shifts the focus from label fusion to data curation. A small multi-agent LLM pipeline, two framers, a critic, and a discriminator, treats disagreement and reasoning quality as epistemic signals and produces instance-level reliability estimates. These estimates guide a QUBO-based subset selection procedure that enforces frame balance while reducing redundancy. Intrinsic diagnostics and an out-of-domain Arabic sentiment transfer test show that the selected subsets are more reliable and encode non-random, transferable structure, without degrading strong text-only baselines.

弱监督多智能体情感分析阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。