用LLM+层次分析法,让用户标准变可解释的评分指标
Transforming User Defined Criteria into Explainable Indicators with an Integrated LLM AHP System
- 用LLM生成具体标准评分,再通过层次分析法算权重
- 在亚马逊评论和抑郁文本评估中表现接近传统方法
- 适合需要快速解释结果的在线推荐系统
跨领域的复杂文本评估需将用户自定义标准转化为可量化的可解释指标,这是搜索与推荐系统中的长期难题。单次提示的LLM评估存在复杂度高、延迟大的问题,而针对特定标准的分解方法则依赖粗糙平均或不透明的黑箱聚合方式。本文提出一种可解释的聚合框架,结合LLM评分与层次分析法(AHP)。该方法通过LLM作为裁判生成各标准得分,利用Jensen-Shannon距离衡量判别力,并通过AHP成对比较矩阵推导出统计上合理的权重。在亚马逊评论质量评估和抑郁相关文本评分任务上的实验表明,该方法在保持相当预测能力的同时,显著提升可解释性与运行效率,适用于对延迟敏感的实时网络服务。
原文摘要 · Abstract (English)
Evaluating complex texts across domains requires converting user defined criteria into quantitative, explainable indicators, which is a persistent challenge in search and recommendation systems. Single prompt LLM evaluations suffer from complexity and latency issues, while criterion specific decomposition approaches rely on naive averaging or opaque black-box aggregation methods. We present an interpretable aggregation framework combining LLM scoring with the Analytic Hierarchy Process. Our method generates criterion specific scores via LLM as judge, measures discriminative power using Jensen Shannon distance, and derives statistically grounded weights through AHP pairwise comparison matrices. Experiments on Amazon review quality assessment and depression related text scoring demonstrate that our approach achieves high explainability and operational efficiency while maintaining comparable predictive power, making it suitable for real time latency sensitive web services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。