MuTSE让人工参与评估LLM文本简化效果,支持多策略实时对比。
MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator

- 构建交互式网页工具,支持任意提示词与模型组合并行测试
- 引入分层语义对齐引擎和线性偏差修正,提升简化效果可视化精度
- 适合语言模型研究者与智能教学系统开发者快速验证优化方案
随着大语言模型在文本简化中的广泛应用,如何系统评估其在不同提示策略和模型架构下的输出效果,成为自然语言处理研究与智能辅导系统中的关键方法挑战。当前研究多依赖静态计算脚本,教育者则受限于标准对话界面,均无法实现提示-模型组合的多维系统评估。为此,我们提出MuTSE——一个面向人机协同的交互式网页应用,可针对任意CEFR水平目标,高效评估大模型生成的文本简化结果。系统支持$P \times M$种提示-模型组合并发执行,实时生成完整对比矩阵。通过集成新型分层语义对齐引擎及线性偏差启发式($λ$),MuTSE将原文与简化句进行可视化映射,显著降低定性分析的认知负担,并支持可复现、结构化的标注,助力下游NLP数据集构建。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become increasingly prevalent in text simplification, systematically evaluating their outputs across diverse prompting strategies and architectures remains a critical methodological challenge in both NLP research and Intelligent Tutoring Systems (ITS). Developing robust prompts is often hindered by the absence of structured, visual frameworks for comparative text analysis. While researchers typically rely on static computational scripts, educators are constrained to standard conversational interfaces -- neither paradigm supports systematic multi-dimensional evaluation of prompt-model permutations. To address these limitations, we introduce \textbf{MuTSE}\footnote{The project code and the demo have been made available for peer review at the following anonymized URL. https://osf.io/njs43/overview?view_only=4b4655789f484110a942ebb7788cdf2a, an interactive human-in-the-loop web application designed to streamline the evaluation of LLM-generated text simplifications across arbitrary CEFR proficiency targets. The system supports concurrent execution of $P \times M$ prompt-model permutations, generating a comprehensive comparison matrix in real-time. By integrating a novel tiered semantic alignment engine augmented with a linearity bias heuristic ($λ$), MuTSE visually maps source sentences to their simplified counterparts, reducing the cognitive load associated with qualitative analysis and enabling reproducible, structured annotation for downstream NLP dataset construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。