Abstract
Growing data volumes and the rise of AI-assisted interpreting solutions increasingly challenge the assessment of interpreting quality. This study investigates whether large language models (LLMs) can support error-based interpreting performance assessment, which has traditionally relied on human-driven approaches. We used transcripts from 42 English-to-Polish interpreting outputs (Broś et al. 2025), evaluated using the NTR model (Romero Fresco & Pöchhacker 2017). Two trained human evaluators conducted detailed analyses, which were then compared against ChatGPT-4o outputs generated through different prompting strategies. Initial results showed an overlap of up to 85% in omission detection between human and AI assessments when a full prompting sequence was applied and the source and target texts were provided in the chat's interface rather than in uploaded text files. However, bulk and iterative analyses revealed significant limitations, including prompt drift and reduced accuracy over successive tasks, with the lowest overlaps below 10%. Despite these drawbacks, the AI model occasionally identified errors overlooked by humans and demonstrated potential for expediting error annotation. The study highlights that, while LLMs provide partial automation and enhance consistency, comprehensive human oversight remains indispensable. Ultimately, integrating AI with human expertise can become a promising hybrid approach to interpreting quality evaluation, balancing efficiency with nuanced judgment.

This work is licensed under a Creative Commons Attribution 4.0 International License.
Copyright (c) 2026 Tomasz Korybski, Karolina Broś, Małgorzata Szupica-Pyrzanowska

