Skip to main navigation Skip to search Skip to main content

EVALUATING INTER-RATER CONSISTENCY OF GENERATIVE AI SYSTEMS IN RANKING MOBILE EDUCATION APPS: A NON-PARAMETRIC EXTENSION

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This study examines the inter-rater consistency of three generative artificial intelligence (GenAI) systems-Microsoft Copilot, Google PaLM, and Assistant-in ranking 100 mobile education applications (apps) across eight evaluative dimensions: (1) content/course quality, (2) pedagogical design, (3) learner support, (4) technology infrastructure, (5) social interaction, (6) learner engagement, (7) instructor support, and (8) cost-effectiveness. Extending earlier rating score-based analyses, the present research focuses on rank-based agreement among the systems, which is less sensitive to calibration differences and more directly relevant to decision-making. Using Spearman's rank correlation coefficient (ρ), Kendall's tau-b (τ-b), Rank-Biased Overlap (RBO), Wilcoxon signed-rank tests, and the ordinal alpha, the study provides a comprehensive assessment of consistency among the GenAI systems. Results indicate that dimensions (1) content/course quality, (2) pedagogical design, (6) learner engagement, and (7) instructor support exhibit moderate-to-strong rank correlations across systems, with ordinal alpha values approaching or exceeding 0.70. By contrast, (5) social interaction and (8) cost-effectiveness show weaker and more inconsistent agreement, with low RBO values highlighting divergence in top-ranked apps. Wilcoxon tests reveal systematic directional differences in several dimensions, confirming that each system applies distinct evaluative tendencies. These findings underscore both the potential and the limitations of GenAI systems as evaluators of educational technology, suggesting that while they can provide reliable rankings in some dimensions, caution is warranted in others. The study highlights the value of rank-based approaches for understanding inter-rater consistency and offers implications for educators, developers, and policymakers. The study concludes by discussing methodological implications, practical applications, and avenues for future research.

Original languageEnglish
Title of host publication22nd International Conference Mobile Learning and 11th Educational Technologies, ML ICEduTech 2026
EditorsInmaculada Arnedillo Sanchez, Piet Kommers, Tomayess Issa, Pedro Isaias, Luis Rodrigues
PublisherIADIS
Pages31-38
Number of pages8
ISBN (Electronic)9798331336851
Publication statusPublished - 2026
Event22nd International Conference on Mobile Learning and 11th International Conference on Educational Technologies 2026, ML ICEduTech 2026 - Zagreb, Croatia
Duration: 7 Mar 20269 Mar 2026

Publication series

Name22nd International Conference Mobile Learning and 11th Educational Technologies, ML ICEduTech 2026

Conference

Conference22nd International Conference on Mobile Learning and 11th International Conference on Educational Technologies 2026, ML ICEduTech 2026
Country/TerritoryCroatia
CityZagreb
Period7/03/269/03/26

Keywords

  • Generative Artificial Intelligence
  • Inter-Rater Consistency
  • Mobile Education Applications
  • Non-Parametric Statistics
  • Ranking Analysis

Fingerprint

Dive into the research topics of 'EVALUATING INTER-RATER CONSISTENCY OF GENERATIVE AI SYSTEMS IN RANKING MOBILE EDUCATION APPS: A NON-PARAMETRIC EXTENSION'. Together they form a unique fingerprint.

Cite this