TY - GEN
T1 - EVALUATING INTER-RATER CONSISTENCY OF GENERATIVE AI SYSTEMS IN RANKING MOBILE EDUCATION APPS
T2 - 22nd International Conference on Mobile Learning and 11th International Conference on Educational Technologies 2026, ML ICEduTech 2026
AU - Chan, Victor K.Y.
N1 - Publisher Copyright:
© ML ICE du Tech 2026.
PY - 2026
Y1 - 2026
N2 - This study examines the inter-rater consistency of three generative artificial intelligence (GenAI) systems-Microsoft Copilot, Google PaLM, and Assistant-in ranking 100 mobile education applications (apps) across eight evaluative dimensions: (1) content/course quality, (2) pedagogical design, (3) learner support, (4) technology infrastructure, (5) social interaction, (6) learner engagement, (7) instructor support, and (8) cost-effectiveness. Extending earlier rating score-based analyses, the present research focuses on rank-based agreement among the systems, which is less sensitive to calibration differences and more directly relevant to decision-making. Using Spearman's rank correlation coefficient (ρ), Kendall's tau-b (τ-b), Rank-Biased Overlap (RBO), Wilcoxon signed-rank tests, and the ordinal alpha, the study provides a comprehensive assessment of consistency among the GenAI systems. Results indicate that dimensions (1) content/course quality, (2) pedagogical design, (6) learner engagement, and (7) instructor support exhibit moderate-to-strong rank correlations across systems, with ordinal alpha values approaching or exceeding 0.70. By contrast, (5) social interaction and (8) cost-effectiveness show weaker and more inconsistent agreement, with low RBO values highlighting divergence in top-ranked apps. Wilcoxon tests reveal systematic directional differences in several dimensions, confirming that each system applies distinct evaluative tendencies. These findings underscore both the potential and the limitations of GenAI systems as evaluators of educational technology, suggesting that while they can provide reliable rankings in some dimensions, caution is warranted in others. The study highlights the value of rank-based approaches for understanding inter-rater consistency and offers implications for educators, developers, and policymakers. The study concludes by discussing methodological implications, practical applications, and avenues for future research.
AB - This study examines the inter-rater consistency of three generative artificial intelligence (GenAI) systems-Microsoft Copilot, Google PaLM, and Assistant-in ranking 100 mobile education applications (apps) across eight evaluative dimensions: (1) content/course quality, (2) pedagogical design, (3) learner support, (4) technology infrastructure, (5) social interaction, (6) learner engagement, (7) instructor support, and (8) cost-effectiveness. Extending earlier rating score-based analyses, the present research focuses on rank-based agreement among the systems, which is less sensitive to calibration differences and more directly relevant to decision-making. Using Spearman's rank correlation coefficient (ρ), Kendall's tau-b (τ-b), Rank-Biased Overlap (RBO), Wilcoxon signed-rank tests, and the ordinal alpha, the study provides a comprehensive assessment of consistency among the GenAI systems. Results indicate that dimensions (1) content/course quality, (2) pedagogical design, (6) learner engagement, and (7) instructor support exhibit moderate-to-strong rank correlations across systems, with ordinal alpha values approaching or exceeding 0.70. By contrast, (5) social interaction and (8) cost-effectiveness show weaker and more inconsistent agreement, with low RBO values highlighting divergence in top-ranked apps. Wilcoxon tests reveal systematic directional differences in several dimensions, confirming that each system applies distinct evaluative tendencies. These findings underscore both the potential and the limitations of GenAI systems as evaluators of educational technology, suggesting that while they can provide reliable rankings in some dimensions, caution is warranted in others. The study highlights the value of rank-based approaches for understanding inter-rater consistency and offers implications for educators, developers, and policymakers. The study concludes by discussing methodological implications, practical applications, and avenues for future research.
KW - Generative Artificial Intelligence
KW - Inter-Rater Consistency
KW - Mobile Education Applications
KW - Non-Parametric Statistics
KW - Ranking Analysis
UR - https://www.scopus.com/pages/publications/105042513500
M3 - Conference contribution
AN - SCOPUS:105042513500
T3 - 22nd International Conference Mobile Learning and 11th Educational Technologies, ML ICEduTech 2026
SP - 31
EP - 38
BT - 22nd International Conference Mobile Learning and 11th Educational Technologies, ML ICEduTech 2026
A2 - Sanchez, Inmaculada Arnedillo
A2 - Kommers, Piet
A2 - Issa, Tomayess
A2 - Isaias, Pedro
A2 - Rodrigues, Luis
PB - IADIS
Y2 - 7 March 2026 through 9 March 2026
ER -