Abstract
This article explores the convergent validity-or, more precisely, the inter-rater consistency-of popular generative artificial intelligence (AI) models in evaluating the quality of project management (PM) software. This study employed three prominent generative AI models-Gemini, CoPilot on Edge, and DeepSeek-to independently assign rating scores (1-10) to a curated list of 50 top-tier PM software systems/tools. The evaluation was structured around eight critical dimensions. Statistical analyses were conducted to assess the consistency of the AI-generated ratings. These included descriptive statistics to measure rating discrimination, analysis of mean absolute differences, paired-samples t-tests to identify erratic and systematic rating biases between AI model pairs, and Cronbach's alpha to determine overall inter-rater consistency for each dimension. The results reveal a high degree of inter-rater consistency across all eight dimensions, a finding juxtaposed with the presence of significant systematic biases between the models. Cronbach's alpha coefficients were found to be acceptable for all dimensions (α = .782 to .897), indicating that the models consistently rate the PM software in a similar order. However, the t-tests confirmed that each model possessed a distinct "rating personality," consistently scoring higher or lower than its counterparts. These findings suggest that while generative AI is a surprisingly trustworthy measure for rating PM software roughly in a particular order, decision-makers must account for the distinct rating biases of each model. Relying on a single model for absolute rating scores is ill-advised, but leveraging multiple models to establish a consensus on relative quality is a viable and powerful new approach for PM software evaluation.
| Original language | English |
|---|---|
| Pages (from-to) | 1967-1974 |
| Number of pages | 8 |
| Journal | Procedia Computer Science |
| Volume | 278 |
| DOIs | |
| Publication status | Published - 2026 |
| Event | International Conference on ENTERprise Information Systems, CENTERIS 2025, International Conference on Project MANagement, ProjMAN 2025, International Conference on Health and Social Care Information Systems and Technologies, HCist 2025 - Hybrid, Abu Dhabi, United Arab Emirates Duration: 26 Nov 2025 → 28 Nov 2025 |
Keywords
- AI consistency
- Convergent validity
- generative AI
- inter-rater consistency
- project management software
- software evaluation
Fingerprint
Dive into the research topics of 'The Convergent Validity of Project Management Software Evaluation by Generative Artificial Intelligence Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver