Subjective ranking decisions are widely used in practice, yet their outcomes may reflect voter preferences, reputation, and contextual considerations that are not fully captured by quantitative indicators, making it difficult to assess how closely final decisions align with measurable evidence. To address this challenge, we developed an interpretable intelligent system with an unsupervised ensemble ranking module for data-driven auditing of subjective ranking decisions. The system consists of two components: (1) an unsupervised feature-weighted ensemble ranking module that integrates heterogeneous feature-relevance estimators to construct a within-season performance ranking of the evaluated players without using historical voting outcomes in feature-importance estimation, ensemble-weight selection, or player scoring; and (2) an interpretable diagnostic comparison module that applies the same rule-based model classes under performance-based and voting-based labels to examine agreement and divergence between the two evaluation perspectives. We used the National Basketball Association (NBA) regular-season Most Valuable Player (MVP) award as a representative case study, with the official MVP winner serving as an external voting-based reference for the performance-based ranking. Experiments covering the 1980–2025 award years showed that the framework produces a stable performance-based benchmark and identifies notable seasons, including 2001, 2005, 2006, and 2008, in which the official voting-based outcome diverges markedly from the data-driven ranking. Robustness analyses further showed that feature-attribution changes caused by correlated variables do not substantially propagate to the final ranking structure, while the integrated ranking remains comparatively stable under feature perturbation and correlated-feature ablation. These results demonstrate the value of integrating label-free ranking construction, external winner-based evaluation, and interpretable diagnostic comparison within a unified auditing framework for subjective ranking decisions.
Citation: Zheng Fang, Wei Liang, Xinyue Weng, Dongding Zhang, Junjie Zhou. An interpretable unsupervised ranking system for auditing subjective ranking decisions[J]. AIMS Mathematics, 2026, 11(9): 29636-29671. doi: 10.3934/math.20261176
Subjective ranking decisions are widely used in practice, yet their outcomes may reflect voter preferences, reputation, and contextual considerations that are not fully captured by quantitative indicators, making it difficult to assess how closely final decisions align with measurable evidence. To address this challenge, we developed an interpretable intelligent system with an unsupervised ensemble ranking module for data-driven auditing of subjective ranking decisions. The system consists of two components: (1) an unsupervised feature-weighted ensemble ranking module that integrates heterogeneous feature-relevance estimators to construct a within-season performance ranking of the evaluated players without using historical voting outcomes in feature-importance estimation, ensemble-weight selection, or player scoring; and (2) an interpretable diagnostic comparison module that applies the same rule-based model classes under performance-based and voting-based labels to examine agreement and divergence between the two evaluation perspectives. We used the National Basketball Association (NBA) regular-season Most Valuable Player (MVP) award as a representative case study, with the official MVP winner serving as an external voting-based reference for the performance-based ranking. Experiments covering the 1980–2025 award years showed that the framework produces a stable performance-based benchmark and identifies notable seasons, including 2001, 2005, 2006, and 2008, in which the official voting-based outcome diverges markedly from the data-driven ranking. Robustness analyses further showed that feature-attribution changes caused by correlated variables do not substantially propagate to the final ranking structure, while the integrated ranking remains comparatively stable under feature perturbation and correlated-feature ablation. These results demonstrate the value of integrating label-free ranking construction, external winner-based evaluation, and interpretable diagnostic comparison within a unified auditing framework for subjective ranking decisions.
| [1] |
B. J. Coleman, J. M. DuMond, A. K. Lynch, An examination of NBA MVP voting behavior: Does race matter? J. Sports Econ., 9 (2008), 606–627. https://doi.org/10.1177/1527002508320653 doi: 10.1177/1527002508320653
|
| [2] | Y. Zhai, T. Xu, Novel metric to predict NBA regular season MVP, In: 2024 10th IEEE International Conference on High Performance and Smart Computing (HPSC), New York: IEEE, 2024, 36–42. https://doi.org/10.1109/HPSC62738.2024.00014 |
| [3] |
C. Cheng, Moral hazard in teams with subjective evaluations, RAND J. Econ., 52 (2021), 22–48. https://doi.org/10.1111/1756-2171.12360 doi: 10.1111/1756-2171.12360
|
| [4] |
S. Ishiguro, Y. Yasuda, Moral hazard and subjective evaluation, J. Econ. Theory, 209 (2023), 105619. https://doi.org/10.1016/j.jet.2023.105619 doi: 10.1016/j.jet.2023.105619
|
| [5] |
V. S. Maas, R. Torres-Gonzalez, Subjective performance evaluation and gender discrimination, J. Bus. Ethics, 101 (2011), 667–681. https://doi.org/10.1007/s10551-011-0763-7 doi: 10.1007/s10551-011-0763-7
|
| [6] |
S. Jacob, F. Willits, Objective and subjective indicators of community evaluation: A Pennsylvania assessment, Soc. Indic. Res., 32 (1994), 161–177. https://doi.org/10.1007/BF01078733 doi: 10.1007/BF01078733
|
| [7] |
B. Anderson, R. Ryan, W. Goudy, Consistency in subjective evaluations of community attributes, Soc. Indic. Res., 14 (1984), 165–175. https://doi.org/10.1007/BF00293408 doi: 10.1007/BF00293408
|
| [8] |
A. I. Chatzimouratidis, P. A. Pilavachi, Objective and subjective evaluation of power plants and their non-radioactive emissions using the analytic hierarchy process, Energy Policy, 35 (2007), 4027–4038. https://doi.org/10.1016/j.enpol.2007.02.003 doi: 10.1016/j.enpol.2007.02.003
|
| [9] |
A. Kuhn, In the eye of the beholder: Subjective inequality measures and individuals' assessment of market justice, Eur. J. Polit. Econ., 27 (2011), 625–641. https://doi.org/10.1016/j.ejpoleco.2011.06.002 doi: 10.1016/j.ejpoleco.2011.06.002
|
| [10] |
J. E. Rockoff, C. Speroni, Subjective and objective evaluations of teacher effectiveness: Evidence from New York City, Labour Econ., 18 (2011), 687–696. https://doi.org/10.1016/j.labeco.2011.02.004 doi: 10.1016/j.labeco.2011.02.004
|
| [11] |
C. A. Parsons, J. Sulaeman, M. C. Yates, D. S. Hamermesh, Strike three: Discrimination, incentives, and evaluation, Am. Econ. Rev., 101 (2011), 1410–1435. https://doi.org/10.1257/aer.101.4.1410 doi: 10.1257/aer.101.4.1410
|
| [12] |
H. Choi, S. Kim, School ties and evaluation outcomes: Evidence from the Korean basketball league, Can. J. Econ., 57 (2024), 1337–1359. https://doi.org/10.1111/caje.12746 doi: 10.1111/caje.12746
|
| [13] |
V. Scoppa, Are subjective evaluations biased by social factors or connections? An econometric analysis of soccer referee decisions, Empir. Econ., 35 (2008), 123–140. https://doi.org/10.1007/s00181-007-0146-1 doi: 10.1007/s00181-007-0146-1
|
| [14] |
J. Lee, Outlier aversion in subjective evaluation: Evidence from World Figure Skating Championships, J. Sports Econ., 9 (2008), 141–159. https://doi.org/10.1177/1527002507299203 doi: 10.1177/1527002507299203
|
| [15] | J. Hollinger, Pro Basketball Forecast, Washington: Potomac Books, 2005. |
| [16] | J. Sill, Improved NBA adjusted $\pm$ using regularization and out-of-sample testing, In: Proceedings of the 2010 MIT Sloan Sports Analytics Conference, Cambridge, 2010. |
| [17] | N. Nguyen, B. Ma, J. Hu, Predicting National Basketball Association players performance and popularity: A data mining approach, In: Computational Collective Intelligence: 12th International Conference, ICCCI 2020, Lecture Notes in Computer Science, Cham: Springer, 12496 (2020), 293–304. https://doi.org/10.1007/978-3-030-63007-2_23 |
| [18] |
Y. Ke, R. Bian, R. Chandra, A unified machine learning framework for basketball team roster construction: NBA and WNBA, Appl. Soft Comput., 153 (2024), 111298. https://doi.org/10.1016/j.asoc.2024.111298 doi: 10.1016/j.asoc.2024.111298
|
| [19] | X. Sun, J. Davis, O. Schulte, G. Liu, Cracking the black box: distilling deep sports analytics, In: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020, 3154–3162. https://doi.org/10.1145/3394486.3403367 |
| [20] |
Y. Wang, W. Liu, X. Liu, Explainable AI techniques with application to NBA gameplay prediction, Neurocomputing, 483 (2022), 59–71. https://doi.org/10.1016/j.neucom.2022.01.098 doi: 10.1016/j.neucom.2022.01.098
|
| [21] |
M. Petkovic, D. Kocev, B. Skrlj, S. Dzeroski, Ensemble- and distance-based feature ranking for unsupervised learning, Int. J. Intell. Syst., 36 (2021), 3068–3086. https://doi.org/10.1002/int.22390 doi: 10.1002/int.22390
|
| [22] |
A. Sarkar, S. S. Goswami, D. K. Behera, D. Bozanic, Criteria weighting methods in multi-criteria decision making: A comprehensive review of subjective, objective, and hybrid approaches, J. Contemp. Decis. Sci., 3 (2026), 1–34. https://doi.org/10.67334/cds202619 doi: 10.67334/cds202619
|
| [23] | A. Sarkar, S. S. Goswami, A comprehensive review: The novel weighting methods for multi-criteria decision-making (MCDM), Spec. Oper. Res., 2026, 1–26. https://doi.org/10.31181/sor202781 |
| [24] |
D. Diakoulaki, G. Mavrotas, L. Papayannakis, Determining objective weights in multiple criteria problems: The CRITIC method, Comput. Oper. Res., 22 (1995), 763–770. https://doi.org/10.1016/0305-0548(94)00059-H doi: 10.1016/0305-0548(94)00059-H
|
| [25] |
M. Keshavarz-Ghorabaee, M. Amiri, E. K. Zavadskas, Z. Turskis, J. Antucheviciene, Determination of objective weights using a new method based on the removal effects of criteria (MEREC), Symmetry, 13 (2021), 525. https://doi.org/10.3390/sym13040525 doi: 10.3390/sym13040525
|
| [26] |
F. Ecer, D. Pamucar, A novel LOPCOW-DOBI multi-criteria sustainability performance assessment methodology: An application in developing country banking sector, Omega, 112 (2022), 102690. https://doi.org/10.1016/j.omega.2022.102690 doi: 10.1016/j.omega.2022.102690
|
| [27] |
P. Emerson, The original Borda count and partial voting, Soc. Choice Welfare, 40 (2013), 353–358. https://doi.org/10.1007/s00355-011-0603-9 doi: 10.1007/s00355-011-0603-9
|
| [28] |
J. H. Friedman, B. E. Popescu, Predictive learning via rule ensembles, Ann. Appl. Stat., 2 (2008), 916–954. https://doi.org/10.1214/07-AOAS148 doi: 10.1214/07-AOAS148
|
| [29] |
A. Sarkar, S. S. Goswami, K. K. Gupta, Sensitivity analysis and validation in MCDM methods: A comprehensive review with advancements, applications, and future directions, Spec. Decis. Mak. Appl., 4 (2026), 1–14. https://doi.org/10.31181/sdmap41202769 doi: 10.31181/sdmap41202769
|
| [30] |
C. Strobl, A. L. Boulesteix, A. Zeileis, T. Hothorn, Bias in random forest variable importance measures: Illustrations, sources and a solution, BMC Bioinform., 8 (2007), 25. https://doi.org/10.1186/1471-2105-8-25 doi: 10.1186/1471-2105-8-25
|
| [31] |
B. Gregorutti, B. Michel, P. Saint-Pierre, Correlation and variable importance in random forests, Stat. Comput., 27 (2017), 659–678. https://doi.org/10.1007/s11222-016-9646-1 doi: 10.1007/s11222-016-9646-1
|