One solution to nearest neighbor outliers is to link them with the class cluster by adding new points between the outlier and the cluster. In the age of Generative AI, new data points are legitimate, and if they are added equally and only when needed, they might help re-balance the dataset. In this paper, we describe a nearest neighbor algorithm called Step Nearest Neighbor (Step NN), where incorrectly identified points can be re-linked by adding new points between them and the main cluster. The step points can also be noted, so that they can, for instance, be ignored as real cases. As new points were added, it worked only with distance-based metrics, however. While other adaptive versions may have to fine-tune parameters, this method was easier to use and as it could be pre-trained, this offers options for big data and distributed environments. Test results showed that accuracy comparisons with the classic version of k-NN is almost equal, but Step NN can generalise. In the cases where k-NN did well, k-NN should be equal to or better than Step NN. In the cases where k-NN did not do well, Step NN is likely to be better.
Citation: Kieran Greer. Adding link points to improve Nearest Neighbor classification[J]. Applied Computing and Intelligence, 2026, 6(2): 186-193. doi: 10.3934/aci.2026011
One solution to nearest neighbor outliers is to link them with the class cluster by adding new points between the outlier and the cluster. In the age of Generative AI, new data points are legitimate, and if they are added equally and only when needed, they might help re-balance the dataset. In this paper, we describe a nearest neighbor algorithm called Step Nearest Neighbor (Step NN), where incorrectly identified points can be re-linked by adding new points between them and the main cluster. The step points can also be noted, so that they can, for instance, be ignored as real cases. As new points were added, it worked only with distance-based metrics, however. While other adaptive versions may have to fine-tune parameters, this method was easier to use and as it could be pre-trained, this offers options for big data and distributed environments. Test results showed that accuracy comparisons with the classic version of k-NN is almost equal, but Step NN can generalise. In the cases where k-NN did well, k-NN should be equal to or better than Step NN. In the cases where k-NN did not do well, Step NN is likely to be better.
| [1] | E. Fix, J. L. Hodges, Discriminatory analysis: nonparametric discrimination, consistency properties, Project Report, 1951, 4. |
| [2] |
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, et al., Generative adversarial nets, Commun. ACM, 63 (2020), 139–144. https://doi.org/10.1145/3422622 doi: 10.1145/3422622
|
| [3] | K. Greer, New ideas for brain modelling 2, In: Intelligent systems in science and information 2014, Cham: Springer, 2014, 23–39. https://doi.org/10.1007/978-3-319-14654-6_2 |
| [4] |
R. Halder, M. Uddin, M. Uddin, S. Aryal, A. Khraisat, Enhancing K-nearest neighbor algorithm: a comprehensive review and performance analysis of modifications, J. Big Data, 11 (2024), 113. https://doi.org/10.1186/s40537-024-00973-y doi: 10.1186/s40537-024-00973-y
|
| [5] |
S. Uddin, I. Haque, H. Lu, M. Moni, E. Gide, Comparative performance analysis of K-nearest neighbour (KNN) algorithm and its different variants for disease prediction, Sci. Rep., 12 (2022), 6256. https://doi.org/10.1038/s41598-022-10358-x doi: 10.1038/s41598-022-10358-x
|
| [6] | L. Ertöz, M. Steinbach, V. Kumar, Finding clusters of different sizes, shapes, and densities in noisy, high dimensional data, Proceedings of the 2003 SIAM International Conference on Data Mining, 2023, 47–58. https://doi.org/10.1137/1.9781611972733.5 |
| [7] | A. Jain, R. Dubes, Algorithms for clustering data, Englewood Cliffs: Prentice Hall, 1988. |
| [8] |
R. Jarvis, E. Patrick, Clustering using a similarity measure based on shared nearest neighbors, IEEE Trans. Comput., C-22 (1973), 1025–1034. https://doi.org/10.1109/T-C.1973.223640 doi: 10.1109/T-C.1973.223640
|
| [9] |
K. Greer, A Pattern-Hierarchy classifier for reduced teaching, WSEAS Transactions on Computers, 19 (2020), 183–193. https://doi.org/10.37394/23205.2020.19.23 doi: 10.37394/23205.2020.19.23
|
| [10] |
N. Chawla, K. Bowyer, L. Hall, W. Kegelmeyer, SMOTE: synthetic minority over-sampling technique, J. Artif. Intell. Res., 16 (2002), 321–357. https://doi.org/10.1613/jair.953 doi: 10.1613/jair.953
|
| [11] |
H. Saadatfar, S. Khosravi, J. Joloudari, A. Mosavi, S. Shamshirband, A new K-nearest neighbors classifier for big data based on efficient data pruning, Mathematics, 8 (2020), 286. https://doi.org/10.3390/math8020286 doi: 10.3390/math8020286
|
| [12] | K. Greer, Step NN source code, GitHub, Inc., 2026. Available from: https://github.com/discompsys/Step-NN. |
| [13] |
R. Fisher, The use of multiple measurements in taxonomic problems, Annals of Eugenics, 7 (1936), 179–188. https://doi.org/10.1111/j.1469-1809.1936.tb02137.x doi: 10.1111/j.1469-1809.1936.tb02137.x
|
| [14] | UCI, UCI machine learning repository, UC Irvine, 2026. Available from: http://archive.ics.uci.edu/ml/. |
| [15] | J. Brownlee, 10 standard datasets for practicing applied machine learning, Guiding Tech Media, 2026. Available from: https://machinelearningmastery.com/standard-machine-learning-datasets/. |