Deep active learning reduces annotation costs by querying only the most informative samples, but it depends on human annotators who are often unavailable in data-scarce settings. We propose deep active learning with artificial intelligence chatbots for semi-supervised learning (DAL-Chat-SSL), a hybrid framework that replaces human annotators with large language model (LLM) chatbots as labeling oracles and extends this paradigm—previously confined to text classification—to tabular, time-series, and image modalities. The framework introduces an embedding-space entropy acquisition function based on inverse-distance weighting in the learned representation, designed to address the miscalibration of softmax outputs that limits conventional predictive entropy under small labeled pools. Two chatbots, ChatGPT and Claude, serve as annotation oracles under a documented prompting protocol; their annotation accuracy is measured on held-out probe sets rather than assumed. We integrate active acquisition with semi-supervised self-training and evaluate the framework on six benchmarks (Sonar, Ionosphere, Optdigits, ECG5000, MNIST binary, MNIST 10-class) over 10 independent runs each, using a held-out test split that never participates in training, acquisition, or model selection. Against a fully supervised reference trained on all available labels, the framework recovers 84.8–99.8% of attainable performance while annotating 4.4–37% of the pool with no human involvement; on Ionosphere, 46 chatbot annotations are statistically indistinguishable from full supervision on 208 ground truth labels. Annotation error rates of 14–33% cost between one and seven points of final performance, and per-class analysis confirms that no class is lost to annotation noise. Acquisition strategies are separate from random sampling on exactly one of the six benchmarks, the one on which the initial model is already close to its ceiling; we relate this to the position of the model on its learning curve rather than to the annotation budget alone. The results indicate that LLM oracles, combined with semi-supervised self-training, offer a practical route to label-efficient learning beyond the text domain.
Citation: Necla Kochan. DAL-Chat-SSL: Deep active learning with LLM oracles and embedding-space entropy for multi-modal classification[J]. AIMS Mathematics, 2026, 11(10): 32452-32483. doi: 10.3934/math.20261277
Deep active learning reduces annotation costs by querying only the most informative samples, but it depends on human annotators who are often unavailable in data-scarce settings. We propose deep active learning with artificial intelligence chatbots for semi-supervised learning (DAL-Chat-SSL), a hybrid framework that replaces human annotators with large language model (LLM) chatbots as labeling oracles and extends this paradigm—previously confined to text classification—to tabular, time-series, and image modalities. The framework introduces an embedding-space entropy acquisition function based on inverse-distance weighting in the learned representation, designed to address the miscalibration of softmax outputs that limits conventional predictive entropy under small labeled pools. Two chatbots, ChatGPT and Claude, serve as annotation oracles under a documented prompting protocol; their annotation accuracy is measured on held-out probe sets rather than assumed. We integrate active acquisition with semi-supervised self-training and evaluate the framework on six benchmarks (Sonar, Ionosphere, Optdigits, ECG5000, MNIST binary, MNIST 10-class) over 10 independent runs each, using a held-out test split that never participates in training, acquisition, or model selection. Against a fully supervised reference trained on all available labels, the framework recovers 84.8–99.8% of attainable performance while annotating 4.4–37% of the pool with no human involvement; on Ionosphere, 46 chatbot annotations are statistically indistinguishable from full supervision on 208 ground truth labels. Annotation error rates of 14–33% cost between one and seven points of final performance, and per-class analysis confirms that no class is lost to annotation noise. Acquisition strategies are separate from random sampling on exactly one of the six benchmarks, the one on which the initial model is already close to its ceiling; we relate this to the position of the model on its learning curve rather than to the annotation budget alone. The results indicate that LLM oracles, combined with semi-supervised self-training, offer a practical route to label-efficient learning beyond the text domain.
| [1] | J. Deng, W. Dong, R. Socher, L. J. Li, K. Li, L. Fei-Fei, ImageNet: A large-scale hierarchical image database, In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009,248–255. |
| [2] |
M. Wu, C. Li, Z. Yao, Deep active learning for computer vision tasks: Methodologies, applications, and challenges, Appl. Sci., 12 (2022), 8103. https://doi.org/10.3390/app12168103 doi: 10.3390/app12168103
|
| [3] |
S. Budd, E. C. Robinson, B. Kainz, A survey on active learning and human-in-the-loop deep learning for medical image analysis, Med. Image Anal., 71 (2021), 102062. https://doi.org/10.1016/j.media.2021.102062 doi: 10.1016/j.media.2021.102062
|
| [4] |
Y. Tang, Y. Zhou, T. Wu, C. Wang, Z. Li, K. Li, AI-driven predictive maintenance for medical imaging equipment: A deep learning framework based on the IoMT data, Reliab. Eng. Syst. Saf., 270 (2026), 112152. https://doi.org/10.1016/j.ress.2025.112152 doi: 10.1016/j.ress.2025.112152
|
| [5] |
C. Ma, H. Li, B. Xue, C. Wang, K. Li, Risk-oriented degradation modeling and maintenance decision support for computed tomography X-ray tubes using mechanism-informed learning, Reliab. Eng. Syst. Saf., 275 (2026), 112837. https://doi.org/10.1016/j.ress.2026.112837 doi: 10.1016/j.ress.2026.112837
|
| [6] |
Y. Zhao, H. Li, Z. Li, K. Li, C. Wang, Hybrid Bayesian and deep learning-based degradation modeling for medical imaging equipment under multi-condition Internet of Medical Things data streams, Reliab. Eng. Syst. Safe., 277 (2027), 112937. https://doi.org/10.1016/j.ress.2026.112937 doi: 10.1016/j.ress.2026.112937
|
| [7] | B. Settles, Active learning literature survey, Computer Sciences Technical Report 1648, University of Wisconsin–Madison, 2010. Available from: https://burrsettles.com/pub/settles.activelearning.pdf. |
| [8] |
P. Kumar, A. Gupta, Active learning query strategies for classification, regression, and clustering: A survey, J. Comput. Sci. Technol., 35 (2020), 913–945. https://doi.org/10.1007/s11390-020-9487-4 doi: 10.1007/s11390-020-9487-4
|
| [9] |
D. Li, Z. Wang, Y. Chen, R. Jiang, W. Ding, M. Okumura, A survey on deep active learning: Recent advances and new frontiers, IEEE Trans. Neural Netw. Learn. Syst., 36 (2025), 5879–5899. https://doi.org/10.1109/TNNLS.2024.3396463 doi: 10.1109/TNNLS.2024.3396463
|
| [10] |
P. Ren, Y. Xiao, X. Chang, P. Y. Huang, Z. Li, B. B. Gupta, et al., A survey of deep active learning, ACM Comput. Surv., 54 (2021), 1–40. https://doi.org/10.1145/3472291 doi: 10.1145/3472291
|
| [11] | X. Zhan, Q. Wang, K. H. Huang, H. Xiong, D. Dou, A. B. Chan, A comparative survey of deep active learning, arXiv preprint, 2022. https://doi.org/10.48550/arXiv.2203.13450 |
| [12] | Y. Gal, R. Islam, Z. Ghahramani, Deep Bayesian active learning with image data, In: Proceedings of the International Conference on Machine Learning (ICML), 2017, 1183–1192. |
| [13] | O. Sener, S. Savarese, Active learning for convolutional neural networks: A core-set approach, In: Proceedings of the International Conference on Learning Representations (ICLR), 2018, arXiv: 1708.00489. |
| [14] | A. Kirsch, J. van Amersfoort, Y. Gal, BatchBALD: Efficient and diverse batch acquisition for deep Bayesian active learning, Adv. Neural Inf. Process. Syst., 32 (2019), 7026–7037. |
| [15] | D. Yoo, I. S. Kweon, Learning loss for active learning, In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 93–102. |
| [16] | S. Sinha, S. Ebrahimi, T. Darrell, Variational adversarial active learning, In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 5972–5981. |
| [17] | K. Sohn, D. Berthelot, N. Carlini, Z. Zhang, H. Zhang, C. A. Raffel, et al., FixMatch: Simplifying semi-supervised learning with consistency and confidence, Adv. Neural Inf. Process. Syst., 33 (2020), 596–608. |
| [18] | B. Zhang, Y. Wang, W. Hou, H. Wu, J. Wang, M. Okumura, et al., FlexMatch: Boosting semi-supervised learning with curriculum pseudo labeling, Adv. Neural Inf. Process. Syst., 34 (2021), 18408–18419. |
| [19] | D. Berthelot, N. Carlini, I. Goodfellow, N. Papernot, A. Oliver, C. A. Raffel, MixMatch: A holistic approach to semi-supervised learning, Adv. Neural Inf. Process. Syst., 32 (2019), 5049–5059. |
| [20] |
X. Yang, Z. Song, I. King, Z. Xu, A survey on deep semi-supervised learning, IEEE Trans. Knowl. Data Eng., 35 (2023), 8934–8954. https://doi.org/10.1109/TKDE.2022.3220219 doi: 10.1109/TKDE.2022.3220219
|
| [21] |
T. Fredriksson, J. Bosch, H. H. Olsson, An empirical evaluation of deep semi-supervised learning, Int. J. Data Sci. Anal., 20 (2025), 4127–4148. https://doi.org/10.1007/s41060-024-00713-8 doi: 10.1007/s41060-024-00713-8
|
| [22] |
Z. Wen, O. Pizarro, S. Williams, Active self-semi-supervised learning for few labeled samples, Neurocomputing, 614 (2025), 128772. https://doi.org/10.1016/j.neucom.2024.128772 doi: 10.1016/j.neucom.2024.128772
|
| [23] |
T. Wan, K. Xu, T. Yu, X. Wang, D. Feng, B. Ding, et al., A survey of deep active learning for foundation models, Intell. Comput., 2 (2023), 0058. https://doi.org/10.34133/icomputing.0058 doi: 10.34133/icomputing.0058
|
| [24] |
K. Wang, D. Zhang, Y. Li, R. Zhang, L. Lin, Cost-effective active learning for deep image classification, IEEE T. Circ. Syst. Vid., 27 (2017), 2591–2600. https://doi.org/10.1109/TCSVT.2016.2589879 doi: 10.1109/TCSVT.2016.2589879
|
| [25] | S. Song, D. Berthelot, A. Rostamizadeh, Combining MixMatch and active learning for better accuracy with fewer labels, arXiv preprint, 2019. https://doi.org/10.48550/arXiv.1912.00594 |
| [26] | J. Guo, H. Shi, Y. Kang, K. Kuang, S. Tang, Z. Jiang, et al., Semi-supervised active learning for semi-supervised models: Exploit adversarial examples with graph-based virtual labels, In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, 2896–2905. |
| [27] |
S. Mittal, J. Niemeijer, Ö. Çiçek, M. Tatarchenko, J. Ehrhardt, J. P. Schäfer, et al., Realistic evaluation of deep active learning for image classification and semantic segmentation, Int. J. Comput. Vis., 133 (2025), 4294–4316. https://doi.org/10.1007/s11263-025-02372-z doi: 10.1007/s11263-025-02372-z
|
| [28] | Y. Xia, S. Mukherjee, Z. Xie, J. Wu, X. Li, R. Aponte, et al., From selection to generation: A survey of LLM-based active learning, In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, 14552–14569. https://doi.org/10.18653/v1/2025.acl-long.708 |
| [29] |
M. Bayer, J. Lutz, C. Reuter, ActiveLLM: Large language model-based active learning for textual few-shot scenarios, Trans. Assoc. Comput. Linguist., 14 (2026), 1–22. https://doi.org/10.1162/TACL.a.63 doi: 10.1162/TACL.a.63
|
| [30] | Y. Qi, X. Yang, J. Lu, G. Guo, J. Enticott, G. Liu, et al., Next generation active learning: Mixture of LLMs in the loop, In: Proceedings of the AAAI Conference on Artificial Intelligence, 40 (2026), 24909–24917. https://doi.org/10.1609/aaai.v40i29.39678 |
| [31] | K. Cho, B. van Merriënboer, C. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, et al., Learning phrase representations using RNN encoder-decoder for statistical machine translation, In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2014, 1724–1734. https://doi.org/10.3115/v1/D14-1179 |
| [32] |
C. E. Shannon, A mathematical theory of communication, Bell Syst. Tech. J., 27 (1948), 379–423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x doi: 10.1002/j.1538-7305.1948.tb01338.x
|
| [33] | D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, In: Proceedings of the International Conference on Learning Representations (ICLR), 2015. |
| [34] | Y. Ji, V. S. Kaza, N. Artham, T. Wang, Deep active learning with manifold-preserving trajectory sampling, arXiv preprint, 2024. https://doi.org/10.48550/arXiv.2410.15605 |