Large language models (LLMs) have achieved unprecedented performance in various tasks, such as classification, summarization, reasoning, and complex decision-making. But, as the size and parameter count of these models has increased exponentially, their computational demand, electricity consumption, and carbon footprint has also risen dramatically. This rapid scaling has put the global research community in front of a serious environmental concern. In response to this growing concern, researchers are gradually shifting their focus from red AI (compute-intensive, resource-heavy systems) to green AI. Particularly, small language models and optimization techniques like quantization, pruning, and knowledge distillation are being explored as sustainable alternatives to LLMs for domain-specific tasks. Covering the literature from 2020 to 2026, this review paper presents an extensive, deep, and rigorous bibliometric analysis. Using an advanced rule-based keyword preprocessing pipeline and Biblioshiny network mapping, we have mapped the annual publication growth, global research hubs, author influence networks, and dominant thematic clusters. Further, this review paper highlights a major research gap: Although many papers have already been written on energy optimization, publications on standardized carbon emission reporting frameworks are still limited.
Citation: Avinash Shrivastava, Anamika Gupta, Sarabjeet Kaur Kochhar. From carbon-intensive to carbon-neutral: sustainability and energy efficiency in small language models: bibliometric analysis (2020–2026)[J]. Applied Computing and Intelligence, 2026, 6(2): 157-172. doi: 10.3934/aci.2026009
Large language models (LLMs) have achieved unprecedented performance in various tasks, such as classification, summarization, reasoning, and complex decision-making. But, as the size and parameter count of these models has increased exponentially, their computational demand, electricity consumption, and carbon footprint has also risen dramatically. This rapid scaling has put the global research community in front of a serious environmental concern. In response to this growing concern, researchers are gradually shifting their focus from red AI (compute-intensive, resource-heavy systems) to green AI. Particularly, small language models and optimization techniques like quantization, pruning, and knowledge distillation are being explored as sustainable alternatives to LLMs for domain-specific tasks. Covering the literature from 2020 to 2026, this review paper presents an extensive, deep, and rigorous bibliometric analysis. Using an advanced rule-based keyword preprocessing pipeline and Biblioshiny network mapping, we have mapped the annual publication growth, global research hubs, author influence networks, and dominant thematic clusters. Further, this review paper highlights a major research gap: Although many papers have already been written on energy optimization, publications on standardized carbon emission reporting frameworks are still limited.
| [1] | T. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, et al., Language models are few-shot learners, Proceedings of the 34th International Conference on Neural Information Processing Systems, 2020, 1877–1901. |
| [2] | J. Deng, W. Dong, R. Socher, L. Li, K. Li, F. Li, ImageNet: a large-scale hierarchical image database, Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2009,248–255. https://doi.org/10.1109/CVPR.2009.5206848 |
| [3] | J. Devlin, M. W. Chang, K. Lee, K. Toutanova, BERT: pre-training of deep bidirectional transformers for language understanding, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019, 4171–4186. https://doi.org/10.18653/v1/N19-1423 |
| [4] |
A. Krizhevsky, I. Sutskever, G. E. Hinton, ImageNet classification with deep convolutional neural networks, Commun. ACM, 60 (2017), 84–90. https://doi.org/10.1145/3065386 doi: 10.1145/3065386
|
| [5] | G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, S. Han, SmoothQuant: accurate and efficient post-training quantization for large language models, Proceedings of the 40th International Conference on Machine Learning, 2023, 38087–38099. |
| [6] |
Y. Liang, Y. Wang, Y. Li, Y. Zeng, Matrix-transformation based low-rank adaptation (MTLORA): a brain-inspired method for parameter-efficient fine-tuning, Neural Networks, 199 (2026), 108642. https://doi.org/10.1016/j.neunet.2026.108642 doi: 10.1016/j.neunet.2026.108642
|
| [7] |
Y. Wang, B. Yin, GCL: group-shared continual learning fine-tuning for sparse LLMs, Neurocomputing, 675 (2026), 132918. https://doi.org/10.1016/j.neucom.2026.132918 doi: 10.1016/j.neucom.2026.132918
|
| [8] | T. Schick, H. Schütze, It's not just size that matters: small language models are also few-shot learners, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021, 2339–2352. https://doi.org/10.18653/v1/2021.naacl-main.185 |
| [9] |
B. Cui, Y. Li, Z. Zhang, Joint structured pruning and dense knowledge distillation for efficient transformer model compression, Neurocomputing, 458 (2021), 56–69. https://doi.org/10.1016/j.neucom.2021.05.084 doi: 10.1016/j.neucom.2021.05.084
|
| [10] | S. Samsi, D. Zhao, J. McDonald, B. Li, A. Michaleas, M. Jones, et al., From words to watts: benchmarking the energy costs of large language model inference, Proceedings of IEEE High Performance Extreme Computing Conference (HPEC), 2023, 1–9. https://doi.org/10.1109/HPEC58863.2023.10363447 |
| [11] | Y. Li, B. Dong, F. Guerin, C. Lin, Compressing context to enhance inference efficiency of large language models, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, 6342–6353. https://doi.org/10.18653/v1/2023.emnlp-main.391 |
| [12] | C. Guo, J. Tang, W. Hu, J. Leng, C. Zhang, F. Yang, et al., Olive: accelerating large language models via hardware-friendly outlier-victim pair quantization, Proceedings of the 50th Annual International Symposium on Computer Architecture, 2023, 1–15. https://doi.org/10.1145/3579371.3589038 |
| [13] |
X. Zhu, J. Li, Y. Liu, C. Ma, W. Wang, A survey on model compression for large language models, Transactions of the Association for Computational Linguistics, 12 (2024), 1556–1577 https://doi.org/10.1162/tacl_a_00704 doi: 10.1162/tacl_a_00704
|
| [14] | W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, et al., Omniquant: omnidirectionally calibrated quantization for large language models, Proceedings of International Conference on Learning Representations 2024, 2024, 1–25. |
| [15] | S. Zeng, J. Liu, G. Dai, X. Yang, T. Fu, H. Wang, et al., FlightLLM: efficient large language model inference with a complete mapping flow on FPGAs, Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays, 2024,223–234. https://doi.org/10.1145/3626202.3637562 |
| [16] |
C. Lee, J. Jin, T. Kim, H. Kim, E. Park, Owq: outlier-aware weight quantization for efficient fine-tuning and inference of large language models, Proceedings of the AAAI Conference on Artificial Intelligence, 38 (2024), 13355–13364. https://doi.org/10.1609/aaai.v38i12.29237 doi: 10.1609/aaai.v38i12.29237
|
| [17] | S. Gholami, M. Omar, Can a student large language model perform as well as its teacher? In: Innovations, securities, and case studies across healthcare, business, and technology, Hershey: IGI Global Scientific Publishing, 2024,122–139. https://doi.org/10.4018/979-8-3693-1906-2.ch007 |
| [18] |
H. Chen, J. Zhang, Y. Du, S. Xiang, Z. Yue, N. Zhang, et al., Understanding the potential of FPGA-based spatial acceleration for large language model inference, ACM Trans. Reconfig. Techn., 18 (2024), 1–29. https://doi.org/10.1145/3656177 doi: 10.1145/3656177
|
| [19] |
S. Pimenow, O. Pimenowa, P. Prus, Challenges of artificial intelligence development in the context of energy consumption and impact on climate change, Energies, 17 (2024), 5965. https://doi.org/10.3390/en17235965 doi: 10.3390/en17235965
|
| [20] | S. Park, K. Kim, J. So, J. Jung, J. Lee, K. Woo, et al., An LPDDR-based CXL-PNM platform for TCO-efficient inference of transformer-based large language models, Proceedings of IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2024,970–982. https://doi.org/10.1109/HPCA57654.2024.00078 |
| [21] |
M. Argerich, M. Patiño-Martínez, Measuring and improving the energy efficiency of large language models inference, IEEE Access, 12 (2024), 80194–80207. https://doi.org/10.1109/ACCESS.2024.3409745 doi: 10.1109/ACCESS.2024.3409745
|
| [22] |
X. Shen, P. Dong, L. Lu, Z. Kong, Z. Li, M. Lin, et al., Agile-Quant: activation-guided quantization for faster inference of LLMs on the edge, Proceedings of the AAAI Conference on Artificial Intelligence, 38 (2024), 18944–18951. https://doi.org/10.1609/aaai.v38i17.29860 doi: 10.1609/aaai.v38i17.29860
|
| [23] | S. Iftikhar, S. Davy, Reducing carbon footprint in AI: a framework for sustainable training of large language models, Proceedings of the Future Technologies Conference (FTC) 2024, Volume 1, 2024,325–336. https://doi.org/10.1007/978-3-031-73110-5_22 |
| [24] | M. Kim, S. Lee, W. Sung, J. Choi, RA-LoRA: rank-adaptive parameter-efficient fine-tuning for accurate 2-bit quantized large language models, Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2024, 15773–15786. https://doi.org/10.18653/v1/2024.findings-acl.933 |
| [25] |
M. Zhang, X. Shen, J. Cao, Z. Cui, S. Jiang, Edgeshard: efficient llm inference via collaborative edge computing, IEEE Internet Things J., 12 (2025), 13119–13131. https://doi.org/10.1109/JIOT.2024.3524255 doi: 10.1109/JIOT.2024.3524255
|
| [26] |
C. Chen, K. Zhao, J. Leng, C. Liu, J. Fan, P. Zheng, Integrating large language model and digital twins in the context of industry 5.0: Framework, challenges and opportunities, Robot. Comput.-Int. Manuf., 94 (2025), 102982. https://doi.org/10.1016/j.rcim.2025.102982 doi: 10.1016/j.rcim.2025.102982
|
| [27] |
R. Zhang, J. He, X. Luo, D. Niyato, J. Kang, Z. Xiong, et al., Toward democratized generative AI in next-generation mobile edge networks, IEEE Network, 39 (2025), 251–260. https://doi.org/10.1109/MNET.2025.3541078 doi: 10.1109/MNET.2025.3541078
|
| [28] |
G. Kim, S. Hwang, B. Jang, Efficient compressing and tuning methods for large language models: a systematic literature review, ACM Comput. Surv., 57 (2025), 1–39. https://doi.org/10.1145/3728636 doi: 10.1145/3728636
|
| [29] |
A. Sharshar, L. Khan, W. Ullah, M. Guizani, Vision-language models for edge networks: a comprehensive survey, IEEE Internet Things J., 12 (2025), 32701–32724. https://doi.org/10.1109/JIOT.2025.3579032 doi: 10.1109/JIOT.2025.3579032
|
| [30] |
Y. Guo, Z. Hao, J. Shao, J. Zhou, X. Liu, X. Tong, et al., PT-BitNet: scaling up the 1-bit large language model with post-training quantization, Neural Networks, 191 (2025), 107855. https://doi.org/10.1016/j.neunet.2025.107855 doi: 10.1016/j.neunet.2025.107855
|
| [31] | J. Fernandez, C. Na, V. Tiwari, Y. Bisk, S. Luccioni, E. Strubell, Energy considerations of large language model inference and efficiency optimizations, Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2025, 32556–32569. https://doi.org/10.18653/v1/2025.acl-long.1563 |
| [32] | R. Aralimatti, S. Shakhadri, K. Kruthika, K. Angadi, Fine-tuning small language models for domain-specific AI: an edge AI perspective, In: Intelligent systems and applications, Cham: Springer, 2025,503–520. https://doi.org/10.1007/978-3-031-99965-9_31 |
| [33] |
F. Jeanquartier, C. Jean-Quartier, P. Rieder, V. Misirlić, C. Pasero, R. Hohensinner, et al., Assessing the carbon footprint of language models: towards sustainability in AI, Resour. Conserv. Recy., 226 (2026), 108670. https://doi.org/10.1016/j.resconrec.2025.108670 doi: 10.1016/j.resconrec.2025.108670
|