Product attributes can be extracted from free-form text content using a large language model (LLM) with high precision. We propose repeating the extraction several times and selecting the most frequent extraction. We measure consistency as the percentage of the most frequent extraction and accept only the most consistent extractions. This approach reduces error rate from 0.5% with the baseline LLM to 0.2%, and, more importantly, eliminates all hallucinations (specificity increases from 99.1% to 100%). The method could also be used for LLM-assisted extraction where a human verifies the results by focusing on the less consistent extractions, and those attributes that the LLM did not find values for.
Citation: Eetu Kyyrö, Pasi Fränti. Using the consistency of LLM for better attribute value extraction[J]. Applied Computing and Intelligence, 2026, 6(2): 173-185. doi: 10.3934/aci.2026010
Product attributes can be extracted from free-form text content using a large language model (LLM) with high precision. We propose repeating the extraction several times and selecting the most frequent extraction. We measure consistency as the percentage of the most frequent extraction and accept only the most consistent extractions. This approach reduces error rate from 0.5% with the baseline LLM to 0.2%, and, more importantly, eliminates all hallucinations (specificity increases from 99.1% to 100%). The method could also be used for LLM-assisted extraction where a human verifies the results by focusing on the less consistent extractions, and those attributes that the LLM did not find values for.
| [1] | U. Asklund, A. Dahlqvist, Implementing and integrating product data management and software configuration management, 2003, Norwood: Artech House. |
| [2] | Akeneo, 2023 B2C survey results report: what shoppers want, Akeneo company, 2025. Available from: https://www.akeneo.com/white-paper/2023-b2c-survey-results-report/. |
| [3] | J. Abraham, Product information management: theory and practice, Cham: Springer, 2014. https://doi.org/10.1007/978-3-319-04885-7 |
| [4] |
G. Lample, M. Ballesteros, S. Subramanian, K. Kawakami, C. Dyer, Neural architectures for named entity recognition, Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2016,260–270. https://doi.org/10.18653/v1/N16-1030 doi: 10.18653/v1/N16-1030
|
| [5] | A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, et al., Attention is all you need, Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017, 5998–6008. |
| [6] |
J. Devlin, M. Chang, K. Lee, K. Toutanova, BERT: pre-training of deep bidirectional transformers for language understanding, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1, 2019, 4171–4186. https://doi.org/10.18653/v1/N19-1423 doi: 10.18653/v1/N19-1423
|
| [7] | Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, et al., RoBERTa: a robustly optimized BERT pretraining approach, arXiv: 1907.11692. https://doi.org/10.48550/arXiv.1907.11692 |
| [8] | J. Kunz, Understandinglarge language models: towards rigorous and targeted interpretability using probing classifiers and self-rationalisation, Ph. D. Thesis, Linköping University, 2024. |
| [9] |
J. Gong, H. Eldardiry, Multi-label zero-shot product attribute-value extraction, Proceedings of the ACM Web Conference 2024, 2024, 2259–2270. https://doi.org/10.1145/3589334.3645649 doi: 10.1145/3589334.3645649
|
| [10] |
J. Dagdelen, A. Dunn, S. Lee, N. Walker, A. Rosen, G. Ceder, et al., Structured information extraction from scientific text with large language models, Nat. Commun. , 15 (2024), 1418. https://doi.org/10.1038/s41467-024-45563-x doi: 10.1038/s41467-024-45563-x
|
| [11] | A. Brinkmann, N. Baumann, C. Bizer, Using LLMs for the extraction and normalization of product attribute values, In: Advances in databases and information systems, Cham: Springer, 2024,217–230. https://doi.org/10.1007/978-3-031-70626-4_15 |
| [12] |
S. Baruah, S. Narayanan, Character attribute extraction from movie scripts using LLMs, Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, 8270–8275. https://doi.org/10.1109/ICASSP48485.2024.10447353 doi: 10.1109/ICASSP48485.2024.10447353
|
| [13] | H. Raj, V. Gupta, D. Rosati, S. Majumdar, Improving consistency in large language models through chain of guidance, arXiv: 2502.15924. https://doi.org/10.48550/arXiv.2502.15924 |
| [14] | X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, et al., Self-consistency improves chain of thought reasoning in language models, arXiv: 2203.11171. https://doi.org/10.48550/arXiv.2203.11171 |
| [15] |
F. Cheng, V. Zouhar, S. Arora, M. Sachan, H. Strobelt, M. El-Assady, RELIC: investigating large language model responses using self-consistency, Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 2024,647. https://doi.org/10.1145/3613904.3641904 doi: 10.1145/3613904.3641904
|
| [16] |
M. Rivera, J. Godbout, R. Rabbany, K. Pelrine, Combining confidence elicitation and sample-based methods for uncertainty quantification in misinformation mitigation, Proceedings of the 1st Workshop on Uncertainty-Aware NLP (UncertaiNLP 2024), 2024,114–126. https://doi.org/10.18653/v1/2024.uncertainlp-1.12 doi: 10.18653/v1/2024.uncertainlp-1.12
|
| [17] | Statistics Finland, Tulorekisterin palkat ja palkkiot (Finnish), Tilastokeskus, 2025. Available from: https://stat.fi/tup/kokeelliset-tilastot/tulorekisterin_palkat_ja_palkkiot/index.html. |