EventsThe 6th International Electronic Conference on Applied Sciences
Published
This submission belongs to the session S3. Computing and Artificial Intelligence of the event The 6th International Electronic Conference on Applied Sciences
Published date
03 Dec, 2025
Academic Editor
author-avatarLucia Billeci
Citation
Abubakar Ado, Abdurrauf Sharifai Garba, Mansir Abubakar, Bashir Salisu, Usman Mahmud, Abdulkadir Bichi Abubakar, Improved Taxonomy Re-structuring using Modified K-means Clustering for Efficient Large-scale Text Classification, in Proceedings of The 6th International Electronic Conference on Applied Sciences, 9 December–11 December 2025, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Improved Taxonomy Re-structuring using Modified K-means Clustering for Efficient Large-scale Text Classification

Abdurrauf Sharifai Garba 1
Abdulkadir Bichi Abubakar 1
1. Faculty of Computing, Northwest University, Kano, Nigeria, Nigeria
2. Faculty of Computer Science and Mathematics, Universiti Teknologi Mara, Shah Alam, Selangor, Malaysia, Nigeria
3. Department of Computer Science, Faculty of Computing and Mathematical Science, Aliko Dangote University of Science and Technology, Wudil, Nigeria, Nigeria
4. Department of Software Engineering, Northwest University, Kano, Nigeria, Nigeria
Abstract

Textual classification for a hierarchical taxonomy of classes is a common and well-known problem associated with Large-Scale Text classifications (LSTCs). Existing approaches simply re-structure the hierarchy of classes prior to classification and have achieved better results. However, when there are many classes with an increased number of features, traditional hierarchy re-structuring tends to produce many nodes with similar granularities. This results in misclassification, and it is computationally expensive or not scalable for many classification models, especially when the hierarchy is longer. In this paper, we propose an improved hierarchy re-structuring algorithm that uses modified k-means clustering. The method uses a k-weight and backtracking, where necessary, to cluster nodes with similar granularities into a few generalized classes, reducing the number of nodes and hierarchy length as well. In addition, the proposed approach can handle overfitting, which usually occurs as a result of the unbalanced nature of LSHT datasets, where the features in each class vary extensively. Experimental results on 20NG, IPC, and DMOZ-small datasets using TD-LR and TD-SVM show that our approach can effectively improve large-scale hierarchical text classification performance over traditional and existing re-structuring approaches. In terms of scalability, our approach increases the number of scalable instances by about 10%; hence, it records the best and fastest running time.

Keywords
Hierarchical Classification
Hierarchy
Large-scale
Re-structuring,
TD-SVM
TD-LR
Poster
Poster-sciforum-140813.pdf
Hybrid VGG19-TCN with Multi-Channel Temporal Attention for Phishing Attack Detection
PLA membranes as functional alternatives to PVC in potentiometric pH electrodes for wine and beverages.