Evaluation on embeddings application for Spanish automatic text clustering
Resumen
The vast amount of information on the Internet, primarily composed of texts, makes clustering reliable information a complicated task. This research aims to improve the automatic clustering of Spanish texts by applying embeddings and unsupervised learning algorithms. Five datasets were used, and embedding generation techniques such as Word2Vec, FastText, Glove, BERT, and GPT-2 were applied. Models like K-means, HDBSCAN, and AutoEncoder combined with K-means were employed for clustering. The results showed that the AutoEncoder model combined with the K-means using Glove embeddings achieved superior performance with an accuracy of 0.92, NMI of 0.79, and ARI of 0.81 on the BBC News dataset. In other datasets, the results varied, but the AutoEncoder with the K-means model consistently outperformed other methods. We conclude that neural network models with AutoEncoder and K-means layer are highly effective for automatically clustering Spanish texts, especially when using high-quality embeddings like Glove.
