Publicación

Evaluation on embeddings application for Spanish automatic text clustering

Evaluación de la aplicación de embeddings para el agrupamiento automático de textos en español
Anthony Wainer Cachay-Guivin · Cachay-Guivin A.W.

Resumen

The vast amount of information on the Internet, primarily composed of texts, makes clustering reliable information a complicated task. This research aims to improve the automatic clustering of Spanish texts by applying embeddings and unsupervised learning algorithms. Five datasets were used, and embedding generation techniques such as Word2Vec, FastText, Glove, BERT, and GPT-2 were applied. Models like K-means, HDBSCAN, and AutoEncoder combined with K-means were employed for clustering. The results showed that the AutoEncoder model combined with the K-means using Glove embeddings achieved superior performance with an accuracy of 0.92, NMI of 0.79, and ARI of 0.81 on the BBC News dataset. In other datasets, the results varied, but the AutoEncoder with the K-means model consistently outperformed other methods. We conclude that neural network models with AutoEncoder and K-means layer are highly effective for automatically clustering Spanish texts, especially when using high-quality embeddings like Glove.

Autores y colaboradores

Authors

Anthony Wainer Cachay-Guivin
Cachay-Guivin A.W.

Palabras clave

Artificial intelligence Datasets Model analysis Natural processing language