Cross-Domain Cyberbullying Detection in Spanish: A Transfer Learning Approach with Theatrical Corpora and Adversarial Validation
Resumen
Current cyberbullying detection models suffer from limited generalizability across social media platforms due to dataset-specific biases. To address this challenge, we propose a robust transfer learning framework that leverages a unique Spanish-language corpus of annotated theatrical scripts. These scripts offer rich, contextually diverse language examples, significantly improving pre-training effectiveness. We further enhance model generalization through adversarial training and data augmentation techniques, generating synthetic examples to enhance robustness against novel cyberbullying patterns. Our approach achieves state-of-the-art performance, with 85% accuracy and 82% f1-score in cross-dataset evaluations, outperforming existing methods. The results demonstrate the effectiveness of combining domain-specific pre-training with adversarial validation for cross-platform cyberbullying detection.
