TY - GEN
T1 - An Annotated Corpus of Arabic Tweets for Hate Speech Analysis
AU - Zaghoauni, Wajdi
AU - Rafiul Biswas, Md
N1 - Publisher Copyright:
© 2025 Incoma Ltd. All rights reserved.
PY - 2025/9/10
Y1 - 2025/9/10
N2 - Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10,000 Arabic tweets and annotated each tweet, whether it contains offensive content or not. If a text contains offensive content, we further classify it into different hate speech targets such as religion, gender, politics, ethnicity, origin, and others. A text can contain either single or multiple targets. Multiple annotators are involved in the data annotation task. We calculated the inter-annotator agreement, which was reported to be 0.86 for offensive content and 0.71 for multiple hate speech targets. Finally, we evaluated the data annotation task by employing a different transformers-based model in which AraBERTv2 outperformed with a microF1 score of 0.7865 and an accuracy of 0.786.
AB - Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10,000 Arabic tweets and annotated each tweet, whether it contains offensive content or not. If a text contains offensive content, we further classify it into different hate speech targets such as religion, gender, politics, ethnicity, origin, and others. A text can contain either single or multiple targets. Multiple annotators are involved in the data annotation task. We calculated the inter-annotator agreement, which was reported to be 0.86 for offensive content and 0.71 for multiple hate speech targets. Finally, we evaluated the data annotation task by employing a different transformers-based model in which AraBERTv2 outperformed with a microF1 score of 0.7865 and an accuracy of 0.786.
UR - https://www.scopus.com/pages/publications/105034118629
U2 - 10.26615/978-954-452-098-4-163
DO - 10.26615/978-954-452-098-4-163
M3 - Conference contribution
AN - SCOPUS:105034118629
T3 - International Conference Recent Advances in Natural Language Processing, RANLP
SP - 1413
EP - 1419
BT - Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, RANLP 2025
A2 - Angelova, Galia
A2 - Kunilovskaya, Maria
A2 - Escribe, Marie
A2 - Mitkov, Ruslan
PB - Incoma Ltd
T2 - 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, RANLP 2025
Y2 - 8 September 2025 through 10 September 2025
ER -