Skip to main navigation Skip to search Skip to main content

An Annotated Corpus of Arabic Tweets for Hate Speech Analysis

  • Wajdi Zaghoauni
  • , Md Rafiul Biswas
  • Northwestern University in Qatar

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10,000 Arabic tweets and annotated each tweet, whether it contains offensive content or not. If a text contains offensive content, we further classify it into different hate speech targets such as religion, gender, politics, ethnicity, origin, and others. A text can contain either single or multiple targets. Multiple annotators are involved in the data annotation task. We calculated the inter-annotator agreement, which was reported to be 0.86 for offensive content and 0.71 for multiple hate speech targets. Finally, we evaluated the data annotation task by employing a different transformers-based model in which AraBERTv2 outperformed with a microF1 score of 0.7865 and an accuracy of 0.786.

Original languageEnglish
Title of host publicationProceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, RANLP 2025
EditorsGalia Angelova, Maria Kunilovskaya, Marie Escribe, Ruslan Mitkov
PublisherIncoma Ltd
Pages1413-1419
Number of pages7
ISBN (Electronic)9789544520984
DOIs
Publication statusPublished - 10 Sept 2025
Event15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, RANLP 2025 - Varna, Bulgaria
Duration: 8 Sept 202510 Sept 2025

Publication series

NameInternational Conference Recent Advances in Natural Language Processing, RANLP
ISSN (Print)1313-8502

Conference

Conference15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, RANLP 2025
Country/TerritoryBulgaria
CityVarna
Period8/09/2510/09/25

Fingerprint

Dive into the research topics of 'An Annotated Corpus of Arabic Tweets for Hate Speech Analysis'. Together they form a unique fingerprint.

Cite this