Skip to main navigation Skip to search Skip to main content

Offensive Hebrew Corpus and Detection using BERT

  • Nagham Hamad*
  • , Mustafa Jarrar
  • , Mohammad Khalilia
  • , Nadim Nashif
  • *Corresponding author for this work
  • Birzeit University
  • 7amleh Center

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Offensive language detection has been well studied in many languages, but it is lagging behind in low-resource languages, such as Hebrew. In this paper, we present a new offensive language corpus in Hebrew. A total of 15,881 tweets were retrieved from Twitter. Each was labeled with one or more of five classes (abusive, hate, violence, pornographic, or none offensive) by Arabic-Hebrew bilingual speakers. The annotation process was challenging as each annotator is expected to be familiar with the Israeli culture, politics, and practices to understand the context of each tweet. We fine-tuned two Hebrew BERT models, HeBERT and AlephBERT, using our proposed dataset and another published dataset. We observed that our data boosts HeBERT performance by 2% when combined with DOLaH. Fine-tuning AlephBERT on our data and testing on DOLaH yields 69% accuracy, while fine-tuning on DOLaH and testing on our data yields 57% accuracy, which may be an indication to the generalizability our data offers. Our dataset and fine-tuned models are available on GitHub and Huggingface.

Original languageEnglish
Title of host publication2023 20th Acs/ieee International Conference On Computer Systems And Applications, Aiccsa
PublisherIEEE Computer Society
Number of pages8
ISBN (Electronic)9798350319439
DOIs
Publication statusPublished - 7 Dec 2023
Externally publishedYes
Event20th ACS/IEEE International Conference on Computer Systems and Applications, AICCSA 2023 - Giza, Egypt
Duration: 4 Dec 20237 Dec 2023

Publication series

NameInternational Conference On Computer Systems And Applications

Conference

Conference20th ACS/IEEE International Conference on Computer Systems and Applications, AICCSA 2023
Country/TerritoryEgypt
CityGiza
Period4/12/237/12/23

Keywords

  • Deep Learning
  • Hate speech
  • Hebrew
  • Offensive
  • Pre-trained model

Fingerprint

Dive into the research topics of 'Offensive Hebrew Corpus and Detection using BERT'. Together they form a unique fingerprint.

Cite this