Skip to main navigation Skip to search Skip to main content

A MORPHOLOGICALLY ANNOTATED CORPUS OF JORDANIAN ARABIC FOR LINGUISTIC ANALYSIS AND DIALECT-AWARE NLP

  • Heyam Salman

Student thesis: Master's Dissertation

Abstract

This thesis presents a morphologically annotated corpus of Jordanian Arabic. A dataset of approximately 36,000 words was collected from publicly available digital sources, including social media posts, blog entries, YouTube comments, and transcripts of Jordanian television series, spanning ten thematic domains to capture naturally occurring dialectal usage across diverse communicative contexts. From this dataset, 17,448 tokens were fully manually annotated with morphological features including part-of-speech tags, normalized forms, prefix, and suffix segmentation, and both dialectal and MSA lemmas, while the remaining tokens received automatic annotations using the ALMA morphological analyzer. All annotations were conducted through the Tawseem annotation portal, following guidelines developed for existing Levantine Arabic corpora to ensure consistency and compatibility with related dialect resources. Inter-annotator agreement reached κ = 0.928 across annotation categories, demonstrating the high quality of the annotation framework. Analysis of the manually annotated data reveals strong morphological similarity between Jordanian and Palestinian Arabic, reinforcing their shared classification within the Southern Levantine dialect continuum. The corpus contributes a new dedicated linguistic resource for an underrepresented Arabic variety and is designed to support future research in Arabic dialectology and dialect-aware natural language processing.
Date of Award2026
Original languageAmerican English
Awarding Institution
  • HBKU College of Humanities and Social Science

Keywords

  • None

Cite this

'