Skip to main navigation Skip to search Skip to main content

Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs

  • Abdellah El Mekki
  • , Samar M. Magdy
  • , Houdaifa Atou
  • , Ruwa AbuHweidi
  • , Baraah Qawasmeh
  • , Omer Nacar
  • , Thikra Al-hibiri
  • , Razan Saadie
  • , Hamzah Alsayadi
  • , Nadia Ghezaiel Hammouda
  • , Alshima Alkhazimi
  • , Aya Hamod
  • , Al-Yas Al-Ghafri
  • , Wesam El-Sayed
  • , Asila Al sharji
  • , Mohamad Ballout
  • , Anas Belfathi
  • , Karim Ghaddar
  • , Serry Sibaee
  • , Alaa Aoun
  • Areej Asiri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Alkathiri, Fadi A Zaraket, Mustafa Jarrar, Yahya Mohamed El Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed
  • University of British Columbia
  • Mohammed VI Polytechnic University
  • Birzeit University
  • Western Michigan University
  • Tuwaiq Academy
  • King Khalid University
  • American University of Beirut
  • Ibb University
  • University of Hail
  • University of Technology and Applied Sciences
  • Arab Open University
  • Minia University
  • Nantes Université
  • Prince Sultan University (PSU)
  • Umm Al-Qura University
  • University of Nouakchott
  • Fatabyyano
  • Independent Researcher
  • Hadhramout University of Science and Technology
  • Cairo University
  • Misurata University
  • Al-Balqa Applied University
  • University of Khartoum
  • Sultan Qaboos University
  • Arab Center for Research and Policy Studies
  • Institut supérieur de l'électronique et du numérique
  • Canada Research Chair (CRC)

Research output: Contribution to conferencePaperpeer-review

Abstract

Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than Modern Standard Arabic (MSA). Despite this, machine translation (MT) systems often generalize poorly to dialectal input, limiting their utility for millions of speakers. We introduce Alexandria, a large-scale, community-driven, human-translated dataset designed to bridge this gap. Alexandria covers 13 Arab countries and 11 high-impact domains, including health, education, and agriculture. Unlike previous resources, Alexandria provides unprecedented granularity by associating contributions with city-of-origin metadata, capturing authentic local varieties beyond coarse regional labels. The dataset consists of parallel English-Dialectal Arabic multi-turn conversational scenarios annotated with speaker-addressee gender configurations, enabling the study of gender-conditioned variation in dialectal use. Comprising 107K total turns, Alexandria serves as both a training resource and as a rigorous benchmark for evaluating MT and Large Language Models (LLMs). Our automatic and human evaluation benchmarks the current capabilities of Arabic-aware LLMs in translating across diverse Arabic dialects and sub-dialects while exposing significant persistent challenges.The Alexandria dataset, the creation prompts, the translation and revision guidelines, and the evaluation code are publicly available in the following repository: https://github.com/UBC-NLP/Alexandria
Original languageEnglish
Pages32567-32592
Number of pages25
DOIs
Publication statusPublished - Jul 2026
EventProceedings of the 64th Annual Meeting of the Association for Computational Linguistics - California, United States
Duration: 2 Jul 20267 Jul 2026

Conference

ConferenceProceedings of the 64th Annual Meeting of the Association for Computational Linguistics
Country/TerritoryUnited States
CityCalifornia
Period2/07/267/07/26

Fingerprint

Dive into the research topics of 'Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs'. Together they form a unique fingerprint.

Cite this