TY - GEN
T1 - AbjadAuthorID
T2 - 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026
AU - Abudalfa, Shadi
AU - Ezzini, Saad
AU - Abdelali, Ahmed
AU - Jarrar, Mustafa
AU - El-Haj, Mo
AU - Durrani, Nadir
AU - Sajjad, Hassan
AU - Adeeba, Farah
AU - Ahmadi, Sina
N1 - Publisher Copyright:
© 2026 Association for Computational Linguistics.
PY - 2026
Y1 - 2026
N2 - Authorship identification is a core problem in Natural Language Processing and computational linguistics, with applications spanning digital humanities, literary analysis, and forensic linguistics. While substantial progress has been made for English and other high-resource languages, authorship attribution for languages written in the Arabic (Abjad) script remains underexplored. In this paper, we present an overview of AbjadAuthorID, a shared task organised as part of the AbjadNLP workshop at EACL 2026, which focuses on multiclass authorship identification across Arabic-script languages. The shared task covers Modern Standard Arabic, Urdu, and Kurdish, and is formulated as a closed-set multiclass classification problem over literary text spanning multiple authors and historical periods. We describe the task motivation, dataset construction, evaluation protocol, and participation statistics, and report official results for the Arabic track. The findings highlight both the effectiveness of current approaches in controlled settings and the challenges posed by lower participation and resource availability in some language tracks. AbjadAuthorID establishes a new benchmark for multilingual authorship attribution in morphologically rich, underrepresented languages.
AB - Authorship identification is a core problem in Natural Language Processing and computational linguistics, with applications spanning digital humanities, literary analysis, and forensic linguistics. While substantial progress has been made for English and other high-resource languages, authorship attribution for languages written in the Arabic (Abjad) script remains underexplored. In this paper, we present an overview of AbjadAuthorID, a shared task organised as part of the AbjadNLP workshop at EACL 2026, which focuses on multiclass authorship identification across Arabic-script languages. The shared task covers Modern Standard Arabic, Urdu, and Kurdish, and is formulated as a closed-set multiclass classification problem over literary text spanning multiple authors and historical periods. We describe the task motivation, dataset construction, evaluation protocol, and participation statistics, and report official results for the Arabic track. The findings highlight both the effectiveness of current approaches in controlled settings and the challenges posed by lower participation and resource availability in some language tracks. AbjadAuthorID establishes a new benchmark for multilingual authorship attribution in morphologically rich, underrepresented languages.
UR - https://www.scopus.com/pages/publications/105040702400
U2 - 10.18653/v1/2026.abjadnlp-1.69
DO - 10.18653/v1/2026.abjadnlp-1.69
M3 - Conference contribution
AN - SCOPUS:105040702400
T3 - EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026
SP - 538
EP - 544
BT - EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script, AbjadNLP 2026
A2 - El-Haj, Mo
A2 - El-Haj, Mo
A2 - Rayson, Paul
A2 - Jarrar, Mustafa
A2 - Ezeani, Ignatius
A2 - Ezzini, Saad
A2 - Ahmadi, Sina
A2 - Haddad, Amal Haddad
A2 - Amol, Cynthia
A2 - Abdelali, Ahmad
A2 - Abudalfa, Shadi
PB - Association for Computational Linguistics (ACL)
Y2 - 28 March 2026
ER -