Skip to main navigation Skip to search Skip to main content

A TRAIN-ONCE, EVALUATE-TWICE PROTOCOL FOR ARABIC-ENGLISH CROSS-LINGUAL DEPRESSION DETECTION WITH LARGE LANGUAGE MODELS

  • Amira Abdalla

Student thesis: Master's Dissertation

Abstract

More than 1B people across the globe are living with mental disorder and depression is among one of them. Considering the digital footprint of mass population in social media, their thoughts, emotions and psychological states can be captured form their social media post. Large language models (LLMs) based digital solutions have been explored leveraging social media posts, but their cross-lingual capability for under-representative language are not heavily explored. In this work, we used "train once, evaluate twice" protocol, to evaluate cross-lingual robustness of eight LLMs in depression detection for English and Arabic language. Six open-source LLMs were adapted using parameter-efficient fine-tuning (PEFT) with QLoRA, and two closed-source API-based models that were fine-tuned via supervised fine-tuning (SFT). In monolingual evaluation, the LLMs perform at near-ceiling levels in English (best score: 0.991 on GPT-3.5 Turbo) and at moderate levels in Arabic (best score: 0.800 on GPT-3.5 Turbo and Fanar-1). The cross-lingual evaluation shows significant variability. The strongest cross-lingual result occurs in AR→EN (Macro-F1 = 0.810 for GPT-4.1 mini), whereas EN→AR remains most challenging (best Macro-F1 = 0.657 for GPT-4.1 mini). The error analysis indicates that EN→AR fails with high false-negative rates, which is critical in clinical practice. At the same time, the AR→EN tends to increase false positives. Interestingly, Arabic-focused LLMs do not outperform general-purpose LLMs in cross-lingual transfer. Measuring the efficiency of PEFT shows that compact adaptation (adapters: 124.6–206.2 MB) with low latency (21-35 ms/sample) and that API-based inference is much slower (762-803 ms/sample by GPT-4.1 mini). Overall, we conclude that existing fine-tuning methods exhibit poor cross-lingual transfer, highlighting the need for robustness of the LLMs and new adaptation methods to provide a fair representation of the diverse linguistic populations in this world.
Date of Award2026
Original languageAmerican English
Awarding Institution
  • HBKU College of Science and Engineering

Keywords

  • Arabic NLP
  • Cross-lingual depression detection
  • Large language models (LLMs)
  • Mental health AI
  • Parameter-efficient fine-tuning (PEFT)

Cite this

'