Skip to main navigation Skip to search Skip to main content

ImageEval 2025: The first arabic image captioning shared task

  • Birzeit University
  • American University of Beirut
  • Arab Center for Research and Policy Studies
  • Northwestern University in Qatar

Research output: Contribution to conferencePaperpeer-review

Abstract

We present ImageEval 2025, the first shared task dedicated to Arabic image captioning. The task addresses the critical gap in multimodal Arabic NLP by focusing on two complementary subtasks: (1) creating the first open-source, manually-captioned Arabic image dataset through a collaborative datathon, and (2) developing and evaluating Arabic image captioning models. A total of 44 teams registered, of which eight submitted during the test phase, producing 111 valid submissions. Evaluation was conducted using automatic metrics, LLM-based judgment, and human assessment. In Subtask 1, the best-performing system achieved a cosine similarity of 65.5, while in Subtask 2, the top score was 60.0. Although these results show encouraging progress, they also confirm that Arabic image captioning remains a challenging task, particularly due to cultural grounding requirements, morphological richness, and dialectal variation. All datasets, baseline models, and evaluation tools are released publicly to support future research in Arabic multimodal NLP.
Original languageEnglish
Pages379-389
Number of pages11
DOIs
Publication statusPublished - Nov 2025
Event Third Arabic Natural Language Processing Conference: Shared Tasks - Suzhou, China
Duration: 8 Nov 20259 Nov 2025

Conference

Conference Third Arabic Natural Language Processing Conference: Shared Tasks
Country/TerritoryChina
CitySuzhou
Period8/11/259/11/25

Fingerprint

Dive into the research topics of 'ImageEval 2025: The first arabic image captioning shared task'. Together they form a unique fingerprint.

Cite this