Cultural heritage institutions face a growing metadata bottleneck: digitization has scaled
rapidly, but expert description has not. This thesis investigates AI-based image captioning for
the photographic collection of the National Museum of Qatar (NMoQ), addressing three
research questions: the impact of prompt configuration on caption quality (RQ1), the influence
of model training background on cultural awareness (RQ2), and the validity of LLM-as-judge
evaluation in this cultural context (RQ3). Five prompts were tested on a 10-image pilot using
GPT-4o mini, and the best-performing prompt was applied across three models: GPT-4o mini,
Fanar Oryx-IVU-2, and Qwen 2.5-7B, generating 1,350 captions from a 450-image dataset. A
150-caption reference subset was manually evaluated, and the full dataset was scored using
Gemini 2.5 Pro as an automated judge calibrated against expert scores. Structured prompt
engineering raised Cultural Awareness scores from 2.70 to 4.10 without fine-tuning. GPT-4o
mini outperformed both regional models across all evaluation methods. Fanar's Gulf-specific
training produced no statistically significant advantage in cultural awareness over Qwen 2.5
under expert evaluation (p = 1.000), challenging the assumption that regional training
translates into improved cultural captioning in English-language tasks. The Geminii judge
achieved moderate post-refinement agreement (Pearson r = 0.481) but exhibited a systematic
leniency bias toward Fanar that rubric instructions alone could not resolve. The study
contributes an evidence-based prompting framework for Gulf heritage captioning, empirical
evidence against the assumed value of regional model training for English-language cultural
tasks, and a replicable LLM-as-judge calibration methodology for non-Western heritage
contexts.
| Date of Award | 2026 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - HBKU College of Humanities and Social Science
|
|---|
- AI image captioning
- Cultural heritage documentation
- Gulf Arab cultural bias
- LLM-as-judge evaluation
- Prompt engineering
- Vision-language models
Automating Cultural Memory: AI Image Captioning for The National Museum of Qatar's Photographic Collection
Saleh, M. (Author). 2026
Student thesis: Master's Dissertation