As large language models become embedded in content moderation, decision-support,
and governance applications, evaluating whether their outputs align with established
human rights principles has become a pressing methodological challenge. This thesis
investigates the reliability of prompt-based LLM judges, operating within the LLM-as
a-Judge framework, when tasked with evaluating AI-generated responses for alignment
with the Universal Declaration of Human Rights (UDHR).
While prior research on LLM-based evaluation has focused on helpfulness, preference
ranking, and safety, structured evaluation against codified normative frameworks
remains largely unexplored. The effects of prompt design and model selection on
normative classification outcomes are not well understood, and no systematic
comparison of LLM judges has been conducted in this domain.
The study introduces a three-condition experimental design comparing a minimal zero
shot prompt, a UDHR-grounded detailed zero-shot prompt, and a few-shot
configuration with four calibration examples. Four judge models, namely ChatGPT,
Grok, Gemini, and Fanar, evaluate the same dataset of human rights related question
answer pairs under all three conditions. The study also introduces the Alignment
Leniency Score (ALS), a metric for quantifying the directional tendency of a binary
judge when resolving normatively ambiguous cases.
The findings show that capable LLM judges can reproduce human alignment labels
with substantial consistency even under minimal prompting. Few-shot conditioning
produces a more meaningful recalibration of classification thresholds than detailed
instruction alone, particularly for borderline cases. Critically, the relative ordering of
model strictness and leniency is preserved across all configurations, indicating that
threshold orientation under ambiguity is a stable, model-level property that prompt
design can moderate but not override.
The thesis concludes that automated human rights alignment evaluation is feasible and
reproducible at scale, but that its outputs reflect model-conditioned thresholds rather
than objective normative verdicts. Model selection and prompt configuration are
substantive methodological decisions with direct consequences for how evaluation
results should be interpreted in governance-sensitive contexts.
| Date of Award | 2026 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - HBKU College of Humanities and Social Science
|
|---|
- AI Alignment
- AI Ethics
- Few-Shot
- Human Rights
- LLM-As-A-Judge
- Zero-Shot
LLM-AS-A-JUDGE FRAMEWORK FOR HUMAN RIGHTS ALIGNMENT
Morssy, N. (Author). 2026
Student thesis: Master's Dissertation