Skip to main navigation Skip to search Skip to main content

LLM-AS-A-JUDGE FRAMEWORK FOR HUMAN RIGHTS ALIGNMENT

  • Noreldin Morssy

Student thesis: Master's Dissertation

Abstract

As large language models become embedded in content moderation, decision-support, and governance applications, evaluating whether their outputs align with established human rights principles has become a pressing methodological challenge. This thesis investigates the reliability of prompt-based LLM judges, operating within the LLM-as a-Judge framework, when tasked with evaluating AI-generated responses for alignment with the Universal Declaration of Human Rights (UDHR). While prior research on LLM-based evaluation has focused on helpfulness, preference ranking, and safety, structured evaluation against codified normative frameworks remains largely unexplored. The effects of prompt design and model selection on normative classification outcomes are not well understood, and no systematic comparison of LLM judges has been conducted in this domain. The study introduces a three-condition experimental design comparing a minimal zero shot prompt, a UDHR-grounded detailed zero-shot prompt, and a few-shot configuration with four calibration examples. Four judge models, namely ChatGPT, Grok, Gemini, and Fanar, evaluate the same dataset of human rights related question answer pairs under all three conditions. The study also introduces the Alignment Leniency Score (ALS), a metric for quantifying the directional tendency of a binary judge when resolving normatively ambiguous cases. The findings show that capable LLM judges can reproduce human alignment labels with substantial consistency even under minimal prompting. Few-shot conditioning produces a more meaningful recalibration of classification thresholds than detailed instruction alone, particularly for borderline cases. Critically, the relative ordering of model strictness and leniency is preserved across all configurations, indicating that threshold orientation under ambiguity is a stable, model-level property that prompt design can moderate but not override. The thesis concludes that automated human rights alignment evaluation is feasible and reproducible at scale, but that its outputs reflect model-conditioned thresholds rather than objective normative verdicts. Model selection and prompt configuration are substantive methodological decisions with direct consequences for how evaluation results should be interpreted in governance-sensitive contexts.
Date of Award2026
Original languageAmerican English
Awarding Institution
  • HBKU College of Humanities and Social Science

Keywords

  • AI Alignment
  • AI Ethics
  • Few-Shot
  • Human Rights
  • LLM-As-A-Judge
  • Zero-Shot

Cite this

'