Skip to main navigation Skip to search Skip to main content

Detecting Subtle Biases: An Ethical Lens on Underexplored Areas in AI Language Models Biases

  • King's College London
  • Amirkabir University of Technology
  • Rutgers University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large Language Models (LLMs) are increasingly embedded in the daily lives of individuals across diverse social classes. This widespread integration raises urgent concerns about the subtle, implicit biases these models may contain. In this work, we investigate such biases through the lens of ethical reasoning, analyzing model responses to scenarios in a new dataset we propose comprising 1,016 scenarios, systematically categorized into ethical, unethical, and neutral types. Our study focuses on dimensions that are socially influential but less explored, including (i) residency status, (ii) political ideology, (iii) Fitness Status, (iv) educational attainment, and (v) attitudes toward AI. To assess LLMs’ behavior, we propose a baseline and employ one statistical test and one metric: a permutation test that reveals the presence of bias by comparing the probability distributions of ethical/unethical scenarios with the probability distribution of neutral scenarios on each demographic group, and a tendency measurement that captures the magnitude of bias with respect to the relative difference between probability distribution of ethical and unethical scenarios. Our evaluations of 12 prominent LLMs reveal persistent and nuanced biases across all five attributes, and Llama models exhibited the most pronounced biases. These findings highlight the need for refined ethical benchmarks and bias-mitigation tools in LLMs.

Original languageEnglish
Title of host publicationLong Papers
EditorsVera Demberg, Kentaro Inui, Lluis Marquez Villodre
PublisherAssociation for Computational Linguistics (ACL)
Pages7352-7379
Number of pages28
ISBN (Electronic)9798891763807
DOIs
Publication statusPublished - 2026
Event19th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2026 - Rabat, Morocco
Duration: 24 Mar 202629 Mar 2026

Publication series

NameEACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, Vol. 1 - (Long Papers)
Volume1

Conference

Conference19th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2026
Country/TerritoryMorocco
CityRabat
Period24/03/2629/03/26

Fingerprint

Dive into the research topics of 'Detecting Subtle Biases: An Ethical Lens on Underexplored Areas in AI Language Models Biases'. Together they form a unique fingerprint.

Cite this