Skip to main navigation Skip to search Skip to main content

Do Large Language Models reflect demographic pluralism in safety?

Usman Naseem*, Gautam Siddharth Kashyap, Sushant Kumar Ray, Rafiq Ali, Ebad Shabbir, Abdullah Mohammad

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference proceeding contributionpeer-review

Abstract

Large Language Model (LLM) safety is inherently pluralistic, reflecting variations in moral norms, cultural expectations, and demographic contexts. Yet, existing alignment datasets such as ANTHROPIC-HH and DICES rely on demographically narrow annotator pools, overlooking variation in safety perception across communities. Demo-SafetyBench1 addresses this gap by modeling demographic pluralism directly at the prompt level, decoupling value framing from responses. In Stage I, prompts from DICES are reclassified into 14 safety domains (adapted from BEAVERTAILS) using Mistral-7B-Instruct-v0.3, retaining demographic metadata and expanding low-resource domains via Llama-3.1-8BInstruct with SimHash-based deduplication, yielding 43,050 samples. In Stage II, pluralistic sensitivity is evaluated using LLMs-as-Raters—Gemma-7B, GPT-4o, and LLaMA-2-7B—under zero-shot inference. Balanced thresholds (δ=0.5, τ=10) achieve high reliability (ICC = 0.87) and low demographic sensitivity (DS = 0.12), confirming that pluralistic safety evaluation can be both scalable and demographically robust.

Original languageEnglish
Title of host publicationFindings of the Association for Computational Linguistics
Subtitle of host publicationEACL 2026
EditorsVera Demberg, Kentaro Inui, Lluís Marquez
Place of PublicationKerrville, TX
PublisherAssociation for Computational Linguistics (ACL)
Pages2042-2052
Number of pages11
ISBN (Electronic)9798891763869
DOIs
Publication statusPublished - 2026
Event19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026 - Rabat, Morocco
Duration: 24 Mar 202629 Mar 2026

Conference

Conference19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
Country/TerritoryMorocco
CityRabat
Period24/03/2629/03/26

Fingerprint

Dive into the research topics of 'Do Large Language Models reflect demographic pluralism in safety?'. Together they form a unique fingerprint.

Cite this