Skip to main navigation Skip to search Skip to main content

Training-Free Cultural Alignment of Large Language Models via Persona Disagreement

  • Huynh Trung Kiet
  • , Dao Sy Duy Minh
  • , Tuan Nguyen
  • , Chi-Nguyen Tran
  • , Phu-Hoa Pham
  • , Nguyen Lam Phu Quy
  • , The Anh Han
  • , Long Tran-Thanh

Research output: Working paperPreprint

3 Downloads (Pure)

Abstract

Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preferences are not culturally neutral. Existing cultural alignment methods either require per-country preference data and fine-tuning budgets or assume white-box access to model internals that commercial APIs do not expose. In this work, we focus on this realistic black-box, public-data-only regime and observe that within-country sociodemographic disagreement, not consensus, is the primary steering signal. We introduce DISCA (Disagreement-Informed Steering for Cultural Alignment), an inference-time method that instantiates each country as a panel of World-Values-Survey-grounded persona agents and converts their disagreement into a bounded, loss-averse logit correction. Across 20 countries and 7 open-weight backbones (2B--70B), DISCA reduces cultural misalignment on MultiTP by 10--24% on the six backbones >=3.8B, and 2--7% on open-ended scenarios, without changing any weights. Our results suggest that inference-time calibration is a scalable alternative to fine-tuning for serving the long tail of global moral preferences.
Original languageEnglish
PublisherarXiv
Number of pages48
Publication statusPublished - 11 May 2026

Bibliographical note

57 pages, 1 figure, 6 MultiTP moral dimensions
Submitted on 11 May 2026 (v1), last revised 18 May 2026 (this version, v2)

Fingerprint

Dive into the research topics of 'Training-Free Cultural Alignment of Large Language Models via Persona Disagreement'. Together they form a unique fingerprint.

Cite this