A Review of the Evolution of Bias Detection Pathways and Fairness Evaluation Dilemmas in Large Language Models: From Surface-Level Auditing to Mechanism-Driven Contextualized Evaluation
DOI: https://doi.org/10.62517/jike.202604307
Author(s)
Bowen Deng
Affiliation(s)
Dongguan University of Technology, Dongguan, Guangdong, China
Abstract
As large language models (LLMs) become deeply integrated into various sectors of society, their potential algorithmic bias and unfairness have emerged as core risks constraining safe deployment. From the distinctive perspective of the “evolution of detection pathways,” this review systematically examines the transition of LLM bias evaluation from early surface-level output auditing to mechanism-driven contextualized evaluation. First, it develops a comprehensive taxonomy covering endogenous and exogenous bias, as well as representational and allocational harms, and compares the fundamental differences between LLMs and traditional NLP systems in how bias is manifested. Second, it analyzes five core detection pathways-surface-level output auditing, counterfactual consistency testing, internal mechanism analysis, attitude probing, and adversarial elicitation-in terms of their measurement dimensions, explanatory power, and limitations. Through analysis of CBBQ, TWBias, and cross-cultural benchmarks, this review identifies three major gaps in current evaluation infrastructure: cultural coverage, dynamic interaction, and harm linkage. The findings show that a significant “harm gap” exists between detection scores and real-world social harm, and that there is a deep circularity challenge between detection methods and definitions of bias. Finally, this review proposes future research directions, including the construction of a layered bias ontology, a tiered fairness evaluation system, and full-lifecycle monitoring, thereby providing theoretical support for building a more inclusive and culturally responsive AI governance framework.
Keywords
Large Language Models; Bias Detection; Fairness Evaluation; Cross-Cultural Context; Harm Gap; Mechanism Analysis; Contextualized Evaluation
References
[1] I. O. Gallegos et al., “Bias and Fairness in Large Language Models: A Survey,” Computational Linguistics, vol. 50, no. 3, pp. 1097–1179, Jan. 2024.
[2] Z. Chu, Z. Wang, and W. Zhang, “Fairness in Large Language Models: A Taxonomic Survey,” ACM SIGKDD Explorations Newsletter, vol. 26, no. 1, pp. 34–48, July 2024.
[3] R. Hida, M. Kaneko, and N. Okazaki, “Social Bias Evaluation for Large Language Models Requires Prompt Variations,” arXiv, July 2024.
[4] Y. Zhao et al., “Mind vs. Mouth: On Measuring Re-judge Inconsistency of Social Bias in Large Language Models,” arXiv, Aug. 2023.
[5] S. L. Blodgett, S. Barocas, H. Daumé, and H. Wallach, “Language (Technology) is Power: A Critical Survey of ‘Bias’ in NLP,” pp. 5454–5476, Jan. 2020.
[6] H. Cyberey, Y. Ji, and D. Evans, “Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?,” pp. 34–45, Jan. 2025.
[7] A. Neumann, E. Kirsten, M. B. Zafar, and J. Singh, “Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs),” pp. 573–598, June 2025,.
[8] R. Shelby et al., “Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction,” pp. 723–741, Aug. 2023.
[9] K. Chehbouni et al., “From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards,” pp. 15694–15710, Jan. 2024, doi: 10.18653/v1/2024.findings-acl.927.
[10] C. May, A. Wang, S. Bordia, S. R. Bowman, and R. Rudinger, “On Measuring Social Biases in Sentence Encoders,” pp. 622–628, Jan. 2019.
[11] A. Caliskan, J. J. Bryson, and A. Narayanan, “Semantics derived automatically from language corpora contain human-like biases,” Science, vol. 356, no. 6334, pp. 183–186, Apr. 2017.
[12] J. Dhamala et al., “BOLD,” pp. 862–872, Mar. 2021.
[13] M. Nadeem, A. Bethke, and S. Reddy, “StereoSet: Measuring stereotypical bias in pretrained language models,” pp. 5356–5371, Jan. 2021.
[14] N. Sivakumar et al., “Bias after Prompting: Persistent Discrimination in Large Language Models,” arXiv, Sept. 2025.
[15] X. Bai, A. Wang, I. Sucholutsky, and T. L. Griffiths, “Measuring Implicit Bias in Explicitly Unbiased Large Language Models,” arXiv, Feb. 2024.
[16] X. Yang, R. Zhan, D. F. Wong, S. A. Yang, J. Wu, and L. S. Chao, “Rethinking Prompt-based Debiasing in Large Language Models,” arXiv, Mar. 2025.
[17] H. Gonen and Y. Goldberg, “Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them,” arXiv, Mar. 2019.
[18] X. Lin, L. Li, and X. Liu, “Inconsistency Between Words and Actions: Implicit Bias in Decision-making with Large Language Models.” Aug. 11, 2025.
[19] A. Parrish et al., “BBQ: A hand-built bias benchmark for question answering,” in Findings of the Association for Computational Linguistics: ACL 2022 , Jan. 2022, pp. 2086–2105.
[20] Y. Huang and D. Xiong, “CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models,” arXiv, June 2023.
[21] H.-Y. Hsieh, S.-C. Huang, and R. T. Tsai, “TWBias: A Benchmark for Assessing Social Bias in Traditional Chinese Large Language Models through a Taiwan Cultural Lens,” pp. 8688–8704, Jan. 2024.
[22] T. Lan et al., “McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models,” pp. 6033–6056, Jan. 2025.
[23] Z. Xu, K. Peng, L. Ding, D. Tao, and X. Lu, “Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction,” arXiv, Mar. 2024.
[24] P. Röttger et al., “Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models,” pp. 15295–15311, Jan. 2024.
[25] P. Röttger, H. R. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy, “XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models,” pp. 5377–5400, Jan. 2024.
[26] J. Vig et al., “Investigating Gender Bias in Language Models Using Causal Mediation Analysis,” in Neural Information Processing Systems , Jan. 2020, pp. 12388–12401. Accessed: Oct. 2025. [Online]. Available: https://proceedings.neurips.cc/paper/2020/file/92650b2e92217715fe312e6fa7b90d82-Paper.pdf
[27] E. Karinshak et al., “LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output,” arXiv, Nov. 2024.
[28] Y. Wang et al., “CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models,” arXiv, Nov. 2023.
[29] E. Perez et al., “Red Teaming Language Models with Language Models,” pp. 3419–3448, Jan. 2022.
[30] J. Cui, W.-L. Chiang, I. Stoica, and C. Hsieh, “OR-Bench: An Over-Refusal Benchmark for Large Language Models,” arXiv, May 2024.
[31] O. van der Wal, D. Bachmann, A. Leidinger, L. van Maanen, W. Zuidema, and K. Schulz, “Undesirable Biases in NLP: Addressing Challenges of Measurement,” Journal of Artificial Intelligence Research, vol. 79, pp. 1–40, Jan. 2024.
[32] A. Z. Jacobs, S. L. Blodgett, S. Barocas, H. Daumé, and H. Wallach, “The meaning and measurement of bias,” pp. 706–706, Jan. 2020.
[33] P. Delobelle, G. Attanasio, D. Nozza, S. L. Blodgett, and Z. Talat, “Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLP,” pp. 21669–21691, Jan. 2024.
[34] A. Khan, S. Casper, and D. Hadfield-Menell, “Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs,” pp. 2151–2165, June 2025.
[35] J. R. Anthis, K. Lum, M. Ekstrand, A. Feller, A. D’Amour, and C. Tan, “The Impossibility of Fair LLMs,” arXiv, May 2024.
[36] Y. Tao, O. Viberg, R. S. Baker, and R. F. Kizilcec, “Cultural bias and cultural alignment of large language models,” PNAS Nexus, vol. 3, no. 9, Sept. 2024.
[37] L. Ouyang et al., “Training language models to follow instructions with human feedback,” arXiv, Mar. 2022.
[38] M. Dabas, S. Chen, C. A. Fleming, M. Jin, and R. Jia, “Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning,” arXiv, July 2025, Accessed: Oct. 2025. [Online]. Available: https://arxiv.org/abs/2507.04250
[39] S. Dev et al., “On Measures of Biases and Harms in NLP,” pp. 246–267, Jan. 2022.