ON
← Back to feed
AI is more likely than humans to form biases when hiring
United States🏛️ PoliticsCenteryesterday

AI is more likely than humans to form biases when hiring

A recent study reveals that large language models (LLMs), such as ChatGPT, Claude, and Gemini, are more prone to forming stereotypes and biases during simulated hiring processes compared to humans. In experiments conducted by researchers from Princeton University and the University of Chicago, these AI systems were tasked with hiring candidates for various roles based on performance feedback. Despite all candidates having equal chances of success, the models began assigning individuals to specific jobs based on early outcomes, leading to segregation by fictional ethnic groups. This tendency to generalize from limited data results in stronger stereotyping than observed in human participants. The findings highlight concerns about AI's potential to reinforce biases in hiring practices.

Artificial intelligence systems used in hiring processes are more prone to developing biases than humans, according to a recent study conducted by researchers from Princeton University and the University of Chicago. The findings suggest that large language models (LLMs), such as ChatGPT, Claude, and Gemini, may inadvertently reinforce stereotypes against job applicants based on limited data, potentially leading to unfair hiring practices. In the study, researchers simulated a hiring scenario where LLMs were tasked with selecting candidates for various roles, including doctors, lawyers, child-care aides, and janitors. The candidates belonged to four fictional ethnic groups: Tufa, Aima, Reku, and Weki. All candidates were equally capable of performing any given job, but the models were not informed of this fact. As the models made hiring decisions, they began associating specific ethnic groups with certain roles based on initial outcomes. If a candidate from the Aima group failed in a role deemed to require high levels of warmth and competence, such as a doctor, the model would avoid hiring other Aimas for similar positions. Instead, it would assign them to roles perceived as requiring fewer skills, such as janitorial work. The results revealed that LLMs exhibited significantly greater stereotyping tendencies compared to human participants in a prior psychological study. Human participants in the original study scored 0.84 on a segregation scale, whereas the models reached scores nearly 65% higher. OpenAI's reasoning model o3 achieved a score of 1.83, approaching the maximum possible value. Ryan Liu, a PhD student at Princeton University and a co-author of the study, explained that LLMs are designed to draw conclusions from minimal data, which can lead to premature generalizations. This tendency, while beneficial in solving logic puzzles, can result in biased decision-making in social contexts. The study highlights a common challenge faced by both human and machine decision-makers known as the "exploration-exploitation dilemma." This refers to the balance between relying on past successes and exploring new possibilities. However, LLMs, due to their training on tasks that emphasize generalization from limited examples, may become overly confident in their assumptions too quickly. Newer models with enhanced reasoning abilities, such as OpenAI’s o3 and DeepSeek’s R1, demonstrated even stronger biases. According to Liu, the rapid formation of stereotypes by these models can lead to problematic outcomes, particularly as chatbots continue to evolve with improved memory and personalization features. Angelina Wang, a computer scientist at Cornell University who was not involved in the study, noted that while chatbots need to retain information from previous interactions, this feature can contribute to the formation of biases. She emphasized that determining the appropriate level of memory retention remains an ongoing challenge. Despite efforts to instruct models to prioritize fairness, the study found that such directives had little impact on the models' behavior. The implications of these findings are particularly pertinent as AI systems increasingly take on roles traditionally held by humans, raising concerns about the potential for systemic bias in automated hiring processes. Further research and development are needed to address these issues effectively, ensuring that AI technologies support equitable opportunities rather than perpetuating existing inequalities.

How each side covered it

The same event, grouped by the political lean of the outlets covering it.

How each side covered it

Support independent, bias-aware news and unlock the social pulse, community voting, and your personalized For You feed.

Become a Supporter

Covered around the world

The same event as reported in other countries.

Covered around the world

Support independent, bias-aware news and unlock the social pulse, community voting, and your personalized For You feed.

Become a Supporter

Claims check

Key factual claims, and how many sources assert vs dispute each.

Claims check

Support independent, bias-aware news and unlock the social pulse, community voting, and your personalized For You feed.

Become a Supporter

Go to the primary sources (2)

The official sources this coverage is built on. Read them directly to bypass framing.

1 reports

MIT Technology Review logoMIT Technology ReviewIndependentCenterFactual 85Objective 78yesterday
AI is more likely than humans to form biases when hiring

A recent study reveals that large language models (LLMs), such as ChatGPT, Claude, and Gemini, are more prone to forming stereotypes and biases during simulated hiring processes compared to humans. In experiments conducted by researchers from Princeton University and the University of Chicago, these AI systems were tasked with hiring candidates for various roles based on performance feedback. Despite all candidates having equal chances of success, the models began assigning individuals to specific jobs based on early outcomes, leading to segregation by fictional ethnic groups. This tendency to generalize from limited data results in stronger stereotyping than observed in human participants. The findings highlight concerns about AI's potential to reinforce biases in hiring practices.

Bias read (Center): The article presents empirical findings from academic research without overt ideological framing. It discusses AI bias in hiring but does not advocate for any particular political stance or agenda. The focus is on technical and psychological aspects rather than partisan issues.

Why factuality (85): The article presents findings from a study conducted by researchers at Princeton University and the University of Chicago, where LLMs like ChatGPT, Claude, and Gemini were tested in a simulated hiring scenario. It accurately describes the methodology and results, noting that models developed biases

Why objectivity (78): The article remains largely neutral in tone, presenting the study's findings without overt bias. However, it uses emotionally charged language such as 'handing them ammunition for forming those biases' which subtly implies negative consequences of AI hiring systems. This slight editorializing reduce

Keep the news honest.

ObjectiveNews is reader-funded and ad-free — we show you the bias instead of hiding it. Support independent journalism for €5/month.

Become a Supporter

Related stories