Brief IA

Recruitment: AI Amplifies Human Biases

🔬 Research·Tom Levy·

Recruitment: AI Amplifies Human Biases

Recruitment: AI Amplifies Human Biases
Key Takeaways
1Researchers have found that AI language models, like ChatGPT, can develop their own biases during recruitment.
2In a simulation, the AIs stereotyped candidates more than humans, influenced by fictitious hiring outcomes.
3The AI models, optimized for generalization, exhibited stronger biases than the human participants in the study.
💡Why it mattersThe increasing use of AI in recruitment could reinforce unfair stereotypes, impacting fairness in the job market.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

AI and Bias in Recruitment

When a candidate submits their resume for a job, it is increasingly likely that an artificial intelligence will be the first to review it. However, this automation raises questions about the fairness of these evaluations. Language models, or LLMs, are known to incorporate human biases present in the data they are trained on. Recent research indicates that these models can also develop their own biases based on their experiences, stereotyping candidates more than human recruiters would.

Researchers from Princeton and the University of Chicago conducted a study using language models such as ChatGPT, Claude, and Gemini in a recruitment simulation. Inspired by a psychological study, this simulation aimed to understand how humans form stereotypes. The models were informed that they were acting as consultants for the mayor of a fictional city, tasked with recruiting for 20 types of jobs, ranging from doctors to janitors. The candidates belonged to four fictional ethnic groups: Tufa, Aima, Reku, and Weki.

In each round of the simulation, a new job was proposed with four candidates, each from a different ethnic group. After each hire, the model learned whether the candidate had succeeded in their role before moving on to the next round. The goal was to maximize the number of successful hires over 40 rounds. The models were unaware that all candidates had equal chances of success in each job.

Quickly, the models began to differentiate candidates based on previous hiring outcomes. For example, if a model found that a candidate from the Aima group failed as a doctor, it would subsequently avoid hiring Aimas for that position, preferring to recruit them for jobs deemed less demanding, such as janitorial roles.

The models proved to be more inclined to stereotype candidates based on their demographic group than the human participants in the original study. On a segregation scale where 2 indicates total segregation, humans scored 0.84, while the models scored approximately 65% higher. OpenAI's model, o3, achieved a score of 1.83, close to the maximum possible.

This trend is explained by the LLMs' tendency to generalize from limited data, explains Ryan Liu, a PhD student at Princeton and co-author of the study presented at ICML in Seoul. LLMs are designed to optimize generalization, making them quick to form stereotypes. In the experiment, newer models, such as OpenAI's o3 and DeepSeek's R1, exhibited even more pronounced biases. When these models generalize too quickly in social contexts, the outcomes can be problematic.

This finding is particularly relevant as chatbots gain memory and personalization, emphasizes Angelina Wang, a computer scientist at Cornell University. A chatbot that relies on its conversation history can reinforce biased behaviors already encountered. Reducing chatbot memory is not a solution, as users want them to remember their interactions. Finding the right balance remains a challenge.

Asking the models to be fair did not significantly change their behavior. According to Liu, either the models cannot apply these values, or the optimization of correct hires takes precedence. However, promising rewards for diverse hires has reduced biases. The challenge, therefore, is to design objectives that incorporate desirable social values.

The models also showed less bias when they had more personal information about the candidates. In another experiment, the models were tasked with reinstating members of different ethnic groups in Canada. With relevant information such as age and education, the models stereotyped candidates less. In contrast, with irrelevant information, such as hair color, biases reemerged.

The question of how much AI systems will stereotype candidates in the real world remains open. In the experiment, the models received immediate feedback on their hires, which is not the case in reality. Companies may take time to evaluate the competence of a new hire. However, once feedback is available, a model could draw hasty conclusions for future hires. The increasing use of LLMs to screen resumes and conduct interviews raises serious concerns about the formation of biases from hiring experiences, concludes Wang.

As LLMs learn from experience to make decisions about hiring, loan approvals, or parole, emerging, untrained biases must be monitored. These biases, although new, are still present, concludes Liu.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.