Uncertain AI Systems Can Be Discriminatory: How Can We Address This?
By Arina Shah, Holli Sargeant, Mackenzie Jorgensen, Adrian Weller and Umang Bhatt - 14 October 2025
When an AI system is uncertain about a decision, what should it do? A common technical solution is simply to abstain and pass the decision to a human. However, our research shows that this seemingly neutral approach can inadvertently worsen discrimination against marginalised groups. AI systems are often most uncertain about people who are underrepresented in the data used to train them. In this blogpost, which draws on our research paper, we compare two methods to deal with uncertain AI systems through the lens of the United Kingdom's anti-discrimination law: selective abstention and selective friction. We conclude that selective abstention still results in the risk of unlawful discrimination, but selective friction may be lawful depending on its implementation and how the decision-maker interacts with the model. Therefore, we argue that selective frictions are a more promising algorithmic intervention than selective abstention.
Introduction
AI systems play a pivotal role in decisions that shape our lives. These systems are meant to make decision-making processes more efficient and accurate, but what happens when these systems are uncertain? “Uncertainty” refers to how confident a model is in its predictions or outputs. In machine learning, people often distinguish aleatoric uncertainty (noise or randomness in the data) from epistemic uncertainty (limits in the model’s knowledge or training data). All AI systems can be uncertain. Uncertainty can arise from messy data or small datasets, or being asked to make a judgment of a situation the system has never seen before. As humans rely more heavily on AI systems, the uncertainty that underpins these systems becomes more than a technical detail. Rather, it has been shown that an AI system’s uncertainty is often disproportionately distributed towards individuals from underrepresented and historically marginalised communities, whose data is consistently sparse or biased.
Consider an AI system trained to diagnose skin cancer from photographs, where the training data consists primarily of images from patients with lighter skin tones, a common limitation in dermatology research. When this system encounters images of skin lesions on darker skin tones, it exhibits high epistemic uncertainty because it has limited training examples of how cancerous versus benign lesions appear on different skin tones and because it lacks knowledge about how factors like melanin content affect the visual presentation of conditions. This leads to darker-skinned patients being classified incorrectly and, as a result, failing to get the diagnosis and treatment that they could need.
There are different ways to design AI systems in a manner that handles uncertain predictions. Some researchers have proposed selective abstention as a solution to handling uncertain predictions and negative impacts on individuals and groups: if the model’s confidence is below a predefined threshold, then the AI output is withheld entirely, deferring the decision to human judgment. Another proposal is selective friction: the low-confidence AI outputs are still shown to the decision maker, but it is delivered with clear warnings, delays, or visual indicators that nudge decision makers to slow down and think critically. We provide the first integrated socio-technical and legal analysis of these two approaches. Drawing on our research, we argue that while it may seem intuitive to hide or friction a low-confidence AI prediction, doing so is fraught with legal risks and can perpetuate the very discrimination we seek to prevent.
The Problem of Unequal Uncertainty
In high-risk decision-making (e.g. in domains like healthcare, finance, and criminal justice), the concern with uncertainty is not solely its presence, but the way it is distributed across demographic groups. Hence, we analysed what happens under each approach to mitigate uncertainty. The first method, selective abstention, is employed to withhold low-confidence predictions. It rests on the assumption that deferring to human judgment will yield a more reliable outcome. However, if certain groups disproportionately trigger abstention, they are subjected to inconsistent decision standards and increased exposure to well-documented potential human cognitive biases. This is not to suggest that the model is an unbiased system; systems trained on historical data can encode and amplify those same biases. Selective abstention can therefore create a two-tier effect for underrepresented groups. Higher model uncertainty can push greater decisions to humans, where error and bias are less transparent and harder to audit. In such cases, a neutral uncertainty threshold can become a mechanism that reproduces or amplifies structural inequalities.
Selective friction, on the other hand, retains the prediction while providing confidence disclosures or introducing procedural pauses that encourage deliberation. This approach enables decision-makers to account for uncertainty actively rather than allowing it to dictate silently which cases are treated differently. A fair concern is that making uncertainty visible can itself bias human reviewers (e.g. anchoring on a “low-confidence” cue or becoming unduly sceptical). The effect is design-dependent: frictions that provide calibrated, interpretable signals (e.g. probability ranges), and those that pair those signals with uniform procedural rules (e.g. “if confidence < x, defer to higher authority or get second opinion), can improve consistency relative to blanket deferral. Unlike selective abstention, which withholds predictions once a threshold is crossed, selective friction keeps the recommendation visible while standardising how uncertainty changes the workflow. If done well, we argue that, contrary to selective abstention, selective friction enhances transparency, promotes consistent decision-making, and mitigates the risk of uncertainty as a hidden driver of discrimination.
Real-World Stakes: Loans & Likes
To ground the analysis, we examine two real-world settings, credit scoring and content moderation, showing how selective abstention and selective friction interact with unequally distributed uncertainty and shape outcomes across demographic groups. We do not assume any ground truth knowledge in both case studies.
Credit Scoring
Take the case of credit scoring. Machine learning has allowed lenders to automate and optimise risk assessments, but not without a trade-off. Historical lending data may contain bias, reflecting decades of unequal treatment in the financial sector. If a model is trained on that data, its confidence in certain applicants, especially women, minorities, or low-income individuals, may be systematically lower. This can lead to troubling patterns. For instance, applicants with low-confidence scores may be subject to human review (selective abstention) or flagged with warnings (selective friction), but the outcome is not always superior. Machines are not uniformly better than humans. However, human judgement is not uniformly fair either. In some cases, human cognitive biases, especially when they fall unevenly across demographic groups, can yield outcomes that are just as unfair, or sometimes even worse, than the model’s own errors.
In these scenarios, we assume a default-risk model (high/low) whose predictions all fall below the confidence threshold, with no ground truth at decision time. Four scenarios illustrate how selective interventions can operate in the face of uncertainty, where we present the model prediction, the final human decision with knowledge of the model’s prediction, and the impact on the individual:
- Model prediction: high-risk of default; Final Human Decision: high-risk of default; Impact: applicant is denied the loan. Agreement between the lender and the system may translate into overreliance on a weak signal without probing its basis, amplifying risk where uncertainty arises from skewed or sparse group data.
- Model prediction: low-risk of default; Final Human Decision: high-risk of default; Impact: applicant is denied. Disagreement between the lender and the system may translate to scepticism toward the model or implicit bias against the applicant, especially if this pattern disproportionately affects certain groups.
- Model prediction: low-risk of default; Final Human Decision: low-risk of default; Impact: applicant is approved. While this is a positive outcome and the lender and the system agree, if the prediction turns out to be wrong, then the lender bears the financial cost, and errors can entrench stereotypes.
- Model prediction: high-risk of default; Final Human Decision: low-risk of default; Impact: applicant is approved. This is a positive outcome, but the lender and the system disagree. Overrides can correct model bias due to data sparsity or representation gaps, but are often inconsistently applied and more likely to benefit already advantaged applicants.
Content Moderation
Another case is content moderation. Automated content moderation is often used to detect hate speech or violent content. But these tools are far from neutral. Studies have shown that posts from communities of colour, LGBTQ+ users, and political activists are disproportionately flagged or “shadow-banned” because the training data underrepresented these groups or mislabelled their speech. Shadow banning, where posts are demoted or hidden without notification, can suppress speech and engagement for creators who are already marginalised.
Similar to the case study before, four scenarios illustrate how uncertainty interacts with human review.
- Model prediction: content violates policy; Final Human Decision: content violates policy; Impact: Agreement between the moderator and system may reflect over-reliance on the model or a lack of cultural understanding, particularly if the content comes from a marginalised group.
- Model prediction: content does not violate policy; Final Human Decision: content violates policy; Impact: Disagreement between the moderator and the system leads to the removal of content. Here, human judgment can reintroduce bias, particularly if the override disproportionately targets certain groups.
- Model prediction: content does not violate policy; Final Human decision: content does not violate policy; Impact: Agreement between the moderator and the system leads to a positive outcome as the content remains online. While this seems benign, it may reflect a superficial review, especially if moderators are simply defaulting to uncertain model predictions.
- Model Prediction: content violates policy; Final Human Decision: content does not violate policy; Impact: Disagreement between the moderator and system in the opposite direction allows for content to remain. This could be a positive correction if the model was biased, but if such overrides are unevenly distributed, say, favouring verified users or creators with more social capital, it may still reinforce inequity.
Are these methods lawful?
Beyond the practical outcomes, these design choices have significant legal implications under the United Kingdom's anti-discrimination law. Both selective abstention and selective friction risk direct and indirect discrimination if their design or effects disproportionately impact individuals with protected characteristics.
Given the seemingly “neutral” nature of the uncertainty-based resignation or friction policy, the most likely legal challenge is indirect discrimination. Indirect discrimination arises when a practice, criterion, or policy that seems neutral on its face ends up putting people with a protected characteristic (like race, sex, or sexual orientation) at a distinct disadvantage compared to others. If the uncertainty threshold disproportionately sends loan applications from women or content from queer creators into a different review process, it creates a disadvantage for those groups. This disadvantage—such as procedural delays or the inconsistency of human review—is what may open the door to a claim of indirect discrimination.
A company can defend an indirectly discriminatory practice if it can prove the practice is a “proportionate means of achieving a legitimate aim” (Art. 19(2)(d) UK Equality Act). The goal of improving decision accuracy is certainly legitimate. The real test is proportionality: is the method used both appropriate and necessary, and is there a less discriminatory way to achieve the same goal? This is where the analysis between selective abstention and selective friction diverges. It is more difficult to argue that completely withholding an AI’s prediction and forcing certain groups into a separate, more burdensome process is a proportionate measure, especially when a less discriminatory alternative exists. For example, a less discriminatory alternative may be selective friction. By transparently communicating the AI’s uncertainty with a warning, friction still serves the goal of encouraging caution and improving accuracy. However, it does so without creating the same procedural penalty as abstention. Because it provides a less harmful way to achieve a similar outcome, we argue that selective friction represents a much more legally defensible position.
Conclusion
The challenge of building fair AI is not simply about creating more accurate models. It is about thoughtfully designing the entire human-AI system to be transparent, accountable, and just. Our research shows that well-intended technical solutions, such as selective abstention, may create procedural inequalities. Worse, they risk embedding unequal treatment under the guise of fairness. We need more than technical fixes. We need transparency, legal safeguards, and empirical research that investigates how these interventions affect real people, especially those from protected groups. We argue that selective friction, if implemented well, offers a more promising path forward. By communicating uncertainty instead of hiding it, we can try to preserve transparency, encourage critical human oversight, and reduce the risk of unlawful discrimination.
Note on the authors
Arina Shah is a NYU ’25 graduate in Data Science & Philosophy.
Holli Sargeant is a Research Fellow in Law at St John’s College, University of Cambridge.
Mackenzie Jorgensen is a Postdoctoral Research Fellow at Northumbria University (LinkedIn/ BlueSky / Twitter).
Adrian Weller is a Research Professor at the University of Cambridge, Head of Safe and Ethical AI at the Alan Turing Institute, and a Fellow in Computer Science at Sidney Sussex College, Cambridge.
Umang Bhatt is an Assistant Professor at the University of Cambridge and a Fellow in Computer Science at King’s College, Cambridge.
This article solely reflects the views of the authors, and does not represent the position of the Faculty or the University.