Responsible AI
AI systems trained on historical data inherit historical inequities. When those systems are deployed at clinical scale, predicting which patients need intervention, which populations carry elevated risk, which deterioration alerts should fire, the hidden assumptions scale with them. Responsible AI is the discipline that makes those assumptions visible: auditing who the system fails, why it fails them, and whether we can explain and correct it before the harm compounds.
Working with a large multi-state health system in the United States, we deployed a model to predict care deferral: patients who would cancel or fail to appear for scheduled appointments (Ahmad et al., 2022). The model performed well by standard accuracy measures. But when we disaggregated its predictions by race and socioeconomic status, a pattern emerged — patients flagged as high-risk for deferral were disproportionately racial minorities. The algorithm had correctly identified a disparity. It could not explain one.
What happened next required something no model can do on its own. Working alongside clinicians, care coordinators, and community health workers, we set out to understand why these patients were deferring. What emerged was not a story about disengagement. These patients were primary caregivers — for children, for elderly parents, for partners with chronic illness — who could not take a half-day from work for an appointment. They were navigating scheduling systems built around a workday that assumed flexibility they did not have.
The solution was a mobile clinic: bringing care to the patients rather than requiring patients to come to the clinic. The program achieved a 12% reduction in deferral rate disparity across demographic groups, narrowing a gap that the model had identified but could not close on its own. The lesson was precise: machine learning can identify who is being left behind. Closing that gap requires understanding why — and that requires talking to people.
My research in this area spans three interconnected problems. First, fairness: AI systems trained on historical data inherit historical biases, and I study how those biases manifest across race, gender, age, and socioeconomic status in high-stakes settings, particularly healthcare (Ahmad et al., 2020; Ahmad et al., 2021). Second, explainability: a model that clinicians cannot understand is a model they cannot trust or correct, and I develop methods that make machine learning legible to the people who depend on it (Ahmad et al., 2018; Ahmad et al., 2019). Third, trustworthy LLMs: large language models introduce new failure modes (hallucination, confident error, catastrophic forgetting) that are particularly dangerous in clinical settings, and I am working on frameworks to detect and reduce them (Ahmad et al., 2023).
This work is motivated by a conviction: AI systems do not just reflect the world as it is. They shape the world as it will be. It has shaped regulatory frameworks for clinical AI (Ahmad et al., 2021), model reporting standards (Ahmad & Eckert, 2023), and research into algorithmic fairness in end-of-life decisions (Ahmad, 2025) and robust clinical decision support under uncertainty (Preuett et al., 2025). I am part of the Responsible AI Systems and Experiences (RAISE) initiative at the University of Washington. Building them responsibly is not a constraint on progress: it is what progress requires.
2025
2023
2022
2021
2020
2019
2018
2021
2020
2018
-
AAAI · 2019–2020
Co-Chair, Symposium on AI for Social Good
-
ACM FAccT
Panel Organizer, Healthcare AI and Algorithmic Fairness
For research collaboration, reach me at maahmad@uw.edu. For advisory engagements and responsible AI governance, reach me at vonaurum@gmail.com.