Equity in Pulse Oximetry
MIT Policy Hackathon 2024 || 48-hour policy case for equitable pulse oximetry
TYPE
Policy research and strategy
CONTEXT
MIT Policy Hackathon 2024, Health Equity Challenge
TEAM
The Decisive Leftovers, 5 members across policy, media, data science, and design
TIMELINE
48 hours, November 15 to 17, 2024
DELIVERABLE
Policy brief addressed to the Director of the FDA Center for Devices and Radiological Health
METHODS
Secondary research, regulatory landscape analysis, precedent case analysis, interpretation of clinical data, policy brief drafting
OUTCOME
Runner-up recognition in the health equity challenge
JUDGED BY
Dr. An-Kwok Ian Wong, Dr. Marie-Laure Charpignon, João Matos, Lisa Abrams, Joanna Brownstein
THE CHALLENGE
A pulse oximeter is clipped onto a person's fingertip and shines light through it to measure the oxygen in blood. The problem arises with higher melanin, as melanin absorbs light, resulting in inaccurate results. The device was never built to handle that difference.
Nearly 16 million American adults have COPD. A lot of them check their oxygen with a common oximeter in a hospital, and make real decisions based on the number it shows.
That number is not equally reliable for everyone. Black patients are three times more likely to get a reading that does not match the actual oxygen in their blood. For 12% of Black patients, the device shows a safe reading while their oxygen is dangerously low. For white patients, that happens 4% of the time. And 80% of hospital oximeter use happens outside the ICU and operating room, where a wrong reading is easy to miss.
We were handed a health equity challenge about device accuracy. But the real problem was in the rules.
The FDA's 2013 guidance says manufacturers should include at least two dark-skinned subjects, or 15% of participants, when testing a device before it goes to market. Should, not must. Two people. A suggestion, not a requirement.
MY ROLE
I was one of five people on a team assembled at the hackathon. Akin came from policy, Jamil from media, Justin and Sujana from computer and data science. I came from design research and strategy. We each worked from our own expertise, so that the brief could hold both technical evidence and a coherent argument.
I led the secondary research and the strategy formation and built our understanding of how the device works, where the disparity originates, and what the FDA has and has not required. I shaped the structure of the policy brief and co-authored it. I was part of every discussion about what the data meant, though I did not write the analysis code. Justin and Sujana did that work.
This was my first time writing policy. The hackathon ran a two-hour workshop on how policy is drafted, and I went into it not knowing the form. That is worth saying plainly, because a lot of what I learned in those 48 hours came from being the person in the room who did not already know the answer.

HOW IT WORKS
An emitter sends red and infrared light through the finger. A sensor on the far side reads how much passes through. The ratio between the two becomes your oxygen number.
Red light (660 nm)
Absorbed more by oxygen-poor blood.
Infrared light (940 nm)
Absorbed more by oxygen-rich blood.
Shown in violet; infrared is invisible
Illustration is schematic, not to scale.
RESEARCH
We had two days. That meant secondary research done fast and read carefully, rather than primary research done thinly.
PROCESS
Four parallel streams of secondary work fed one synthesis. The data stream came back with nothing, and that sent us back to the question.

The reframe is the project. A flat result was not a dead end, it was a question about the instrument.
The null result specifically drove Recommendation 3. Recommendations 1 and 2 came out of the regulatory and precedent streams.
ORIGIN OF THE FLAW
The disparity is not an accident of manufacturing. In the 1980s the medical industry chose a simpler, more commercially viable oximeter design over an earlier one that accounted for skin tone variation. The earlier device used eight wavelengths of light and could be adjusted per individual. The design that won uses two, and was validated largely on light-skinned populations. The bias was designed in, decades ago, and has been carried forward ever since.
READING THE ARGUMENTS
We spent real time debating the articles we found, on the physics of melanin absorption, on hidden hypoxemia, on the FDA's own request for comment. Disagreement was useful. Two of us would read the same paper and take different things from it, and the argument that followed usually surfaced the assumption neither of us had noticed.

MAPPING THE REGULATORY MACHINERY
The FDA clears pulse oximeters through the 510(k) pathway, which asks manufacturers to show their device is substantially similar to another already on the market. Hundreds of clearances now trace back to designs green-lit before health equity was a consideration. Accuracy is judged by an ARMS threshold under three percent, a single number that averages across everyone and therefore hides differences between groups. Over-the-counter oximeters, the ones people bought during the pandemic, receive no FDA review at all.
INTERPRETING THE CLINICAL DATA
Our data members analyzed the BOLD dataset, which links 49,099 paired measurements: SpO₂ from the oximeter, SaO₂ from arterial blood gas, the invasive gold standard. We reviewed the outputs together as a team.
One fact shaped how carefully we worked.
The BOLD dataset was assembled by João Matos, An-Kwok Ian Wong, and their collaborators. Those researchers were on our judging panel. We were analyzing their data, citing their papers, and preparing to tell them what we thought it showed. There was no room to be approximately right.
Distribution of Discrepancy (Sp02 - Sa02)
Fig1. Source: Analysis done on BOLD

Distribution By Age
Fig 2. Source: Analysis done on BOLD

Box Plot for Discrepancy by Race
Fig 3. Source: Analysis done on BOLD
SYNTHESIS
The regression returned a null result. Using self-reported race as the variable, there was no strong correlation with the SpO₂ and SaO₂ discrepancy. The box plots looked flat across racial categories.
For a moment this read as a failure. The literature says the disparity is real. Our data would not show it.
THE REFRAMING
The same 60 people, plotted on the same melanin axis. On the left they are sorted into racial categories. On the right they are measured.

Illustrative. Melanin values shown are conceptual and are not measured data. The overlap between racial categories is the argument, not a finding from the BOLD dataset.
Then we looked at the individual measurements. Across fifty thousand data points, the deviations were large and clearly present. They simply did not sort by race.
Race was not failing to explain the discrepancy because the discrepancy was not there. Race was failing because race is the wrong variable.
Two people who check the same box on a form can have very different melanin levels. The category is too coarse to capture what the light is actually doing to the skin. Every study, every FDA standard, and every clinical protocol that uses racial categories as a proxy for pigmentation is measuring the wrong thing and then reporting the noise.
The second insight came from looking for precedent. The Americans with Disabilities Act, passed in 1990, took a problem that had been treated as individual misfortune and made it a design requirement enforced across public and private sectors.
181
countries have since passed disability rights legislation modeled on the ADA
It is proof that a federal standard can raise the floor for a vulnerable population and drive industry innovation at the same time. We used it in the brief as the model for what medical device reform could look like, so that our recommendations would read as achievable rather than aspirational.
That finding became Recommendation 3, the only one of the three that came directly out of the data.
Recommendations 1 and 2 came from the regulatory research and the precedent case.
IDEATION
Ideas on the table included calibrating the device with an additional melanin sensor, moving the reading site away from the fingertip to avoid the epidermis, returning to a multi-wavelength design, using temperature sensors to detect blood flow, and taking measurements at two locations so that a large differential would trigger an invasive check.
Two options we considered seriously and rejected in writing:
Rely on arterial blood gas testing for high-melanin patients.
Rejected.
-
ABG is more invasive, more expensive, and requires specialized training.
-
It would create a new barrier to care and is impractical for continuous monitoring or home use.
-
A fix that only works for people already inside a well-resourced hospital is not an equity fix.
Build software correction algorithms to adjust readings.
Rejected.
-
It would entrench racial categories as the input rather than replacing them, and it adds a layer of complexity on top of a hardware problem without solving the hardware problem.
-
It also puts a model between the patient and the measurement, in a setting where the measurement is meant to be trusted.
STRATEGY
Three recommendations, addressed to Dr. Michelle Tarver, Director of the FDA's Center for Devices and Radiological Health.
Pulse oximeters read less accurately on higher-melanin skin, and nothing in the system requires them not to.
THE PROBLEM
PREMARKET TESTING
01
Mandate representative testing standards
-
Match testing demographics to the latest US Census racial breakdowns
-
Replace the current minimum of two “darkly pigmented subjects”
-
Require manufacturers to document and report demographic testing data
More rigorous pre-market testing. Better accuracy across diverse populations. Manufacturer accountability.
CLINICAL PRACTICE
02
Establish a dual-verification safety protocol
-
Primary reading with the standard hand-clip pulse oximeter
-
Secondary verification: ear lobe reading, or a minimally invasive earlobe blood drop
-
The earlobe prick is less invasive than an arterial blood gas draw
Direct blood oxygen measurement. Less invasive than ABG. Personalized calibration.
MEASUREMENT SCIENCE
03
-
Objective melanin measurement tools, such as the colorimeter used in CHOP’s pediatric study
-
A standardized melanin scale to replace subjective racial categories
-
Require manufacturers to test against the scale, and build a database linking melanin to accuracy
Prototype: patient risk assessment
Better identification of at-risk patients. Testing standards that hold as the population changes.
Recommendations as submitted in the policy brief to the FDA Center for Devices and Radiological Health, 17 November 2024. Wording condensed from the brief’s key components.
Who gets tested, and who gets treated
The FDA suggests a floor. The population is the actual distribution. Recommendation 1 asks that the first match the second, and that it be a requirement.
.png)
42%
of the US population is not non-Hispanic white. The guidance suggests testing 15%, and does not require it.
Recommendation 1: match the testing population to the Census breakdown, and require it rather than suggest it.
Population: US Census Bureau, 2023, as cited in Figure 1 of the policy brief. The 42% figure is derived from the same breakdown. Testing guidance: FDA, Pulse Oximeters, Premarket Notification Submissions [510(k)s], 2013. Bar colors are categorical and do not represent pigmentation values.
Standardized melanin metrics Prototype
Alongside the brief, we built a prototype risk assessment platform, so that an individual could estimate their own likelihood of receiving an overestimated reading. The concept was mine. Policy moves slowly, and a person using an oximeter at home today has no way of knowing the number might be wrong for them. Justin built the working version. We framed it in the brief as an MVP still in development, because that is what it was. It was a demonstration that this research could reach the public, not a validated tool.

MVP, in development
OUTCOME
We presented the brief on Sunday, at the end of 48 hours, to a panel that included two of the researchers who built the dataset we had spent the weekend inside. Our team received runner-up recognition in the health equity challenge.
The challenge was won by a team called OxEquity, made up of graduate students in public health, biomedical engineering, and economics. Their proposal reweighted the FDA's ARMS accuracy metric and quantified device performance at the level of demographic subgroups. One judge described what made it win: it was accessible and actionable, and it required less political capital to implement in the clinic.
That contrast is the most instructive thing about this project, so it belongs here rather than buried.
OxEquity worked within the racial subgroups. We argued the subgroups themselves were the flaw. Their recommendation could be adopted by building on processes the FDA already runs. Ours asks the FDA to change the variable it measures, which is slower, costlier, and harder to sell. Both readings of the evidence are defensible. Theirs was more implementable. I still think ours was more correct.
The brief itself is the artifact. It was written as a real document to a real named official, in the form the FDA reads, with sources, rejected alternatives, and a precedent case. It was not implemented, and I will not claim otherwise. What it demonstrates is a chain of reasoning that holds: a documented problem, a regulatory gap, evidence read carefully enough to reframe the question, and three recommendations that answer the reframed question rather than the original brief.
PERSONAL LEARNINGS
01.
The most useful result was the one that looked like nothing
A null correlation is easy to discard. We nearly did. What made the difference was asking why the data refused to cooperate, instead of asking what else we could run. The answer, that our variable was wrong, became the strongest argument in the brief. I now treat a flat result as a question about the instrument rather than a dead end.
02.
Categories carry assumptions, and the assumptions get built into hardware
A form asks for race because race is what forms have always asked for. Then a study uses that field, then a standard cites that study, then a device is cleared against that standard, and thirty years later a patient's low oxygen goes undetected. The category was a shortcut nobody revisited. That chain is a design problem before it is a medical one.
03.
Constraints clarified the strategy
With 48 hours and no time for primary research, we could not go wide. We had to decide early what question we were answering and defend it. I have run longer projects that were less focused.
04.
The correct answer and the adoptable answer are different answers, and a strategist has to know which one is being asked for
We were asked for policy recommendations, and we gave the FDA the recommendation the evidence pointed to. The winning team gave the FDA a recommendation it could act on next quarter. I have thought about that judging comment for a long time. Political capital is a design constraint, the same way cost and manufacturing tolerance are design constraints. I had treated it as someone else's problem. The strongest version of our brief would have made the same melanin argument and then sequenced it, showing the FDA a first step it could take within existing processes and a path to the harder change. The insight was right. The implementation runway was missing.
05.
What I would do differently
I would learn the analysis rather than interpret it secondhand. I sat in every discussion of the BOLD results and helped decide what they meant, and I could not have produced them. That gap is real, and closing it is the skill I most want to add. The reframing insight came from reading the output. I want to be the person who can also generate it.
Sources: CDC Chronic Disease Indicators, COPD (2024); US Census Bureau (2023); FDA 2013 guidance on pulse oximeter premarket notification; Wong et al., JAMA Network Open (2023); Matos et al., BOLD dataset, PhysioNet (2023); Johns Hopkins Bloomberg School of Public Health (2024).
VA Red Coat
Coming Soon