How We Built Our Synthetic Dataset to Reduce Bias
Avoiding bias and ensuring equity were the reasons we built this dataset ourselves.
We wanted Leo to practice helping people who are too often overlooked by healthcare tools: people with public insurance, people without insurance, people with limited English, people with low health literacy, people in rural areas, disabled people, caregivers, veterans, and people managing chronic illness.
We also wanted to protect privacy from the start. That meant building a synthetic dataset instead of relying on real patient records or scraped conversations.
No Real Patient Records in This Dataset
This dataset was built from scratch to help Leo practice patient advocacy in a safer, more equitable way. No real patient records. No real conversations. No real names, insurance numbers, or medical histories. Every scenario in this dataset is synthetic: invented on purpose, to teach the right behavior without touching a single real person's health data.
This is not Leo's only source of learning. It is one important dataset we created because equity needed to be built into the foundation, not patched in later.
We could have taken a faster path by relying on real patient records or scraped conversations. We chose not to, because your medical history isn't training material for a product.
We also chose not to let a generic model decide what a “typical” patient looks like. In healthcare, “typical” too often means the patient who already has the most access.
Built With Every Kind of Patient in Mind
Most healthcare tools learn from the patients who already have the easiest path through the system: English speakers, people with strong insurance, people who already know how to push back.
This dataset was not built that way. We designed it to put harder, more overlooked patient experiences closer to the center.
The scenarios cover a wide range of real situations: different insurance types, different languages, different levels of health literacy, disability, rural access, limited English proficiency, low income, low digital literacy, veterans, caregivers, and people managing chronic illness.
That matters because bias does not always show up as obvious harm. Sometimes it shows up as an answer that assumes you have transportation, a specialist, paid time off, strong insurance, medical vocabulary, or someone who can help you fight a denial.
Leo needs to work for patients who do not have those advantages.
A patient advocate that only knows how to help the easiest cases isn't much of an advocate.
Addressing Bias was Intentional
We built this dataset because building in equity changes the product.
It changes the questions Leo practices answering. It changes the reading level. It changes how Leo responds when someone is scared, confused, uninsured, dismissed, or unsure what to ask next.
It also changes what we check before a scenario is good enough. We look for whether Leo stays in scope, uses plain language, respects uncertainty, handles crisis moments safely, and gives support that makes sense across different patient experiences.
The goal is not to claim Leo is bias-free. No serious healthcare AI company should say that.
The goal is to keep finding bias, reducing it, and building better support for the people most likely to be harmed when healthcare tools are built around narrow assumptions.
Leo Knows Where the Line Is
Leo does not diagnose you. Leo does not prescribe medicine, change your treatment, or give you definitive legal advice. Those limits are built into the training, including how Leo responds in crisis or safety situations. Leo is built to recognize those moments and respond the right way, every time.
A Separate Process Reviews Every Scenario
Building a dataset is one problem. Trusting it is another. We didn't let the same process that generated the scenarios also decide if they were good enough.
A separate review checks the data against patient-centered standards: is it safe, does it reflect a real range of patients, is the language plain enough for someone with no medical background, is Leo honest about what it can and can't do, and does Leo handle crisis moments correctly.
Equity is part of that review. A scenario is not strong just because it sounds medically fluent. It has to support the kind of patient who may already be fighting confusion, cost, bias, language barriers, disability, or lack of access.
This Keeps Changing
Healthcare changes. Language changes. The patients who need Leo change. We keep testing and adding to the training data so Leo keeps up, instead of stopping on day one and calling it finished.
Your Health Information Stays Yours
Leo works for you. Not your insurance company. Not a hospital. Not us.
Get ProjectLeo on the App Store or on the web.
Frequently Asked Questions
Did this dataset use my data or another patient's medical records?
No. This dataset is 100% synthetic. We did not use real patient records, real conversations, or real identifying information to create it.
What does "synthetic data" mean?
It means every scenario in this dataset was created on purpose, not pulled from a real person's life. We built situations that reflect what real patients go through, without using anyone's actual health information.
Can Leo diagnose me or tell me what medicine to take?
No. Leo does not diagnose conditions, prescribe medicine, or change your treatment plan. Leo helps you understand your options, prepare for appointments, and speak up when something's wrong. Always talk to a licensed healthcare provider about medical decisions.
How do you know the training data is fair to different kinds of patients?
We built scenarios covering a wide range of patients on purpose: different languages, insurance types, disabilities, income levels, and health literacy. A separate review checks whether the dataset reflects patients who are often left out, not just the patients healthcare tools usually center.
We do not claim any dataset is perfect or bias-free. We do the work, check for gaps, fix what we find, and keep improving.
Is this a one-time thing, or does it keep changing?
It keeps changing. We review and add to the training data on an ongoing basis, so Leo keeps working for real patients, not just the patients we imagined at the start.
⚠️ Leo provides health information and guidance, but does not replace professional medical advice, diagnosis, or treatment. Always consult with a qualified healthcare provider about your health concerns. In an emergency, call 911 immediately.