Synthetic data · research/education only · not a diagnostic tool · do not enter real health data
I fine-tuned two open medical AIs to reason about health trends — safely.
Same task, same synthetic data, same recipe on two models — Gemma 3 and its medical
sibling MedGemma. The question: what does fine-tuning fix, and does a "medical" model actually do better?
✅
Fine-tuning fixed the safety problems.
Both models went from breaking the safety rules to following them every time.
🔍
The medical model didn't win.
After fine-tuning it was tied with the general one — a capable base + fine-tuning is what mattered, not the medical pretraining.
Overall quality score, before vs after fine-tuning
Higher is better (0–1). Automatic score: safe framing + correct structure + no made-up numbers.
before fine-tuningafter fine-tuning
Code & full write-up →
Hugging Face ↗
See a real before → after example (one synthetic patient)
MedGemma
before fine-tuning
same model
after fine-tuning