Synthetic data · research/education only · not a diagnostic tool · do not enter real health data

I fine-tuned two open medical AIs to reason about health trends — safely.

Same task, same synthetic data, same recipe on two models — Gemma 3 and its medical sibling MedGemma. The question: what does fine-tuning fix, and does a "medical" model actually do better?

Fine-tuning fixed the safety problems. Both models went from breaking the safety rules to following them every time.
🔍
The medical model didn't win. After fine-tuning it was tied with the general one — a capable base + fine-tuning is what mattered, not the medical pretraining.

Overall quality score, before vs after fine-tuning

Higher is better (0–1). Automatic score: safe framing + correct structure + no made-up numbers.

before fine-tuningafter fine-tuning

See a real before → after example (one synthetic patient)

MedGemma

before fine-tuning

same model

after fine-tuning