Multimodal foundation models and agents for medicine. I develop generalist medical AI that learns jointly from biomedical data and navigates the healthcare systems, and I study how to compose these foundation models into agents that can reason over a case, gather the evidence they need, and collaborate with clinicians across the diagnostic workflow — moving from single-task predictors toward systems that participate in the full arc of clinical care.
Rigorous evaluation environments and benchmarks for medical AI. Impressive benchmark scores rarely tell us whether a model is safe to use in the clinic. I build the evaluation infrastructure needed to close this gap: dynamic clinical environment simulators that test models in realistic, interactive settings, and benchmarks that probe grounding, robustness, and behavior across diverse populations — so that progress in medical AI is measured by clinical readiness, not leaderboard position.
Bridging AI to patients, clinicians, and health systems. A model only matters when it reaches the people it is meant to serve. I create generative tools that translate complex diagnostic findings into language patients can understand, interpretable and reliable decision support that clinicians can trust and verify, and methods for monitoring and continuously improving models after deployment — treating AI not as a replacement for clinicians, but as infrastructure for a learning healthcare system.
Google Scholar
LinkedIn
GitHub
Twitter