
Ethan Goh, MD
Executive Director, ARISE (AI Research and Science Evaluation)
Focused on how frontier AI models are safely translated into real world care.
Engaged in multi-institutional efforts shaping the evaluation and alignment of medical AI systems.
Advises healthcare, life sciences, technology teams, and investors on clinical AI strategy, translational pathways, and commercialization.
Current Initiatives
- ARISE
Stanford-Harvard research network advancing evaluation and real-world deployment of AI systems in healthcare
- Medical AI Superintelligence Test (MAST)
Unified evaluation framework for frontier models across clinical reasoning, safety, and decision-making
- Specialist AI Guiding Experts (SAGE)
AI-based clinical decision support system designed for prospective evaluation in provider workflows
Selected Work
- Research
- GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trialEditorial CommentaryMaking large language models into reliable physician assistants
- Large language model influence on diagnostic reasoning: a randomized clinical trialEditorial CommentaryLarge language models — misdiagnosing diagnostic excellence?
- Performance of a large language model on the reasoning tasks of a physicianEditorial CommentaryAI can reason like a physician—what comes next?
- GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial
- Clinical AI Benchmarks
- First, do NOHARM: towards clinically safe large language models
- Automated evaluation of large language model response concordance with human specialist responses on physician-to-physician eConsult cases
- Asking the right questions: benchmarking large language models in the development of clinical consultation templates
Reports & Synthesis
- State of Clinical AI Report
Synthesis of themes across high-impact clinical AI research, translating latest findings into implications for real-world deployment
- HAI AI Index: Science & Medicine
Benchmarking major advances in AI within biomedical research and clinical care
- FDA public comment: measuring AI medical device performance
Invited commentary on real-world evaluation frameworks, including LLM-as-judge approaches for post-deployment oversight
Perspectives
- How physicians actually use AI
Training doctors to use AI matters less than task specific workflow and product design
- Omission harms are the dominant safety risk in clinical AI
The overlooked danger of AI failing to suggest critical actions
- Why AI agents inside the EHR remain unreliable
Technical barriers to autonomous clinical actions in electronic health records
Selected Media & Commentary
- A.I. Chatbots Defeated Doctors at Diagnosing Illness
The New York Times
- AI Helps Prevent Medical Errors in Real-World Clinics
TIME Magazine
- Meet ARISE: A Stanford- and Harvard-Backed Lab Dedicated to Objectively Validating AI in Healthcare
Forbes
- Multiple Reasoning Models and the Future of AI Chatbots
JAMA
Affiliations
- Stanford Division of Computational Medicine
Executive Director, ARISE Network
- Harvard Medical School
Faculty Director, Generative AI and Agentic Systems
- BMJ Digital Health & AI
Founding Editorial Board
- Google
Visiting Faculty Researcher
Speaking
- Selected Talks
- AI is Better Than Doctors (At Some Things). Now What?
HIMSS AI, 2025
- Moving AI Into The Real World
Stanford AIMI Symposium, 2025
- Foundation Model Roadmap: What Health AI Teams Need to Know
OpenAI, Google Industry Panel, 2025
- The State of Clinical AI
Accelerate AI Executive Summit, 2026
- AI is Better Than Doctors (At Some Things). Now What?
- Teaching
- Stanford Healthcare AI Leadership & Strategy Program
Stanford Division of Computational Medicine
- Harvard Generative AI and Agentic Systems Executive Course
Harvard Medical School
- MED 279: Stanford Health Consulting Group
Stanford Medicine
- Stanford Healthcare AI Leadership & Strategy Program
Advisory
Select advisory and board roles with healthcare, life sciences, and technology organizations.
Contact
