openai
Realistic multi-turn health conversations graded against physician-written rubrics across multiple axes (accuracy, completeness, communication) — an open evaluation benchmark for AI assistants in medicine.
Last 30 days
Explore
Read the real rows without downloading anything
Reading rows…
LiveRead from openai/healthbench at the moment you asked. Nothing is cached or stored — every row above came from that request.