openai
Professional-graded subset of HealthBench: physician evaluators score model responses to clinically realistic conversations, targeting expert-level health assessment.
Last 30 days
Explore
Read the real rows without downloading anything
Reading rows…
LiveRead from openai/healthbench-professional at the moment you asked. Nothing is cached or stored — every row above came from that request.
Column statistics
Over 525 rows of default/test