Domain
Benchmark yourself
One human.
Fifty-seven fields.
A reproducibly random sample of 100 multiple-choice questions from the MMLU test set, spanning mathematics, law, medicine, history, and more.
Your saved progress is still here.
Your MMLU estimate
—%
— / 100 correct
95% interval: —
This is a 100-question human, zero-shot estimate—not the full model benchmark.
Breakdown