AI vs Human

AI vs Human Progress — how fast LLMs caught up

Every year a new model posts a higher score on the tests we made to measure human knowledge. This page plots those scores against the published human baselines, then sets them next to the years of learning a person actually needs.

Chart 1The race to the human line

Best reported score each year, one line per benchmark. Colour shows which benchmark a line belongs to; every legend entry names the human baseline that benchmark is measured against. Dashed lines are the human-expert baselines.

Dashed line = published human-expert baseline for that benchmark. Benchmarks with no published baseline are drawn against the 100% ceiling and carry no dashed line.

Chart 2Childhood vs machine

Two timelines on the same scale of years. The top lane is when a person typically reaches each stage of learning; the bottom lane is when an LLM reached comparable tasks, counted from 2018 (the first generative pre-trained transformer). It is an illustrative comparison of learning time, not a benchmark score.

A person
An LLM
Both lanes share one x-axis: years from the start of learning (age 0 for a person, 2018 for the LLM). A person takes about 22 years to a degree and 26 to a PhD; the LLM crossed human-expert baselines in about five to six years.

Chart 3Where it stands today

The latest score on each benchmark against the human baseline. The bar is the AI score; the orange tick marks the human baseline. A ✓ means the AI score has reached or passed the human baseline, ✗ means it has not, and "no baseline" means the benchmark has never published a human percentage to compare against.

Blue bar = best reported AI score. Orange tick = human-expert baseline.

How to read this

Three things matter when you look at these lines.

A benchmark is a test, not a person

Scoring 90% on MMLU means answering 90% of a set of multiple-choice exam questions correctly. It does not mean the model would do 90% as well as a human at a real job. Benchmarks saturate: once models cluster at the top, the test stops telling us much.

"Human baseline" means a specific group of people

MMLU's 89.8% is an estimate of expert performance, GSM8K's 95% is a human accuracy figure, GPQA's 65% is the score of PhD holders answering in their own field, and MMMU's 88.6% is its best human expert. A different group, or a harder subset, gives a different line — always check who the "human" was.

The tilde (~) marks an approximate value

Scores published between benchmarks, or rounded in a summary, are shown with a ~ in their label. Values without a tilde are the score as reported by the cited source.

Frequently asked questions

How fast have LLMs improved compared with humans?

On MMLU, the best model went from 43.9% in 2020 (GPT-3) to roughly 93% in 2026, crossing the 89.8% human-expert baseline by 2025. On GPQA Diamond, the best model went from 38.8% in late 2023 to about 94% by 2026, passing the 65% PhD-expert baseline in 2024. AI reached expert-level scores in about five to six years, where a person spends roughly twenty-two years in education to reach the same level.

Where do these AI benchmark numbers come from?

Human baselines and early scores come from the Stanford HAI Artificial Intelligence Index 2024 report and the original benchmark papers (MMLU 2020, GSM8K 2021, HumanEval 2021, GPQA 2023, MMMU 2023). Later scores come from lab announcements and public leaderboards such as Epoch AI. Every point is labelled with the model that set it, and approximate values are marked with a tilde symbol.

What does the human learning curve show?

The second chart lines up the years of study a person needs (reading at about age 6, algebra at about 13, a degree at about 22, a PhD at about 26) against the years an LLM needed to reach expert-level scores on the same kind of tasks. It is an illustrative comparison of learning time, not a benchmark score, and it is labelled as such on the page.

Have AI models beaten humans on every benchmark?

No. Of the four benchmarks here with a published human baseline, AI has clearly passed MMLU, GSM8K and GPQA Diamond but still sits below the best human-expert score on MMMU (about 83% versus 88.6%). HumanEval and MATH have no published percentage human baseline, so they are shown against a 100% ceiling without a human marker. Beating a benchmark is not the same as matching a person in the real world.

Sources & method

  • Human baselines and early scores — Stanford HAI, Artificial Intelligence Index Report 2024, Chapter 2 (Technical Performance): MMLU human baseline 89.8%, GSM8K human 95%, MMMU best human expert 88.6%.
  • GPQA Diamond — Rein et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023): PhD expert 65% (74% after adjustment), non-expert 34%. Model scores from the same paper and from Epoch AI.
  • MMLU — Hendrycks et al. (2020), with the 89.8% expert estimate from the GPT-4 technical report (OpenAI, 2023).
  • HumanEval — Chen et al., Evaluating Large Language Models Trained on Code (2021). MATH — Hendrycks et al. (2021).
  • Later model scores — lab announcements and public leaderboards (Epoch AI, Papers With Code, and vendor model cards). Reported figures differ between runs because of prompt format, few-shot settings and scaffolding; the highest reported score for a leading model in that year is used.

This page is a summary for general understanding. It selects the single best reported score each year to show the frontier; it is not a controlled re-evaluation of every model on the same harness.

Related tools