AGI: One Word, Two Definitions That Do Not Agree

✓ sourced to official docs · Published 2026-08-19

“Are we close to AGI?” is asked constantly and answered confidently in both directions by people looking at the same systems. That is a clue. When capable, informed people disagree that sharply about a factual question, it is usually not a factual question.

The word names at least three different things, and the two most-cited definitions do not agree on which of them count.

Two definitions that do not line up

The OpenAI Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work”.

Morris et al. (2023), from Google DeepMind, propose a framework built the opposite way — deliberately pulling apart what that sentence merges. It sets out six performance levels crossed with two generality columns, and states as a design principle that a definition should “Focus on Potential, not Deployment”. Autonomy is a separate scale entirely, from AI as a Tool through to AI as an Agent.

Read them side by side and the disagreement is structural, not a matter of emphasis:

OpenAI CharterLevels of AGI
Capabilityimplied by “outperform humans”explicit, six graded levels
Generality”most economically valuable work”explicit axis: Narrow vs General
Autonomyfolded in — “highly autonomous”separate axis, six levels
Deploymentrequired — it must do the workexplicitly excluded

One treats autonomy as part of the definition; the other treats it as an independent dial. One requires the system to actually be doing economically valuable work; the other explicitly says potential is what counts. Two people using these two definitions can look at an identical system and correctly reach opposite conclusions.

Draw your own line

Pick the cell you would call AGI. The tool reports what you have just required — and what you have not.

Draw your own line which cell do you call AGI?

The Levels of AGI framework separates two things people usually merge: performance (Emerging at roughly an unskilled human, Competent at the 50th percentile of skilled adults, Expert at the 90th, Virtuoso at the 99th, Superhuman above every human) and generality (a narrow scoped task versus a wide range of non-physical tasks including learning new skills). Autonomy — tool, consultant, collaborator, expert, agent — is a third axis again.

A chess engine sits at Superhuman Narrow and has for decades, which is why narrow performance alone never satisfied anyone as a definition of AGI. Where exactly the line falls on the general column is the thing people disagree about, usually without noticing they are disagreeing about a definition rather than about the evidence.

Every level name and threshold here is from Morris et al. Nothing marks where current systems sit, because that is exactly the contested claim — the point is to see what your own definition demands.

Two things usually surface within a few clicks.

Narrow performance was never the argument. A chess engine has sat at Superhuman Narrow for decades and nobody called it general intelligence. Whatever AGI means, it is a claim about the General column — which is why benchmark scores on scoped tasks move the debate so little.

Autonomy is a choice, not a consequence. The framework separates it deliberately: a system can be highly capable and used purely as a consultant, or modestly capable and given wide latitude to act. Bundling autonomy into the definition — as the Charter does — means arguing about capability and deployment in the same breath, which is a large part of why the argument goes in circles.

What this article will not tell you

Whether we are close.

Not because it is unknowable in principle, but because the honest answer is it depends which of those definitions you meant, and this site does not publish claims it cannot support. There is no agreed measure here, no benchmark whose passing settles it, and no way to check a forecast until it has already happened. A confident prediction would be the least verifiable sentence on this entire site.

What is worth saying is narrower and firmer. The rest of this cluster describes what these systems demonstrably do: text becomes tokens, tokens are positioned by the company they keep, the model sees only what fits in one request, and the next word is drawn from a reshaped distribution. None of that mechanism settles the AGI question either way — it is a description of the machinery, not a measure of the mind. But it does mean that when someone tells you a system is or is not close, you can ask the question that actually resolves things:

Which definition are you using, and what would change your mind?

If there is no answer to the second half, the first half was not a prediction.

What this page simplifies

Sources