Grades measure the wrong thing once intelligence is abundant
For more than a century, grades have been education's main unit of account. They record how well a student performed on a task the institution chose, under conditions it controlled, which made sense when knowledge was scarce and holding it was a fair proxy for using it. AI is removing that scarcity. When a model can write the essay and pass the bar exam, a grade says less about the student and more about which tools were allowed in the room.
Grades also measure a stock: what a person knows at one moment, against a fixed curriculum. As we approach AGI, valuable capabilities shift faster than curricula can be rewritten, and much of what a person knows can be summoned on demand. What matters now is what a person can do with machine intelligence, how reliably they can do it, and how quickly they can learn to do something new.
Time to Capability measures exactly that with one question: how long does it take a person, starting without expertise, to reliably accomplish a meaningful real-world objective? Time is a unit everyone understands, and it resists the inflation that has hollowed out grades. Timing the same capability with and without AI makes the effect of AI visible, turns agency into something observable, and gives policy a clear target: shrinking the time and cost to capability for every learner, starting with the poorest.
Agency becomes measurable once you put a clock on it
None of education's familiar numbers, whether enrollment, years completed, credentials awarded, or test scores, tells you whether a person can go out into the world and make something happen.
Time to Capability (TTC) asks a narrower and more useful question. How long does it take the median person, starting without the relevant expertise, to become capable of accomplishing a meaningful real-world objective safely and reliably?
The question moves the finish line. A credential is not the finish line, and neither is mastery. The finish line is the first moment a person can independently produce a real-world outcome and keep producing it. Agency stops being an aspiration and becomes an elapsed time that anyone can observe, compare, and try to shrink.
The timing matters because AI changes the clock unevenly. Some capabilities that used to take years may now take weeks, while others barely move. A good metric should show both, and it should show which parts of the pathway AI actually compressed.
Capability is a gate you pass, not a score you accumulate
TTC needs a threshold, or it collapses into a measure of how fast a machine can generate something impressive. A person who generated one working app with Claude once is not yet capable of building software. Capability means that, given a new problem in the domain, the person can repeatedly produce acceptable outcomes, notice failures, recover from mistakes, and recognize when they need outside expertise.
Four dimensions define the threshold:
These four dimensions should work as gates rather than as a weighted average. An outcome that is brilliant nine times and dangerous the tenth does not average out to "mostly capable." The logic mirrors the Core Gates of the Abundance Index, where every gate must clear before a domain counts.
TTCij = the first time t at which Cj(i) = 1 on a task variant person i has never seen The thresholds q*, r*, a*, and s* are set in advance by working practitioners for each objective j.
The thresholds belong to practitioners, not to learners or tool vendors. A panel of working professionals should define, before any learner starts, what an acceptable outcome looks like for each objective. The "never seen" condition matters just as much, because it separates capability from rehearsal.
Traditional pathways are slow because they are bundled and queued
For most of history, we treated the educational pathway itself as a proxy for capability. The sequence ran from school to degree to apprenticeship to junior work, and only then to professional capability. Each stage was long because each was scarce: teachers, labs, mentors, and trusted credentials were all rationed.
TTC becomes more useful once you break it into the terms that make it up:
- Taccess
- is the time it takes to get admitted, relocate, or afford the pathway at all.
- Tprereq
- is time spent on material taught in advance of any need for it.
- Tinstruction
- is time spent receiving explanation.
- Tpractice
- is the repetitions the person must perform personally, including physical ones.
- Tfeedback
- is the latency between attempting something and learning what went wrong.
- Tcredential
- is the time a regulator or institution requires regardless of demonstrated ability.
- Tqueue
- is cohort pacing, semester calendars, and waiting for a mentor's attention.
A four-year computer science degree contains a great deal of valuable material, but only a fraction of its hours go into building software that real people use. Much of the rest goes to prerequisites taught ahead of need, cohort pacing, waiting for graded feedback, and the institution's calendar. Traditional TTC is long less because the capability is intrinsically hard and more because the pathway is bundled and queued.
AI attacks these terms selectively. Patient tutoring compresses instruction. Instant critique compresses feedback latency. Just-in-time explanation shrinks the prerequisite block, and near-zero marginal cost shrinks access. For many cognitive tasks, an agent can also carry out part of the execution, which lowers the bar the person must personally reach. AI does much less to practice when the practice is physical, and almost nothing to credential time when a regulator sets the clock.
Compression is uneven, and the unevenness is the finding
Press play to run the clock on ten capabilities. Grey bars follow the traditional pathway, and blue bars follow an AI-leveraged pathway to the same capability threshold. The logarithmic axis spreads out the short pathways. Select any row to see its stages.
All durations are order-of-magnitude hypotheses for a motivated median learner, drawn from typical program lengths. They exist to be replaced by measured values from the benchmark battery described below.
Three bands appear. Symbolic and analytic capabilities, such as software, market analysis, and bounded legal research, compress by close to an order of magnitude or more. Creative and interpersonal capabilities compress by roughly three to ten times, because taste and live conversation still demand the person's own repetitions. Regulated and embodied capabilities, such as clinical practice and electrical work, compress the least, because motor learning, case volume, and licensure each run on their own clocks.
That spread could become an index in its own right. A compression map of the economy would show where AI is turning years into weeks and where it is not. Learners, universities, and policymakers need exactly that map when they decide what to teach, what to fund, and which credential clocks to revisit.
Every capability has two clocks
The framework already separates unaided capability, which is what I can accomplish myself, from AI-leveraged capability, which is what I can accomplish with machine intelligence. TTC should measure both.
An eighteen-year-old learning to code might become economically capable in three months with AI, even though independent expertise still takes two years. That is not fake capability. A modern pilot's capability includes the avionics, a carpenter's includes power tools, and a mathematician's includes computation. AI becomes part of the person's extended cognitive system.
TTCU still matters because it sets the oversight floor, which is the minimum unaided understanding a person needs to recognize when the machine is wrong. Below that floor, AI-assisted output looks like capability but fails the reliability and safety gates. I call this counterfeit capability. It is the autopilot problem applied to learning: the system works until the moment it doesn't, and the person has no way to notice.
AI shortens the two clocks through different mechanisms. As a tutor, it raises the learning rate, which shortens TTCU. As a tool, it lowers the unaided level a person must reach before the combined system clears the threshold, which shortens TTCAI further. The oversight floor keeps the second effect honest.
The two clocks, simulated
Adjust how much the tool can do, how good the tutoring is, and how much of the capability has to live in the person's own body and judgment. The shaded band is output the learner cannot yet verify, so it does not count toward capability.
Model: unaided capability grows as 1 − e−kt. AI adds L·(1−E) of the remaining gap, scaled by min(1, unaided capability ÷ oversight floor) so that unverifiable output does not count. The model is a thinking tool, not a forecast.
Two results stand out when you move the sliders. Raising tool leverage without raising unaided understanding produces an impressive raw-output curve and only a modest early change in verified capability. Raising tutoring quality shortens every clock at once, because verification depends on understanding. The strongest pathway improves both together, which argues for redesigned pedagogy rather than tool access alone.
The compression ratio shows what AI did to a domain
Once TTC exists for both pathways, their ratio becomes the headline number.
A statement like "AI has produced a 20× compression in the time an ordinary person needs to reach this capability threshold" says far more than "87% of students have access to an AI tutor." Access is an input. Compression is an outcome.
Running three experimental arms lets you split the ratio into its sources:
The tutoring effect measures how much faster people learn. The tooling effect measures how much less they need to learn before the combined system clears the bar. A domain where tutoring dominates is producing durable human capability. A domain where tooling dominates deserves closer attention to oversight, because more of its capability lives outside the person.
The leverage gap, TTCU minus TTCAI, is worth tracking across cohorts as well. A shrinking gap means people are catching up to their tools. A widening gap means the tools are outrunning the people who use them.
Agency is the rate of change, not the stock
Today we define an educated person by the capabilities that person already has. When valuable capabilities keep shifting, a more important property may be how quickly a person can acquire a capability they do not yet have.
Consider two graduates. Person A knows a hundred things extremely well but needs two years to become useful in an unfamiliar domain. Person B knows fewer things at the start but can meet an unfamiliar problem, learn its fundamentals, use AI intelligently, find experts, run experiments, and become operationally capable in about six weeks.
If the world never changes, Person A wins, because the stock keeps paying. Once the frontier moves, every shift resets the game, and Person B resets faster. The simple model below suggests the crossover arrives sooner than intuition says: stock wins only in a nearly static world, and even then narrowly. The deciding variable is the half-life of valuable knowledge, and AI is shortening that half-life.
Stock versus velocity over ten years
Each vertical line marks a shift in the frontier, when the valuable capability moves to a new domain. Person A carries a deep stock into the first domain and learns slowly. Person B starts shallow and learns fast. Stretch the half-life toward ten years to find the point where stock still wins.
The measurable version of this idea is a personal TTC profile: the median TTC across a battery of objectives the person has never attempted before. The profile also has a second derivative. A person who has learned how to learn should reach capability faster in the fifth unfamiliar domain than in the first. Education could then aim at the rate at which humans expand their capability frontier, and it could show evidence that it had succeeded.
The question stops being what a person knows. It becomes how quickly that person can become able to do something new.
Distance to capability is the abundance metric
The essay asks one question that I would elevate into the Abundance Index itself: how far is she, in time and resources, from capability?
Someone born in rural Kenya in 2000 who wanted to become capable of software engineering was effectively ten years, $50,000, and a relocation away. By 2030, the same capability may sit six months, about $100, and an internet connection away. That shift is educational abundance, and a count of years of schooling cannot see it.
Dividing cost by local income makes the metric equity-aware. The same $3,000 is a rounding error in one place and nearly two years of income in another. Counting barriers separately keeps visible the gates that money alone does not open.
Inside the Abundance Index, DtC fits naturally in the Education and Knowledge domain. One candidate Core Gate would require the median DtC on the benchmark battery to fall below a set level for learners in the bottom income quintile, rather than for the median learner alone.
Distance to software capability, 2000 to 2030
Move through the eras to watch the distance shrink for two learners. Blue is learning time, grey is cost measured in months of local income, and amber is the delay added by structural barriers.
Figures are illustrative round numbers chosen to show the shape of the metric. The 2030 values are projections. Replace them with measured cohort data before citing.
The grand objective follows directly. Drive the distance between human intention and human capability toward zero, without sacrificing judgment, autonomy, safety, or understanding.
A benchmark battery turns TTC into an experiment
A question like "how long does it take to become a lawyer?" is contaminated by credentialing and regulation. A cleaner approach defines real-world capability challenges, each with a threshold that practitioners set in advance.
| Challenge | Starting condition | Capability threshold |
|---|---|---|
| Build | No programming experience | A working application, built and deployed, that ten real people use; rebuilt on a new spec without help. |
| Research | No domain knowledge | A synthesis of a scientific question that blinded domain experts rate as accurate and appropriately uncertain. |
| Enterprise | No business experience | A genuine customer problem identified, a solution built, and a first paying customer secured. |
| Language | Beginner in the target language | A spontaneous 30-minute conversation with a native speaker on an unscripted topic. |
| Finance | No financial training | An accurate, defensible plan for a real household's situation, reviewed by a certified planner. |
| Civic | An unfamiliar policy question | The strongest arguments on multiple sides, explained and weighed against the evidence, as judged by partisans of each side. |
| Physical | No trade experience | A previously unseen mechanical failure diagnosed and repaired safely. |
Learners should be randomly assigned to three arms: traditional tools, AI tools, and AI with redesigned pedagogy. Each arm is measured in hours until the learner crosses the predetermined threshold. Active hours and calendar time should be logged separately, because calendar time captures the access and queue terms that active hours hide. A practice like the 500-hour challenge, where people track their hours of active AI use, already produces the right kind of log.
TTC is a time-to-event measure, so it should use the statistics built for time-to-event data. Survival curves such as Kaplan–Meier estimates handle learners who have not yet crossed the threshold when the study ends. Reporting only the median of those who finished would flatter every arm. A sound report gives the 25th percentile, the median, the 75th percentile, and the share who never cross.
Three probes keep the measurement honest. An unaided probe removes AI at regular intervals and retests, which traces TTCU alongside TTCAI. A transfer probe uses a task variant the learner has never seen. A retention probe repeats the test after ninety days, because capability that decays within a season was never capability.
TTC has to be built to resist its own gaming
Any metric that matters will be optimized, so the design should anticipate how. One-shot demos are the most obvious risk, and the reliability gate and transfer probe exist to defeat them. Threshold drift is subtler: as AI improves, a fixed threshold becomes trivially easy, so thresholds should be anchored to real outcome value and rebased periodically, the way any index is rebased.
Deskilling is the risk I would watch most closely. If TTCAI falls while TTCU rises across cohorts, the system is producing dependence rather than capability. Selection bias is another risk, because motivated volunteers make every arm look good, so randomization and intent-to-treat reporting matter. Task pools should rotate with held-out variants so that nobody can memorize the benchmark. Finally, averages hide the people who never make it, so every report should include the 90th percentile and the bottom income quintile.
| Measure | What it captures | Unit |
|---|---|---|
| TTCU | Time until the person clears the threshold alone | hours or months |
| TTCAI | Time until the person, working with AI, clears the threshold with verified output | hours or months |
| CCR | How much AI compressed a domain, split into tutoring and tooling effects | ratio |
| Leverage gap | How far the person's own understanding trails the combined system | hours or months |
| CAV | How fast a person moves from novice to threshold in new domains, and whether that speed rises | capability per month |
| DtC | How far a specific learner sits from capability in time, cost, and access | month-equivalents |
The best education metric measures generative capacity
The ultimate education metric may not be the percentage of people who possess capability X. It may be how quickly an ordinary person can become capable of achieving a new worthwhile objective. That second question measures the generative capacity of a person, which is the ability to become more capable, rather than the stock of knowledge that person happens to hold today.
TTC gives the framework a sharp bridge from AI to education, from education to capability, and from capability to agency and abundance. It is observable, it is experimental, and it points policy at the right target: the distance between what a person intends and what that person can do.
About
Time to Capability is a concept developed by Max Song and Abundance Education. It draws on the thinking and writing of Peter Diamandis.