All insights
AI & Search Intelligence
7 min read8 September 2026Nathan Mzumara

AI Models Have a 60-Day Shelf Life: What 261 Releases Show

AI Models Have a 60-Day Shelf Life: What 261 Releases Show

Three quarters of all AI models ever announced are already dead. Of 261 significant releases between June 2018 and September 2026, 197 are superseded or retired. That is 75.5%. Only 62 are still the model their maker is actually selling.

The median age of a model you can buy today is 60 days.

I built a workbook of every significant large language model released across 15 labs, then read the release history instead of the marketing. This piece is what that history says, in plain terms, and what it means if you are the person signing the contract.

The market replaces, it does not accumulate

Most software markets pile up. Old versions stay available. You can still buy something released three years ago.

This market does not work like that. When a new model arrives, the old one is usually withdrawn from the API, redirected to its successor, or quietly given a shutdown date in a changelog nobody reads.

So the 2023 flagship you saw compared somewhere is not a slower option today. It is gone.

And the replacement is speeding up. Across all 15 labs, the average gap between significant releases fell from 9.12 days in 2023 to 2.98 days in 2026. That is a 3.1-fold compression in three years, and it has not levelled off.

Why this matters to you: if it takes you longer than about two months to go from shortlist to signature, the shelf will have turned over before you sign. Not "there will be a newer option". The specific model you tested may not be the one you end up calling.

What is actually on the shelf right now

Sixty-two models are available today, either generally or on restricted access. Of those, 82.3% shipped in 2026.

The oldest survivor is Llama 4 Scout at 520 days. It survives for an unglamorous reason: it has open weights, so nobody can withdraw it. Not because anyone is still improving it.

Here is a detail that surprises most people. The lab with the most models available is not OpenAI, Anthropic or Google. It is Mistral AI, with nine. Mistral ships a wide catalogue of small, specialised models. The frontier labs ship a short ladder of large ones.

That gap matters, because "how many models does this vendor have" and "how good is this vendor" are two different questions with two different answers. Vendors quote whichever one flatters them. A long list is a catalogue strategy. A short list can be a focus strategy. Neither proves capability.

One more number for scale: the median context window on today's shelf is one million tokens. In 2022 the working figure was about two thousand. Capability did move. It just did not move in a way you can read off a model count.

Three things changed underneath the churn

The turnover is only interesting because of what changed while it was happening.

1. The plain text model stopped being made. In 2023 this field shipped 25 models whose category was simply "text". In the 250 days of 2026 so far, it has shipped zero. Meanwhile 60.5% of this year's releases accept more than text – images, PDFs, audio, video – and 74.1% list reasoning as a capability. In 2023, none did. The first reasoning model appeared in 2024, and there are now 110.

2. The cheap end got much cheaper. The expensive end got dearer. The cheapest capable model fell from $0.60 to $0.05 per million output tokens between 2023 and 2026 – twelve times cheaper. Frontier pricing went the other way, up 25%.

This is the most commercially important line in the whole dataset. Cheap intelligence is why an assistant now sits inside products where nobody would ever have paid for one. Expensive intelligence is where the frontier labs make their margin. Two markets moving in opposite directions, usually reported as one.

3. Some models now ship restricted. Seven models on the current shelf carry a category naming cyber capability, and seven were never generally released at all – trusted access or government-only terms. That is new. Anthropic and OpenAI both now ship models under restriction, and the first model graded critical on cyber capability by its own maker arrived this year. A field that restricts its own output has decided some of that output is dangerous, and that becomes your compliance question whether you asked for it or not.

How to choose between AI models without getting caught out

Four things, in order of how much money they save you.

Buy the capability, not the model. Put a swap clause in the contract. If a vendor cannot tell you what happens when your model is deprecated, you have not finished negotiating. Treat a 60-day median shelf life as your planning assumption, not a worst case.

Test the workflow, not the model. Public benchmark scores go stale as fast as the models do, which is to say in weeks. A test built around your actual task, with your actual data, survives a model swap. A spreadsheet of benchmark numbers does not.

Expect the floor to keep falling and the ceiling to keep rising. If your costs depend on cheap inference, they are improving. If they depend on frontier inference, budget for the opposite.

Watch the money underneath the prices. Disclosed funding across the private labs runs to roughly $356bn, about 12.4 times the revenue they have actually booked. Put plainly: for every $1 of revenue these labs have actually booked, they have raised about $12 of investment. Only 7 of 15 labs have an audited revenue figure at all. Today's prices are being paid for by investment, not income. The first commercial layer that does not depend on subscriptions has only just arrived: ChatGPT advertising reached a $1bn annualised run rate inside 200 days.

For a wider view of where the field is heading, Stanford's AI Index Report tracks capability, investment, adoption and policy in one place each year, and is the closest thing this industry has to an independent scoreboard.

Why most "best AI models" articles are out of date

Almost every model comparison you will read is a snapshot dressed up as a verdict. Written in a week, published, left alone.

At a 2.98-day average release gap, a comparison is materially out of date about a fortnight after it goes live. Many of the pages ranking today were written a year ago.

That is not the writers' fault. It is what covering a field at this speed does. But it does mean the useful thing is not the ranking. It is the method: what you measured, on what task, with what data, and how quickly you can run it again. Publish the method and the ranking becomes something you can regenerate. Publish only the ranking and you have written something with a two-week half-life.

Common questions about AI models

Here are the questions people most often search alongside AI models, answered from the same 261-release dataset rather than from general knowledge.

How many AI models are there? This workbook counts 261 significant releases from 15 labs since 2018. Only 62 are currently available. There are many thousands of fine-tuned variants beyond that, but 62 is the number that matters if you are choosing something to build on.

What are AI models, in simple terms? A model is a trained system that takes an input and produces an output. In this dataset almost all of them are large language models: you give them text, images, documents or audio, and they return text. The useful distinctions between them are the size of input they accept, whether they can reason step by step, whether you can hold the weights, and what they cost per million tokens.

How often are new AI models released? On average, one significant release every 2.98 days across the 15 labs tracked, as of 2026. In 2023 that figure was 9.12 days.

What is the best AI model? There is no stable answer, which is the honest one. With a 60-day median shelf life and a release every three days, any "best" list is a snapshot. Test two or three against your own task, keep the test, and re-run it each quarter.

Are AI models getting more expensive? Both. The cheapest capable model fell twelve-fold between 2023 and 2026, to $0.05 per million output tokens. Frontier pricing rose 25% over the same period.


Figures come from a workbook of 261 model releases across 15 labs, compiled from vendor documentation, launch posts, API release notes and Wikipedia, captured on 7 September 2026. Every figure was computed once and independently re-derived through a second code path before publication. Full method in the report, How AI Changed the Way People Search.

Tags

AI modelsmodel releasesprocurementLLM pricing

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.