Recency Is Close to a Filter: 86.3% of What ChatGPT Read Carried a 2026 Date
Of the 2,680 results ChatGPT retrieved across 60 software buying questions, 1,249 carried a machine-readable publication date. Of those, 90.4% were published within twelve months and 86.3% carried a date in the current year. Only 63 results – 5.0% – were more than two years old. Where a retrieved page declared its age, it was almost always recent.
That is a striking distribution, and it is easy to over-read. The honest version of this finding is more specific and more useful than "fresh content wins", so let me set out both the pattern and the two caveats that constrain it.
The distribution
| Age of retrieved result | Results | Share of dated results |
|---|---|---|
| 0–30 days | 212 | 17.0% |
| 31–90 days | 337 | 27.0% |
| 91–180 days | 389 | 31.1% |
| 181–365 days | 191 | 15.3% |
| 1–2 years | 57 | 4.6% |
| 2+ years | 63 | 5.0% |
The mode sits in the 91–180 day band. Content roughly three to six months old is the single largest slice of what the retrieval layer surfaced. Content over two years old is a rounding error at 5.0%.
The two caveats that keep this honest
First: 53.4% of retrieved results carried no date at all, and were evidently not penalised for it. Of 2,680 retrievals, 1,431 had no machine-readable publication date and were still pulled into context. So this is not a hard recency filter that excludes undated pages. Whatever the mechanism is, it does not require a date to admit a page.
Second: a "2026" date is a claim made by the publisher. It is not verified freshness. And the comparison-site long tail identified elsewhere in this study has an obvious incentive to restamp pages – those are exactly the sites publishing "Best ERP Software 2026" pages that get republished with a new date each cycle.
Put those together and the finding is narrower than "fresh content wins". The finding is that the retrieval layer strongly prefers pages that assert freshness. That is a slightly different and much more actionable statement.
Why the distinction matters operationally
If the mechanism were "recently created content ranks better", the implication would be to publish more. If the mechanism is "pages asserting a current date get surfaced", the implication is quite different: your existing pages need a visible, current date and a genuine refresh cycle, and the value of a page is not fixed at publication.
There is a defensible split here, and it turns on what the page is for.
An undated evergreen page competes on vendor terms. That is fine if you are the vendor and the page is the canonical source for a fact. Your pricing page does not need a publication date to be the primary source for your prices. In this corpus, pricing pages accounted for 278 of the 2,680 retrieved results and vendor-owned domains took 65.9% of all citations, largely on that basis.
Any page whose job is to be compared against others needs a visible, current date and a real refresh cycle, or it will lose retrieval to a site that restamps monthly. Best-of listicles, category landscapes, "alternatives to" pages, versus pages – all of these are competing directly with the comparison-site long tail, and that tail restamps.
That is the whole practical read. Two content classes, two different date policies.
How this interacts with what actually gets cited
Recency is a retrieval-stage effect. It governs which pages enter the context window, not which pages survive into the visible answer. The survival step is brutal and independent: 88.91% of unique URLs read never became a citation, and only 11.09% of read URLs and 20.9% of read domains appeared in the answer at all.
So a freshness programme buys you entry, not victory. It gets your page into the 44.7 results the model reads per answer. Whether it converts into one of the 6.1 citations rendered depends on whether the page states a checkable fact in a structured form – which is why vendor pricing pages convert at 16.4% and review aggregator pages, when they are surfaced at all, convert at 28.8%.
| Source type | Retrieval share | Citation share | Conversion |
|---|---|---|---|
| Vendor-owned | 55.1% | 65.9% | 16.4% |
| Independent comparison site | 9.3% | 14.4% | 21.3% |
| Review aggregator | 1.9% | 4.1% | 28.8% |
| Analyst / market research | 3.2% | 4.4% | 18.4% |
| Editorial media | 3.0% | 4.1% | 18.5% |
| Community / UGC | 1.2% | 0.0% | 0.0% |
Read that table alongside the recency data and the sequencing is clear. Freshness affects the first column. Structure affects the third.
A refresh policy that follows from the data
Date everything that makes a comparative claim. Visible on the page, in the markup, and in the title where it fits the format. "Best X software in 2026" is not a cliché in this context; it is a machine-readable claim about currency in a category where 86.3% of dated retrievals carried a current-year stamp.
Set a refresh cadence by content class, not by calendar. Comparative pages – versus, alternatives-to, best-of, category landscape – on a quarterly cycle at minimum. Canonical fact pages – pricing, feature availability, integrations, supported platforms – updated when the fact changes, and dated accordingly.
Do not restamp without substance. The observation that restamping works is not an argument for doing it dishonestly. It is an argument for having a refresh cycle real enough to justify the date. The sites winning this in the corpus are winning on format and cadence together, and one of the four I read directly had a published methodology and a dated review cycle.
Do not date-stamp your canonical fact pages defensively. A pricing page with a stale-looking date is worse than an undated one. 53.4% of retrieved results had no date and were not penalised, so undated is a viable state for pages that are the source of record.
Kill or consolidate content over two years old. Only 5.0% of dated retrievals were older than two years. If a comparative page has not been touched in that long, it is very close to invisible to this retrieval layer, and it is diluting the internal signal of the pages that are current.
FAQ
Does ChatGPT prefer recent content?
Strongly, where a date exists. Of 1,249 retrieved results carrying a machine-readable publication date, 90.4% were published within twelve months and 86.3% carried a current-year date. Only 63 results, 5.0%, were more than two years old.
Does content without a publication date get penalised by ChatGPT?
Not in this corpus. 53.4% of the 2,680 retrieved results carried no machine-readable date at all and were still pulled into context. Undated pages are not excluded, so the effect is a preference for asserted freshness rather than a hard filter.
How often should I refresh content for AI search?
By content class. Comparative pages - versus, alternatives-to, best-of, category landscapes - warrant a quarterly refresh with a visible current date, because they compete against comparison sites that restamp on a cycle. Canonical fact pages such as pricing should be updated when the fact changes.
Does a fresh date guarantee a citation?
No. Recency is a retrieval-stage effect and gets a page into the context window. Only 11.09% of unique URLs read ever became a citation. Converting a retrieval into a citation depends on whether the page states a checkable fact in a structured, extractable form.
Should old content be deleted for AI visibility?
Comparative content older than two years is close to invisible - only 5.0% of dated retrievals were that old. Consolidate or refresh it rather than leaving it live. Canonical fact pages are a different case and can remain undated indefinitely.
Related in this series
- How ChatGPT Shortlists Software Brands, the full report as a PDF
- Forty-Five Reads, Six Links: How ChatGPT Actually Cites Software Brands
- The Page Shapes AI Retrieval Actually Finds: 525 Versus Pages, 17 Reviews
- The Sites Beating G2 in ChatGPT: 43 Comparison Sites, 53 Citations, 3.5x the Aggregators
- Vendor Pages Take 65.9% of ChatGPT's Citations. G2 Takes 4.1%.
About the research. Nathan Mzumara is an organic growth and AI search practitioner. Publication dates were taken from the retrieval set in the ChatGPT response stream, not inferred from page content. Method and limitations are stated in the pillar report.
Tags