All issues
Discovery Digest · 25 September 2026

Issue 19. retrieval decides who gets recommended, agents start checking out & more

TL;DR

**Retrieval, not persuasion, now decides whether an AI engine names your brand: 2.8% mention rate without it, 91.4% with it.** Google logged the September 2026 spam update on 24 September and added multimodal Search reporting to Search Console the same day, so camera-led demand finally has numbers attached. Meanwhile Meta's Muse agent is negotiating and checking out with one-time cards, OpenAI put email, calendar and Slack behind ChatGPT Voice, and frontier token prices split in two directions at once.

Issue 19. retrieval decides who gets recommended, agents start checking out & more
01 · Google Search

Google logs the September 2026 spam update, global, all languages, up to two weeks

What
Google filed a single incident on the Search Status Dashboard: the September 2026 spam update, service affected Ranking, scope global and all languages. That is the entire public record. A spam update is not a core update. Core updates reassess how useful content is across the board, while spam updates change the systems that detect and demote behaviour breaching Google's published spam policies. The incident entry links straight to that spam policy documentation, which tells you where the diagnosis starts. There is no named target, no list of affected sites and no regional carve-out. Anything beyond the dashboard entry that you have read this week is inference dressed as reporting.
When
Logged at 09:15 US/Pacific on 24 September 2026, with a stated rollout of up to two weeks, so completion lands on or before roughly 8 October 2026.
How it shifts discovery
The two week window matters more than the start date. Movement during a rollout is partial, uneven and often reverses, which is exactly why most post-update analysis goes wrong: a reading on day three, a board briefing on day four, and a decision made on a picture that was never finished. The expensive mistakes I see are always mid-rollout ones. Pages get rewritten, redirects get fired, and afterwards nobody can tell whether the recovery came from the fix or from the update simply completing. What I'd do: freeze a dated baseline of impressions, clicks and average position for 24 September at segment level, not site level. Split by template, market and language, because a global trend line will hide the markets that actually moved. Then audit against the published spam policies rather than instinct, and hold all structural changes until the rollout is confirmed complete.
Questions to ask
  • Do we have a timestamped, segment-level baseline from 24 September, or are we about to argue from memory?
  • If we lost visibility, can we point to a specific spam policy breach, or are we assuming this is a quality problem?
  • Who has authority to block structural changes until the rollout completes?
Sources
02 · AI agents

Meta launches Muse, an agent that browses, negotiates and checks out

What
Muse is a personal AI agent that runs on its own dedicated cloud computer, the Muse Secure VM, and works across apps a person connects. It plans, opens a browser, fills out forms, negotiates on the user's behalf, persists after the app is closed, and returns for approval before sending an email or making a purchase. It is powered by Muse Spark, which Meta calls its most capable model to date. It can check out using Link built by Stripe, and Meta says it is the first AI agent covered by Link's purchase protections, which include damaged or lost items, price drops, no-fee returns and a return guarantee on eligible purchases. Link's wallet for agents issues a one-time-use card so the buyer's real card details stay hidden. Shop Pay and 1Password support are listed as coming soon. Meta states Muse conversations and VM data are not shared with its ad systems.
When
Announced 8 September 2026, rolling out in the US on iOS, Android and muse.ai, with AI glasses described as coming soon.
How it shifts discovery
Two of the three outcomes Meta showcases, selling a car for more and lowering a bill, are negotiations conducted against a business. Automated haggling is now a shipped feature, not a lab demo. The quieter problem is the one-time card. If your checkout sees a fresh card on every order, card-based repeat customer matching, stored-card loyalty and fraud rules that reward a known card will all misread agent purchases as new, risky buyers. Price-drop cover and no-fee returns also lower the cost of buying early and returning, which will move your return rate before it moves your revenue. What I'd do: ask your payments and fraud owners this week whether one-time-use cards trip your risk rules, and change your customer identity key from card to email or account ID. Then check that your product pages carry the specification detail an agent needs to decide, because navigation and hero imagery are being skipped.
Questions to ask
  • Does our fraud stack decline or friction one-time-use cards, and how many orders would that cost us?
  • Is our repeat-customer logic keyed on a stored card rather than an account identifier?
  • Can an agent extract dimensions, materials, stock and returns terms from our PDP without rendering our design?
Sources
03 · Analytics and measurement

Search Console adds web multimodal reporting, so camera-led demand finally has numbers

What
Google added a multimodal search type to Search Console, covering searches where the query is an image rather than typed text. Four surfaces roll into it: Google Lens, Circle to Search on Android, image uploads to Google Search, and Chrome's right-click 'Search this image'. The data appears in two places, the report on performance on Search results and the report for Generative AI features, and you access it through the search type filter. Until this release those visits sat inside your general Search totals with no way to separate them. Google says metrics only appear if your site is already receiving traffic from these queries, so an empty report is itself a finding. The change was announced on the Google Search Central Blog by Harsh Kharbanda, Product Manager Lead for Google Lens, and Moshe Samet, Product Manager Lead for Search Console.
When
Announced Thursday 24 September 2026, rolling out globally from that date.
How it shifts discovery
This sounds like a minor reporting tweak. It is not, because it measures a demand channel that no keyword tool can see. There is no phrase, no match type and no query the user had to think of first, which means the image a person photographed is effectively the query. That promotes product shots, packaging and in-situ photography from decoration to entry point, and it makes image alt text, file naming and structured product data commercial assets rather than accessibility housekeeping. What I'd do: open the Performance report, apply the multimodal filter, export the data and compare it against your text query set for the same period. Then check the same filter inside the Generative AI features report to see where multimodal and AI surfaces overlap. If multimodal traffic is landing on pages with thin or generic imagery, that is your first optimisation queue.
Questions to ask
  • Which pages already receive multimodal traffic, and do their images match what a buyer would realistically photograph?
  • Are our product and packaging images unique, or are we publishing the same supplier shots as every competitor?
  • How much of our AI feature traffic is multimodal, and is anyone reporting that separately to the board?
Sources
04 · AI search

ChatGPT Voice reaches email, calendar and Slack, and ships finished files

What
OpenAI updated ChatGPT Voice in three parts. Voice can now use plugins including email, calendar and Slack; it can be powered by GPT-6 Astra, Sol and Luna; and it can be used in ChatGPT Work on web and mobile, where you can create docs, decks, sites and spreadsheets or tackle complex browser tasks by talking. OpenAI's announcement opens with 'We heard you loud and clear', framing the release as a response to user pressure. No pricing or plan detail was given, only that the rollout is global in the latest version of the app. One open question went unanswered in the thread: whether GPT-6 Sol and Luna are coming to Chat itself rather than only ChatGPT Work and Codex.
When
23 September 2026, rolling out globally the same day.
How it shifts discovery
Email, calendar and Slack are not content surfaces. They are systems of record and coordination, so once a spoken request can reach them the model stops retrieving a page and reading it aloud and starts doing things inside your stack. Putting the newer model tier behind speech also raises the chance that one spoken answer is treated as sufficient, which is the quiet threat to click-through from voice-led discovery. And a voice session that used to end in words now ends in a deliverable. What I'd do: test how your brand is described when a spoken query is answered by GPT-6 tier models, because voice answers are shorter and less forgiving than text ones. Then talk to IT about which plugins are enabled in your workspace, since Slack and email access changes your data governance posture, not just your marketing one.
Questions to ask
  • What does a voice-only answer about our category actually say, and are we in it?
  • Which connected apps has our workspace authorised for Voice, and who approved that?
  • If a voice session ends in a finished deck or spreadsheet, does our content still get cited, or just consumed?
Sources
05 · AI search research

Brand visibility in AI search: 2.8% without retrieval, 91.4% with it

What
A new arXiv paper by Benjamin Tannenbaum, 'From Prompt to Recommendation: A Fitted Stage Model of Brand Visibility in AI Search' (arXiv:2609.23162, Information Retrieval), analyses 34,960 unbranded prompt-engine observations from 75 anonymised Aiso projects, covering 2,854 distinct monitored prompts across repeated GPT and Gemini runs. Two retrieval gates decide almost everything: own-domain exposure, meaning your domain appears as a citation in the observable live retrieval path, and branded fan-out, meaning the engine issues a follow-up search containing your brand name. With neither, mention rates are 2.8% on GPT and 3.8% on Gemini. With own-domain citation but no fan-out, 49.0% and 58.4%. With both, 91.4% and 100%. Within-prompt checks, holding fan-out absent, show own-domain exposure associated with a mean mention-rate lift of 40.2 points on GPT and 49.0 on Gemini. Prior visibility persists too: a previous non-mention with no current exposure gives next-run mention rates of 1.6% and 1.9%, while a previous mention with current exposure gives 80.5% and 83.7%.
When
Submitted to arXiv on 19 September 2026. Observations were collected between June and September 2026.
How it shifts discovery
This is the clearest evidence yet that AI visibility work is retrieval engineering, not copywriting. If your domain is not in the retrieval path you are playing for roughly a one-in-thirty-five chance of being named, and no amount of persuasive brand language changes that. The persistence finding is the uncomfortable part: visibility compounds, so absence compounds too. What I'd do: audit crawlability and citation-worthiness for the pages that answer your top unbranded prompts, since those are the queries where the user never types your name. Track own-domain citation presence as a leading indicator in your AI visibility reporting, above mention rate, because it is the gate that moves mention rate. Then treat branded fan-out as the second objective: it depends on the engine already associating your name with the category, which is an entity and coverage problem, not a page problem.
Questions to ask
  • For our top 50 unbranded prompts, how often does our own domain appear in the retrieval path?
  • Are we reporting mention rate alone, when own-domain citation is the upstream lever?
  • Which competitors are already in the compounding loop of prior mention plus current exposure?
Sources
06 · AI economics

LLM API pricing: the spread between tiers now matters more than the average

What
BenchLM's frontier token price index stood at 18.7 on 22 September 2026 against a March 2023 base of 100, a fall of 81.3%. The same snapshot shows the index up 16.9% month on month. Both are true because the index tracks the median blended price across active frontier models, at three parts input to one part output, so it moves when the cohort changes and not only when providers cut rates. Five rows moved in opposite directions: GPT-5.6 Luna fell from $2.25 to $0.450 blended, GPT-5.6 Sol from $11.25 to $8.00, GPT-5.6 Terra from $5.63 to $4.50, while Claude Opus 5.5 entered at $8.00 blended and Claude Fable 5.1 entered at $20.00. The gap between the dearest and cheapest model in the table is roughly 600 to 1 on output price, from GPT-5.4 Pro at $180 per million down to Gemini 1.5 Flash at $0.30.
When
Snapshot published 22 September 2026, covering 511 models and 454 benchmarks.
How it shifts discovery
Most budget models assume a smooth decline and then get repriced upward because the definition of frontier moved, not because anyone raised a price. My read is that model choice is now a bigger lever on unit economics than any discount you are realistically going to negotiate. A 44x blended gap inside a single snapshot means routing decisions, not procurement, decide your bill. What I'd do: split your AI workloads into cheap-tier and frontier-tier buckets and be honest about which genuinely need the frontier. Bulk classification, extraction, clustering and first-pass drafting rarely do. Then rebuild your cost model on the monthly series rather than on a single assumed trend, and set a review trigger for when a new model enters your tier.
Questions to ask
  • What percentage of our token spend runs on a frontier model that the task does not require?
  • Is our forecast built on an assumed price decline that the index does not support?
  • Who owns model routing, and how quickly can we switch tiers when prices move?
Sources
07 · AI economics

GPT-6 Sol and Luna land at half the price of the 5.6 tier

What
OpenAI expanded the GPT-6 family with GPT-6 Sol and GPT-6 Luna, cutting API prices by 50% against GPT-5.6 Sol and GPT-5.6 Luna promotional pricing. GPT-6 Sol is $2 input and $10 output per million tokens, down from $4 and $20. GPT-6 Luna is $0.10 and $0.50, down from $0.20 and $1.20. OpenAI attributes the cut to caching and inference improvements passed on to customers, and says Sol is trained with similar methods to GPT-6 Astra, bringing Astra's advances in professional work, factuality, coding, computer use and alignment into the cheaper tier. Astra remains OpenAI's best model overall. On AutomationBench 1.0.6, which tests agents on end-to-end business workflows using 47 tools, GPT-6 Sol at xhigh effort scores 33.2% at $0.27 per task, outperforming Claude Opus 5 at max effort at 9% of its cost per task.
When
No publication or availability date is stated in OpenAI's announcement, which only notes GPT-6 Astra was introduced 'earlier this month'. Pricing is API only, with no consumer plan described.
How it shifts discovery
The benchmark table is not the story. The story is that a class of marketing work which was too expensive to run at scale, things like full-catalogue content audits, per-SKU description generation or continuous AI visibility monitoring, just got roughly twice as affordable. Cost per task, not raw score, is now the right comparison for routing decisions. What I'd do: reprice every token-based workflow you costed last quarter, because the business case for the ones you shelved may now clear. Note the 90% cached-input discount when you model repeated prompts against a stable corpus, since that matters more than the headline rate for monitoring workloads. And keep model names exact in finance models, docs and prompts: it is GPT-6 Sol, not GPT Sol, and retrieval systems match entities on exact strings.
Questions to ask
  • Which shelved AI projects clear the business case at half the previous token cost?
  • Are we structuring prompts to benefit from cached input on repeated monitoring runs?
  • Do our internal docs and finance models still reference superseded GPT-5.6 pricing?
Sources
08 · AI models

Grok 4.7 arrives at Grok 4.6 prices, and flat frontier pricing is the real news

What
SpaceXAI released Grok 4.7, calling it its most capable model for coding and knowledge work, served at $2 per million input tokens and $6 per million output tokens, identical to Grok 4.6. A faster variant with twice the output speed costs twice the price. Three things changed underneath: a new, larger base model, a longer reinforcement learning run on a harder task mix weighted toward multi-hour problems, and native training on the Grok Bot harness. Availability on day one covers Cursor, Grok Build, the Grok API, third-party coding harnesses, and model routers and cloud platforms, with Grok Build free to try. Red-team capabilities are invite-only. On published benchmarks, Grok 4.7 xHigh scores 46.3% on CursorBench 4.0 against Grok 4.6 High at 40.4%, 64.0% on EEBench against 53.0%, and 38.0% on Terminal-Bench 4.0 against 20.3%.
When
Released 21 September 2026.
How it shifts discovery
For search and growth teams the interesting line is not the coding scores, it is native understanding of the Grok Bot harness. A model tuned for the conversational harness is a model tuned for the surface where consumers actually ask questions, which is where your brand mentions live or die. The pricing point matters just as much: at $6 per million output tokens, unchanged across a capability jump, the cost of running frequent AI visibility measurement stops being the constraint. Compare that to $20 output on GPT-5.6 Sol Max and $50 on Fable 5.1 Max. What I'd do: if you have been sampling AI visibility monthly because measurement was expensive, move to weekly on a cheap frontier tier and keep the prompt set fixed so the series is comparable. Add Grok to your monitored engine list if consumer discovery in your category touches it.
Questions to ask
  • Is Grok in our monitored engine set, or are we only tracking GPT and Gemini?
  • How often could we measure AI visibility if cost per run fell by two thirds?
  • Are our benchmark comparisons using cost per task, or raw scores in isolation?
Sources
09 · Paid and AI search

OpenAI begins testing ChatGPT Sponsored Agents, plus HubSpot and Shopify integrations

What
OpenAI introduced four separate things at once, and they are at very different stages. Sponsored Agents let a user click an ad in ChatGPT and open a clearly labelled conversation with a business-sponsored agent, asking follow-up questions before following a link to the site. OpenAI states that conversation is distinct from ChatGPT's independent answers and separate from the user's original conversation, and that its ads principles are unchanged. Sponsored Agents are limited to select advertisers in the United States. Three campaign management changes are rolling out now rather than in test: an Ads Manager plugin inside ChatGPT for creating, updating and analysing campaigns by prompt; AI assistance inside Ads Manager suggesting copy and imagery based on your landing page and objective, with advertiser review before anything is added; and opt-in AI-powered text customization that adapts existing headlines and descriptions to conversational context and auto-translates into the user's preferred language. HubSpot is the first CRM partner and Shopify the first ecommerce partner.
When
Announced 16 September 2026.
How it shifts discovery
The partnerships are the real news, because they are a distribution decision rather than a feature. Putting ChatGPT Ads inside HubSpot and Shopify places it where mid-market marketing teams already work, which is how ad platforms win share without winning an argument. The pattern elsewhere is familiar: the platform absorbs production, and your remaining levers are the brief, the landing page and the objective. That makes landing page quality a paid media input, not just an organic one, since the AI drafts from it. What I'd do: before enabling text customization, decide what claims and phrasing you will not allow a model to rewrite, and document it. Audit the landing pages that would feed ad generation, because weak pages now produce weak ads automatically. And press for measurement detail: the announcement leaves the reporting and control questions on Sponsored Agent conversations largely open.
Questions to ask
  • If AI drafts our copy from our landing page, is that page accurate enough to be a source of truth?
  • What claims, pricing or regulated language must never be auto-adapted or auto-translated?
  • How would we attribute a conversion that starts in a sponsored conversation rather than a click to site?
Sources
10 · Analytics and measurement

Google Analytics publishes a Q4 measurement checklist for pre-peak conversion tracking

What
Google Analytics is pushing teams to audit their GA setup, key events and conversion signals before the Q4 peak trading period. The guidance points at Help Center documentation on editing and managing key event setup, which is the mechanism: key events are what feed your conversion reporting and, downstream, your attribution and any bidding that consumes those signals. The practical scope is unglamorous and familiar. Duplicate or misnamed key events, events that were created for a test and never retired, and conversion definitions that no longer match how the site actually works after a year of releases.
When
Published ahead of the Q4 2026 peak season.
How it shifts discovery
This lands in the same fortnight as a spam update rollout and a new multimodal search type, which is the point. Your reporting is about to be scrutinised during the highest-attention trading weeks of the year, at exactly the moment your organic numbers are moving for reasons outside your control. If your key events are messy, you will spend December arguing about the data instead of acting on it. What I'd do: this week, list every key event in the property and mark each one keep, rename or retire. Check that each still fires on the current template, not the one it was built for. Then reconcile GA conversion counts against your order or CRM system for a recent seven day window and document the variance now, before peak, so you have an agreed margin of error to work from rather than a December surprise.
Questions to ask
  • When did we last reconcile GA key events against the order or CRM system of record?
  • Which key events exist only because someone ran a test and never cleaned up?
  • Do our conversion definitions still match the current checkout and lead flows after this year's releases?
Sources

Key takeaways

What to walk away with this week

  1. Retrieval is the gate. Own-domain citation in the live retrieval path lifts GPT mention rates from 2.8% to 49.0%, and adding branded fan-out takes it to 91.4%. Report citation presence, not just mention rate.

  2. Do not conclude anything about the September 2026 spam update before roughly 8 October. Freeze a dated, segment-level baseline from 24 September and audit against published spam policies rather than content quality instinct.

  3. Apply the new multimodal filter in Search Console this week. If Lens and Circle to Search are sending visits, your product imagery is the query, and an empty report is still a finding.

  4. Agent checkout breaks card-based identity. One-time-use cards from Link mean repeat customer matching, stored-card loyalty and fraud rules will misread agent purchases unless you key on account ID or email.

  5. Model choice now moves unit economics more than procurement does. With a 44x blended price gap inside one snapshot and GPT-6 Sol at half the 5.6 rate, reprice the AI workflows you shelved last quarter.

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.