Issue 11. AI discovery crosses the price line, and GA4 finally makes it countable
TL;DR
OpenAI cut Luna 80% and Anthropic shipped Opus 5 at last year's prices, which means always-on conversational agents and long-running discovery workflows just cleared the cost hurdle. At the same time GA4 finally made AI-referral traffic legible with a new Source Group dimension, so you can measure the channel you could not see before. This week is about the economics of AI search flipping, and the tooling catching up to prove it.
OpenAI cuts Luna 80%, pushing AI agents to the front of the funnel
- What
- OpenAI slashed GPT-5.6 Luna prices by 80% and Terra by 20%, and the cuts also apply to how usage is counted inside Codex and ChatGPT Work. Luna is not a stripped-down chatbot: it uses tools and completes multi-step workflows, which is exactly what a customer-facing shopping agent needs to check stock, compare specs and hand off to checkout. The live proof is a 24/7 multilingual retail agent (avatarin and Yamada Denki) that served 30,000 shoppers in two weeks with 92% positive responses.
- When
- 30 July 2026.
- How it shifts discovery
- This moves real-time conversational AI across the price threshold where it stops being an internal efficiency tool and becomes a viable customer-facing discovery channel. Front-of-funnel discovery is a high-volume, low-margin game, and at Luna's new price the unit economics finally clear. Map one high-volume customer question your team fields daily, route it through Luna, reserve more capable models only for steps where extra intelligence changes the outcome, and instrument the agent with the same attribution rigour as any other channel.
- Questions to ask
- Which single high-volume customer question could we route through Luna first?
- Do we have attribution in place to treat a live agent as a discovery channel, not just support?
- Where does multilingual coverage collapse a barrier to markets we do not staff for?
- Sources
Claude can now record a task and replay it as a reusable skill
- What
- Anthropic added a 'Record a skill' capability to Claude Cowork that lets you demonstrate a workflow once and have Claude save it as a named, durable, shareable skill. The shift is from one-off prompts that die in a chat window to standing procedures your team can rerun, tweak and hand off. That quietly closes the gap between prompting and actual process automation.
- When
- 30 July 2026.
- How it shifts discovery
- The real unlock is not speed, it is standardisation. When a schema audit or a campaign QA check runs the same way every time, you can trust the output and audit it later. Treat skills like code, not notes: name and document each one, assign an owner, review on a schedule as search rules move, and log which skill produced which output. Pick one repeatable workflow (content briefs, reporting pulls, schema audits) and record it first, then pair the library with real monitoring because reusable skills repeat without oversight.
- Questions to ask
- Which workflow do we repeat most often and lose to inconsistent prompting?
- Who owns each recorded skill and keeps it current as guidelines change?
- How do we monitor and audit skills that run without a human triggering each one?
- Sources
GA4 makes AI referral traffic countable with Source Groups
- What
- Google Analytics launched a new Source Group dimension and hostname filters. Source Groups roll fragmented values like 'chatgpt.com', 'chat.openai.com' and 'perplexity.ai' into clean, platform-level categories, while hostname filters strip out spam and unauthorised domains that distort conversion data. This is the first native GA4 signal built with emerging AI traffic in mind, and crucially it applies retroactively.
- When
- 28 July 2026.
- How it shifts discovery
- AI-assistant referral traffic has been growing but nearly impossible to measure cleanly, so most teams treat it as noise inside 'referral' or 'direct'. Now you can see AI discovery as a coherent category and compare it against organic, paid and social on equal terms, which is the foundation for real budget decisions. Audit how GA4 mapped your sources, apply hostname filters to whitelist legitimate domains, re-run the last 6 to 12 months to see how much 'direct' was actually AI discovery, then rebase your attribution and stop under-investing in a channel you could not previously see.
- Questions to ask
- How much of our historic 'direct' and 'referral' traffic was actually AI discovery?
- Are AI assistants mapped into the right Source Groups before we report upward?
- Which spam hostnames are inflating our sessions and conversions right now?
- Sources
Meta splits Marketplace selling into its own Seller app
- What
- Meta took selling on Facebook Marketplace out of the main app and launched Seller, a dedicated app with AI listing creation, a unified buyer inbox, inventory management and performance insights. Upload photos and Meta AI fills in the title, description, price suggestion and category, with a bulk feature to create multiple listings at once. Marketplace lists 430 million items every month globally, and Seller syncs so existing listings, messages and history carry over.
- When
- 24 July 2026 (App Store, US users aged 18 and over; web and Android in testing).
- How it shifts discovery
- The strategic point is discovery fragmentation: product visibility no longer sits neatly in Google Shopping plus your site, it is splintering across purpose-built apps. That creates three jobs. Audit where your product actually surfaces (Marketplace, Seller, Google and AI answer engines), treat Seller's performance insights as a discovery signal not just a sales dashboard, and fix your omnichannel measurement so a sale that starts on one surface is not lost to another.
- Questions to ask
- Do we know every surface where our product currently gets discovered?
- Are we treating Seller insights as a discovery signal or just a sales report?
- Can our measurement follow a sale that starts on one Meta surface and finishes elsewhere?
- Sources
Meta makes 'real human' a free verification badge
- What
- Meta launched Facebook Verified, a free selfie-based badge that confirms a real person sits behind a profile, decoupled from the paid Meta Verified subscription. It verifies a human, not a business: Pages and Pro Mode accounts are explicitly excluded, and the badge appears first on Marketplace, Dating, Groups and Profile, with feed post badges planned later. You verify once and the badge travels with you across Facebook.
- When
- Announced 24 July 2026, rolling out in phases starting in select markets.
- How it shifts discovery
- In an AI-flooded feed, provable humanity is becoming a distribution and trust signal, and Meta rarely ships a free trust layer unless it plans to weight it. Unverified brand-adjacent accounts risk quiet deprioritisation, especially in Marketplace and Groups where transactions and community trust live. Audit which team and creator profiles genuinely represent people, flag them for verification the moment your market goes live, prioritise anyone active in Marketplace, Groups or direct messaging, and keep brand handles on the paid track where relevant.
- Questions to ask
- Which of our team and creator profiles represent real people and should be verified?
- Where does verification reduce friction that converts, such as Marketplace and DMs?
- How do we handle brand Pages that the free badge will not cover?
- Sources
Anthropic ships Opus 5 at Opus 4 prices, making agents cheap
- What
- Anthropic released Claude Opus 5 at the same price as Opus 4.8 while more than doubling coding performance and topping knowledge-work benchmarks. It sets the state of the art on GDPval-AA and OSWorld 2.0, and on Zapier AutomationBench its pass rate is around 1.5 times the next-best model for the same cost. It also introduces an 'effort' setting that lets you dial intelligence up or down per task, turning model selection into a two-axis decision.
- When
- 24 July 2026 (available now, default on Claude Max).
- How it shifts discovery
- The headline is not the intelligence, it is the maths: frontier reasoning is now cheap enough to run long, multi-step agents at scale. Three workflows move from pilot to production overnight: automated content audits that crawl and score pages for decay, GEO monitoring that queries AI answer engines daily to track brand citations, and multi-step research workflows across a whole content calendar. Route bulk sweeps to low effort, reserve max effort for ambiguous high-stakes tasks, and deploy with oversight because longer-running agents drift.
- Questions to ask
- Which always-on agents were blocked purely by token cost and now clear?
- Where do we set low effort for bulk work versus max effort for high-stakes tasks?
- What monitoring stops a long-running agent from drifting into quality erosion?
- Sources
OpenAI details the GPT-5.6 price-performance mechanics
- What
- OpenAI published the mechanics behind the Luna and Terra cuts, framing them as advancing the price-performance frontier by making every layer more efficient. Luna delivers performance comparable to models that were frontier-class a year ago at roughly six cents on the dollar per task, at nearly nine times the speed. OpenAI also introduced Fast mode in the API, which replaces Priority Processing and delivers up to 2.5 times faster GPT-5.6 Sol at twice the price with no change in intelligence.
- When
- 30 July 2026.
- How it shifts discovery
- This matters for the cost of running content generation, GEO monitoring and agentic search workflows at scale, not just customer-facing agents. The practical mechanism is matching intelligence to the outcome: define your quality standard, then use evaluations to find where extra intelligence materially improves the result and where cheaper processing delivers the same quality. Build a routing map now, sending well-specified implementation steps to Luna and reserving Sol for resolving uncertainty and defining the plan.
- Questions to ask
- Have we run evaluations to find where cheaper models match our quality bar?
- Which steps in our AI workflows genuinely need frontier intelligence?
- Does Fast mode change our cost case for latency-sensitive tasks?
- Sources
Two API settings tripled ARC-AGI-3 scores with 6x fewer tokens
- What
- OpenAI showed that turning on two API settings it uses in ChatGPT and Codex, retained reasoning and compaction, tripled GPT-5.6 Sol's scores on the ARC-AGI-3 benchmark and cut output tokens by 6x. With the official harness Sol scored 13.3% on the public set; with retained reasoning and compaction it scored 38.3%. The failure was not the model but the harness: it discarded private reasoning after each action and used a rolling truncation window, so the agent kept relearning the task from scratch.
- When
- 29 July 2026.
- How it shifts discovery
- For teams building agentic discovery and AI-search workflows, memory settings are a practical efficiency lever with a direct effect on both quality and cost. Agents do best when they remember what they have done, so if you run long tasks, retain reasoning across steps and compact rather than truncate history. Audit your own harness or tooling for silent truncation before you conclude a model is underperforming, because you may be measuring your configuration, not the model.
- Questions to ask
- Does our agent tooling discard reasoning or truncate history between steps?
- Are we blaming the model when the harness is the real bottleneck?
- Where would retained reasoning and compaction cut our token bill?
- Sources
GA4 custom events turn default metrics into a real journey
- What
- Google Analytics reminded teams that default data only tells half the story, pointing to custom events like video plays and watch time. Setting up custom events lets you measure specific engagement moments that default GA4 tracking does not capture out of the box. That turns raw pageviews into a usable view of how people actually behave on your content.
- When
- 31 July 2026.
- How it shifts discovery
- If you are trying to prove content value or map a customer journey, default metrics leave you guessing. Track the interactions that signal real intent (video plays, watch time, scroll depth, key clicks) so you can connect on-site behaviour to conversion rather than assuming it. Pick the two or three engagement signals that matter most for your content, set them up as custom events, and feed them into your reporting before your next content review.
- Questions to ask
- Which on-site interactions actually signal intent for our content?
- Are we measuring engagement or just counting pageviews?
- How do these custom events connect to our conversion reporting?
- Sources
OpenAI opens free frontier access to academic researchers
- What
- OpenAI launched ChatGPT for Academic Researchers, giving scientists, mathematicians and engineers free access to its frontier GPT-5.6 models, starting with 10,000 researchers and expanding to 100,000 through 2027. Each workspace includes business-grade privacy, researcher data is not used to train models by default, and participants can invite up to four collaborators. It is built to accelerate discovery across disciplines, from grant applications to testing hypotheses.
- When
- 29 July 2026.
- How it shifts discovery
- This expands a high-authority query and citation surface that GEO teams should watch, especially publishers and brands targeting research-led queries. As frontier tools reach 100,000 researchers, the volume and depth of research-oriented AI queries grows, and with it the value of being a citable, authoritative source in those answers. If you serve research-led or technical audiences, audit whether your content is structured, sourced and specific enough to be cited when these users query AI on their subject.
- Questions to ask
- Do we publish research-led content that AI would cite for technical queries?
- Is our content structured and sourced well enough to be a reliable citation?
- Which research-adjacent audiences are moving their discovery into AI tools?
- Sources
Key takeaways
What to walk away with this week
AI agent economics flipped this week: Luna is 80% cheaper and Opus 5 ships at last year's prices, so always-on conversational agents and long-running discovery workflows finally clear the cost bar.
GA4's new Source Group dimension makes AI-referral traffic countable and applies retroactively, so pull the last 6 to 12 months and rebase attribution now.
Discovery is fragmenting across surfaces (Meta's Seller app, verified human profiles, AI answer engines), so audit every place your product actually surfaces.
Configuration beats model choice more often than teams think: retained reasoning and compaction tripled benchmark scores with 6x fewer tokens, so audit your harness before blaming the model.
Standardise your AI work: record repeatable skills, assign owners, and pair the library with monitoring because agents drift the longer they run.