Issue 08. The week AI search became the default
TL;DR
This week, the separation between AI search and search ended. Google completed the global rollout making Gemini 3.5 Flash the default output for every query, turning AI Mode from an opt-in feature into the baseline for every user worldwide. OpenAI repositioned ChatGPT as a work-delivery platform with ChatGPT Work, an agent that ships finished spreadsheets and documents rather than text replies, while Meta released Muse Image, its first agentic image generation model, across Instagram and WhatsApp. Behind those transitions, Bloomberg confirmed Google has scrapped and rebuilt Gemini 3.5 Pro after coding failures, and Anthropic opened chip talks with Samsung: signals that the infrastructure race below the product layer is as contested as the product race above it.
Google makes AI-generated answers the default output for every search query worldwide
- What
- Google completed the global rollout on 10 July 2026 that makes Gemini 3.5 Flash the default primary output for every Google search query. A prose answer, synthesised in real time from multiple web sources with inline citations, now appears before the traditional ranked list of links for all users, across all markets and languages. The ranked link results remain on the page but are secondary. There is no permanent user-facing toggle to restore the pre-July experience. The transition had been announced at Google I/O on 19 May 2026, with Liz Reid, Google's Vice President of Search, describing it as the biggest upgrade to the search box in over 25 years.
- When
- Rollout completed globally on Fri 10 July 2026. Reported by Search Engine Land and TechTimes between 10 and 13 July 2026.
- How it shifts discovery
- The change that search and marketing teams have been planning for over two years was completed last week, and it arrived without a pause. The ranked list of blue links is no longer the primary output of Google Search for any query, for any user. Every reporting framework built around organic impressions, clicks, and CTR is now measuring the secondary layer of the search page. The questions that practitioners have been treating as future-planning items, how to appear in AI Mode, how to interpret AI citation signals, how to attribute traffic from AI surfaces, are operational requirements from this week forward. Teams that activated Google's Search Console AI performance reports in June have a baseline. Teams that did not are building one retrospectively against a rollout that is already complete.
- Questions to ask
- Our organic click data in Search Console reflects the secondary link layer, not the primary AI answer. Do we know the share of queries where our brand or content appears in the Gemini prose output versus the ranked list below it, and have we updated our reporting frameworks to reflect which layer is now primary?
- Google's Search Console AI performance reports have been rolling out since June. Have we activated them, established a pre-July baseline, and confirmed who in the organisation owns that data as a standing measurement responsibility?
- There is no permanent opt-out for search users. Do we know how Gemini 3.5 Flash represents our brand across our five highest-volume query categories, and are those representations accurate, current, and consistent with our brand positioning?
- Sources
OpenAI launches ChatGPT Work, an agent that ships finished spreadsheets, documents, and web apps
- What
- OpenAI launched ChatGPT Work on 9 July 2026 for Pro, Enterprise, and Education users, with a rollout to Plus and Business planned over the following days. ChatGPT Work is a GPT-5.6 agent that connects to a user's existing tools via plugins, including Slack, Microsoft Teams, Google Drive, SharePoint, email, calendar, CRM, and project trackers, breaks a goal into steps, executes those steps across connected apps, and returns completed work: spreadsheets, slides, documents, and interactive web applications. Scheduled Tasks extend the system to automation, allowing the agent to perform an action once, repeat it on a schedule, or monitor for changes in connected apps over time. The launch merges Codex into the same desktop application.
- When
- Launched on Thu 9 July 2026. Announced on openai.com and reported by Bloomberg and BNN Bloomberg on 9 July 2026.
- How it shifts discovery
- ChatGPT's competitive position has been built around being the best answer engine. ChatGPT Work changes the competitive category: the product now competes with Microsoft Copilot Cowork, Google Workspace AI, and Notion AI as a work-delivery layer, not a conversation layer. The strategic implication for brand and marketing teams is direct. The outputs a user receives from ChatGPT Work, a research brief, a competitor analysis, a campaign deck, are synthesised from whatever sources the agent retrieves, and that retrieval is governed by the same citation and ranking logic as a standard ChatGPT query. Brands absent from ChatGPT's retrieval layer are absent from the finished work product, not merely from the conversational reply. Scheduled Tasks compound this: a user who sets an agent to monitor a competitor or category does not need to re-query. The agent does it continuously on a schedule. At scale, that is a persistent monitoring layer over the web, operating on behalf of millions of users at once.
- Questions to ask
- ChatGPT Work retrieves content to complete tasks such as competitor analysis, market research, and campaign planning. Do we know how our brand, products, and category appear in the material ChatGPT Work would retrieve for those tasks, and is that material accurate and current?
- Scheduled Tasks let users automate ongoing research and monitoring across connected apps. Have we considered the scenarios where a competitor, analyst, or journalist could use ChatGPT Work to monitor our brand, product launches, or pricing changes, and are we confident our public-facing information would be represented accurately in those outputs?
- ChatGPT Work merges Codex into the same agent as GPT-5.6 and handles task execution across connected enterprise apps. For teams with existing Codex or GPT-5.6 deployments, does the merged architecture change our data-handling configuration, permission model, or cost structure?
- Sources
Meta releases Muse Image, its first agentic image generation model, across Instagram and WhatsApp
- What
- Meta released Muse Image on 7 July 2026, the first image generation model from Meta Superintelligence Labs, available immediately in Meta AI on Instagram and WhatsApp, with Facebook, Messenger, and Advantage+ creative integration to follow. Unlike standard diffusion models, Muse Image operates as an agent: it invokes search and code-execution tools to improve factual accuracy, generates accurate QR codes and text-embedded infographics, and combines elements from multiple reference images into a single output. At launch, the model holds second place in the Artificial Analysis Arena for text-to-image, single-image editing, and multi-image editing on human-preference Elo rankings. Muse Video is in development. Everyday creation is free; a subscription applies once users exceed a usage threshold.
- When
- Released on Mon 7 July 2026. Announced on the Meta Newsroom and ai.meta.com, reported by TechCrunch and CNBC on 7 July 2026.
- How it shifts discovery
- Muse Image's reach is what makes it strategically significant. Instagram, WhatsApp, and Facebook together cover five to six billion monthly users. As the model extends across Meta's full surface, it becomes the largest distribution network for AI image generation in the world, accessible through a messaging interface rather than a dedicated creative tool. For brands operating on Meta's platforms, the Advantage+ integration is the operational question to resolve first. When Muse Image generates and iterates advertising creative at scale, a brand's existing visual assets are the source material it works from. Brand image libraries that are accurate, licensed, and representative of current visual identity are the input the model will use. Muse Image's agentic accuracy capabilities close a long-standing limitation of image generation in advertising contexts: readable text in images, correct QR codes, and factual accuracy are functional requirements for most performance creative, and they are now built into the generation step rather than corrected in post-production.
- Questions to ask
- When Muse Image reaches Advantage+ creative, it will generate and iterate ad creative from our existing brand assets. Have we audited our Meta creative library for accuracy, licensing, and alignment with our current visual identity, so the AI-generated output reflects the brand we intend to represent?
- Muse Image is live in Instagram and WhatsApp today. Do we know how our brand is represented when users describe or request imagery related to our product category inside Meta AI, and is there anything in that output we need to address before the Facebook rollout expands its reach?
- Muse Image uses agentic tool-calling to improve factual accuracy and produce text-embedded imagery correctly. Does that capability change the economics of producing performance creative at scale for our Meta campaigns, and have we mapped how it compares to our current production workflow?
- Sources
Google scraps and rebuilds Gemini 3.5 Pro after coding failures, no release date named
- What
- Bloomberg reported on 16 July 2026 that Google has delayed Gemini 3.5 Pro, its planned flagship AI model, after internal testing revealed the system fell short of expectations in coding performance and complex reasoning tasks. Google announced at I/O in May that Gemini 3.5 Pro would ship in June 2026. When June passed without a release, Google updated the model's training data to improve coding skills, but the results were disappointing. According to ten current and former Google employees cited by Bloomberg, the model was subsequently scrapped and rebuilt from scratch. A Google spokesperson confirmed to Bloomberg that the company is currently testing Gemini 3.5 Pro with partners, alongside an upgraded Flash model, but gave no new release date. Alphabet's share price fell on the day of the report.
- When
- Reported by Bloomberg and CNBC on Thu 16 July 2026. Confirmed by Search Engine Journal on the same date.
- How it shifts discovery
- Gemini 3.5 Flash is Google's most capable publicly available model, and as of 10 July it is the model generating the default answer for every Google Search query. Gemini 3.5 Pro was the frontier tier intended to sit above it. Without a Pro release date, Google is running the world's most used AI search surface on a mid-tier model while its flagship is rebuilt. That gap is least visible in conversational search, where Flash performs well. It is most visible in the agentic, multi-step, and complex-reasoning workloads where Anthropic's Claude Opus 4.8, OpenAI's GPT-5.6 Sol, and xAI's Grok 4.5 each have commercially available frontier tiers. The sourcing Bloomberg used, ten current and former employees describing frustration and concern about competitive standing, is a signal worth weighing independently of the release timeline. Google has structural advantages in search distribution, data, and inference scale that no competing lab can replicate. If those advantages are not producing a competitive frontier model on schedule, the reasons are worth examining before treating the delay as temporary.
- Questions to ask
- Gemini 3.5 Pro is the model Google needs to match GPT-5.6 Sol and Claude Opus 4.8 at the frontier capability tier. If our AI strategy assumes Google will close that gap within a specific planning horizon, how does the absence of a release date affect our near-term model selection and deployment decisions?
- Google's AI search default runs on Gemini 3.5 Flash for all queries. For searches in our product category that involve complex reasoning, multi-step research, or agentic task completion, do we know how Flash-level capability affects the quality and completeness of the answers our brand appears in?
- Bloomberg sourced ten current and former Google employees expressing concern about competitive standing. Does our vendor risk assessment for Google's AI stack reflect the possibility of a longer-than-expected gap between Flash and Pro capability levels, and how does that affect any Google-dependent AI workloads we have in production or in planning?
- Sources
Anthropic opens talks with Samsung to develop a custom 2nm AI inference chip
- What
- Anthropic is in early-stage talks with Samsung Electronics to design and manufacture a custom AI chip using Samsung's 2-nanometer process node and advanced packaging facilities, according to reporting by The Information and Bloomberg published on 2 July 2026. The project is at a preliminary stage: Anthropic has not begun detailed chip design, testing, or manufacturing, and is still determining the processor's specifications and how it would fit into a server configuration. Anthropic recently hired Clive Chan, a founding member of OpenAI's custom chip team, as part of a broader hardware buildout. The talks follow OpenAI's earlier announcement of its own custom inference processor, Jalapeño, developed in partnership with Broadcom.
- When
- Reported by The Information and Bloomberg on Thu 2 July 2026. Further reported by TechCrunch and UPI on 2 to 3 July 2026.
- How it shifts discovery
- Anthropic's compute expenditure is reported at approximately $1.25 billion per month. At that scale, a custom chip programme is not a research initiative: it is a commercial response to a cost structure that external silicon cannot resolve. Both Anthropic and OpenAI are responding to the same fundamental pressure. Inference costs at frontier model scale are high enough that owning the silicon layer changes the economics of the business in a way that no combination of GPU procurement deals can fully replicate. For organisations building AI strategies on Claude, the medium-term implication is that Anthropic's capacity to run and price frontier models is constrained by external chip supply today but will be partially self-determined within two to three years if the Samsung partnership advances. The confidence signal here is the more immediate read: Anthropic is making multi-year infrastructure commitments at a moment when most of its revenue is less than two years old. That is a planning posture, not a survival posture.
- Questions to ask
- Anthropic's compute costs are reported at roughly $1.25 billion per month. If custom silicon materially reduces those costs over the next two to three years, how does that change our expectation for Claude API pricing, and have we built a scenario for lower inference costs into our AI cost modelling for 2028 onwards?
- Both OpenAI and Anthropic now have active custom chip programmes. As frontier labs move toward silicon ownership, how does that change the vendor dependency and concentration risk in our AI infrastructure strategy, and do we have a position on building critical workflows around labs that control their own compute versus those that remain externally dependent?
- Anthropic's chip project is at an early stage with no confirmed design, testing, or manufacturing underway. What is the typical lag between a 2nm chip design commitment and volume production capacity, and does that two-to-three-year horizon matter for the planning cycle we are currently working in?
- Sources
- The Information, Anthropic in Talks With Samsung to Manufacture Custom AI Chip, 2 July 2026
- Bloomberg, Anthropic in Talks With Samsung for Custom AI Chip, 2 July 2026
- TechCrunch, Anthropic is discussing a new custom chip with Samsung, 2 July 2026
Key takeaways
What to walk away with this week
The separation between AI search and search ended on 10 July. Every Google query now returns an AI-generated answer by default. Establish your AI impressions baseline in Search Console this week and update your reporting framework to treat the Gemini prose layer as the primary search surface, not a secondary feature.
ChatGPT Work's Scheduled Tasks mean AI agents are now monitoring competitors, categories, and brand mentions on behalf of users around the clock. Audit the accuracy and completeness of your public-facing information, because a persistent agent audience is operating on that material continuously, not occasionally.
Gemini 3.5 Pro has no release date. The model powering Google Search is a mid-tier Flash model, and Google's planned frontier tier is still being rebuilt. Reassess any strategy or AI procurement decision that assumes Google's frontier capability is comparable to GPT-5.6 Sol or Claude Opus 4.8 in the near term.
Meta Muse Image will generate advertising creative from your existing brand assets at scale through Advantage+. Audit your Meta creative library for accuracy, licensing, and visual identity alignment before that integration arrives, because the model's outputs will be shaped by whatever you have already given it to work from.
Both Anthropic and OpenAI now have active custom chip programmes. Silicon ownership is becoming a prerequisite for frontier AI at scale. Factor infrastructure independence into your AI vendor assessment alongside model performance and pricing.