All issues
Discovery Digest · September 7, 2026

Issue 15. The week GPT-6 arrived and data sovereignty became infrastructure

TL;DR

In three days starting 1 September, the AI frontier delivered more substantive releases than most months have managed. OpenAI launched GPT-6 Astra on 3 September, the most capable model the company has shipped, rated Critical for cybersecurity and gated behind the Daybreak programme before broader access followed. On 1 September, Anthropic released Claude Fable 5.1 with a 75 per cent reduction in cache-read costs and announced Enterprise Frontier Safeguards, a framework that stores activity logs in the customer's own cloud while Anthropic's detection layer still runs across them. Google published Gemini 3.8 Flash on 2 September, its third Flash model in six weeks, alongside a restricted cybersecurity variant for government and infrastructure operators. And Perplexity launched Hybrid Compute on Mac the same day, routing sensitive data to an on-device model while the cloud handles reasoning and planning. The week that opened September set the terms for the rest of the quarter.

Issue 15. The week GPT-6 arrived and data sovereignty became infrastructure
01 · OpenAI

OpenAI launches GPT-6 Astra, the first frontier model rated Critical for cybersecurity

What
OpenAI launched GPT-6 Astra on 3 September 2026, its first GPT-6 model and the most capable model the company has shipped publicly. The model carries a 1.05 million token context window, 128,000 token maximum output, text and image input, and a knowledge cutoff of 30 April 2026. On FrontierMath Tier 4 it scores 97.6 per cent; on ARC-AGI-3, 99.9 per cent; on ExploitBench, 100 per cent. It is the first OpenAI model rated Critical for cybersecurity capability. Pricing is $10 per million input tokens and $50 per million output tokens, 2.5 times the GPT-5.6 Sol rate. A Fast mode runs at double the standard per-token rate. Access began with Daybreak programme enterprises, followed by ChatGPT Plus, Pro, Business, Enterprise, and the API over the days following launch.
When
Launched Wed 3 September 2026. Announced on openai.com. Reported by Al Jazeera, Yotta Labs, and DataCamp on 3 to 4 September 2026.
How it shifts discovery
GPT-6 Astra's 100 per cent ExploitBench score and Critical cybersecurity rating mean OpenAI has shipped a model that saturates the published exploit benchmark, and placed access controls on it that match the capability level. The Daybreak programme vetting gate, deployed again here after GPT-5.6 Sol in July and GPT-5.6-Cyber in August, is now a repeating access pattern rather than a one-off precaution. At $50 per million output tokens, Astra is priced for workloads where the performance differential over Sol justifies the cost premium: complex agentic tasks, long-horizon research, and multi-step computer-use workflows where Sol's error rate produces downstream cost that exceeds the model premium. The 99.9 per cent ARC-AGI-3 score is the more strategically significant figure. ARC-AGI-3 is designed to test genuine generalisation rather than benchmark pattern-matching. A score at that level signals a qualitative step rather than incremental improvement. Every model pricing and capability decision made in August against GPT-5.6 Sol as the frontier ceiling is now made against a model one tier above it.
Questions to ask
  • GPT-6 Astra carries a Critical cybersecurity rating and a 100 per cent ExploitBench score. For teams with Daybreak programme membership or plans to apply, what is the named owner and evaluation timeline for assessing whether Astra's cybersecurity capability is relevant to our red-team or vulnerability research function?
  • Astra is priced at $50 per million output tokens, 2.5 times GPT-5.6 Sol. For the agentic and long-horizon tasks we currently run on Sol, have we identified which produce downstream errors costly enough to justify the step up to Astra pricing, and is there a benchmarking process ready to run this week?
  • GPT-6 Astra's 99.9 per cent ARC-AGI-3 score signals a generalisation capability qualitatively above the GPT-5.6 family. What is the trigger in our model evaluation framework for reassessing production workload allocations when a new frontier tier arrives, and was that trigger set before this week?
Sources
02 · Anthropic

Anthropic releases Fable 5.1 with a 75% cache-read cost cut and a restricted Mythos 5.1 variant

What
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on 1 September 2026. Fable 5.1 is generally available across the Claude API, Amazon Web Services, Google Cloud, and Microsoft Azure as claude-fable-5-1. It keeps Fable 5's headline pricing of $10 per million input tokens and $50 per million output tokens but reduces the cache-read rate by 75 per cent to $0.25 per million tokens. Anthropic describes the model as delivering improved performance over Fable 5 in coding and knowledge-work tasks. Mythos 5.1 is a restricted-access variant with safeguards selectively reduced for vetted cybersecurity and life-sciences organisations requiring capabilities normally constrained by the standard safety layer.
When
Released Mon 1 September 2026. Announced on anthropic.com. Reported by VentureBeat and MacRumors on 1 September 2026.
How it shifts discovery
The headline pricing staying flat while the cache-read rate drops 75 per cent is a specific signal about where Anthropic sees growth in Fable-tier workloads: not in new query volume but in sustained, high-context, multi-session deployments where a prompt cache is doing significant work. For teams running Fable 5 in production agentic pipelines with large system prompts or document context reused across requests, the cache-read reduction changes the economics of those workflows immediately, without any re-engineering required. The Mythos 5.1 variant follows the same pattern OpenAI established with GPT-5.6-Cyber and now GPT-6 Astra: a restricted tier with selectively reduced safeguards for vetted specialised use cases. Three of the four leading frontier labs, OpenAI, Anthropic, and Google with its Cyber series, now ship restricted-access models alongside their standard tiers in the same week. The pattern is structural: access controls are the industry's chosen mechanism for managing the most sensitive AI capability.
Questions to ask
  • Fable 5.1's cache-read rate drops 75 per cent to $0.25 per million tokens. For workloads running on Fable 5 with large reused system prompts or document context, have we calculated the revised cost and confirmed whether the economics of staying on Fable tier change materially compared with Opus 5 for any affected pipelines?
  • Mythos 5.1 is available to vetted cybersecurity and life-sciences organisations with selectively reduced safeguards. Does our organisation have use cases blocked by Fable 5's standard safety layer, and is there a named owner evaluating whether a Mythos programme application is relevant this quarter?
  • Fable 5.1 is now callable as claude-fable-5-1 across AWS, GCP, and Azure. For teams accessing Fable 5 through cloud provider APIs, what is the process for updating model identifiers in production, and have we confirmed the rollout timeline with our specific cloud provider?
Sources
03 · Anthropic

Anthropic launches Enterprise Frontier Safeguards, keeping customer logs in the customer's cloud

What
Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, a framework that resolves the conflict between zero data retention and misuse detection in regulated AI deployments. Under EFS, activity logs, prompt transcripts, and agent session data are stored in cloud infrastructure the customer controls, with the customer holding the encryption keys and managing human review. Anthropic's detection algorithms run against that data within the customer's environment rather than requiring the data to transfer to Anthropic's servers. The system was developed with more than 100 customers across financial services, healthcare, manufacturing, telecom, law, retail, and the public sector, and with AWS, Google Cloud, and Microsoft Azure. Eligible customers receive zero data retention on Fable 5 and Fable 5.1 until the phased rollout begins in autumn 2026.
When
Announced Mon 1 September 2026. Published on anthropic.com/news. Reported by Unite.AI and MarkTechPost on 1 to 2 September 2026.
How it shifts discovery
The conflict EFS addresses has blocked regulated enterprise AI adoption more consistently than any capability gap. Data residency requirements, legal privilege, and regulated data handling rules have forced organisations in financial services, healthcare, and law to choose between zero data retention (which eliminates misuse detection) and misuse detection (which requires the vendor to hold the data). EFS resolves this at the infrastructure layer: the customer controls the vault, Anthropic runs the analysis within it. The 100-customer development collaboration and three-cloud-provider integration at launch signal production-ready infrastructure rather than a pilot programme. For enterprise teams managing AI governance frameworks, EFS makes a specific deferred decision available this quarter: a Claude deployment that satisfies data residency requirements and passes a compliance audit is no longer a theoretical future state. The phased rollout starting autumn 2026 means teams that need to factor EFS into a procurement or governance review should be contacting Anthropic now rather than waiting for general availability.
Questions to ask
  • EFS resolves the zero-data-retention versus misuse-detection conflict that has blocked regulated Claude deployments. Which deferred Claude use cases in our organisation were blocked by the inability to satisfy both requirements simultaneously, and is there a named owner contacting Anthropic about EFS eligibility this week?
  • EFS stores activity logs in the customer's own cloud infrastructure with customer-held encryption keys. Does our cloud provider require specific configuration to participate in EFS, and has our cloud security team been briefed on the architecture before any deployment decision is made?
  • One hundred customers across financial services, healthcare, law, and the public sector contributed to EFS design. Are any peer organisations or regulators in our sector among them, and does that shift our expectation of EFS becoming a compliance standard rather than a differentiator in our industry?
Sources
04 · Google AI

Google ships Gemini 3.8 Flash, its third Flash model in six weeks, with a cybersecurity variant for government use

What
Google released Gemini 3.8 Flash on 2 September 2026, available via the Gemini API, Google AI Studio, and Android Studio under the model identifier gemini-3.8-flash, and rolling out to select consumer products. The model is Google's third Flash release in six weeks, following Gemini 3.6 in late July and Gemini 3.7 Flash on 13 August. It improves on 3.7 Flash across all benchmarks Google published, including software engineering, long-horizon agentic tasks, and multi-step reasoning in specialised domains. It supports text, image, audio, video, and PDF input with a 1 million token context window and 64,000 token maximum output. Introductory pricing matches 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026. Alongside the standard model, Google released Gemini 3.8 Flash Cyber, a restricted variant with enhanced cybersecurity capabilities available only to government agencies and critical infrastructure operators.
When
Released Tue 2 September 2026. Announced on the Google AI Blog. Reported by 9to5Google and Unite.AI on 2 September 2026.
How it shifts discovery
Google has shipped three Flash models in six weeks. At that iteration pace, the evaluation question shifts: it is less about which Flash model to use than how to build an evaluation process that keeps pace with a three-week release cycle. Each iteration directly affects the quality of AI answers across Google's consumer surfaces, and the improvement from 3.7 to 3.8 is measurable across published benchmarks. For practitioners, two decisions are time-bounded. The $0.75 introductory pricing expires 31 December 2026 and doubles from 1 January 2027: any workload migration to 3.8 Flash needs to be benchmarked and completed before then. The Gemini 3.8 Flash Cyber restricted variant extends the week's clearest pattern: Google joins OpenAI and Anthropic in shipping a cybersecurity-enhanced restricted tier alongside the standard model in the same week. All three labs have drawn the boundary between standard and restricted capability at access controls rather than at capability limits. The AI infrastructure layer is bifurcating, and the tier separation is structural, not temporary.
Questions to ask
  • Gemini 3.8 Flash is now available in the API and rolling out across consumer products. Have we re-tested how our brand and content appear in Google AI answers since the update, and have those outputs changed compared with the Gemini 3.7 Flash baseline we established after the 13 August release?
  • The $0.75 introductory price for Gemini 3.8 Flash expires 31 December 2026 and doubles from 1 January 2027. For workloads running on other providers that 3.8 Flash could handle, what is our deadline for completing a benchmark comparison and completing any migration before the introductory window closes?
  • Google, OpenAI, and Anthropic have each shipped restricted cybersecurity-enhanced model variants in the same week. Does our AI vendor risk framework track restricted capability tiers within each provider separately from the standard offering, and does it include a process for determining whether those restricted tiers are relevant to our organisation?
Sources
05 · AI Search

Perplexity launches Hybrid Compute on Mac, routing sensitive data to an on-device model while the cloud handles reasoning

What
Perplexity launched Hybrid Compute for the Perplexity Mac app on 1 September 2026, available to Pro, Max, and Enterprise subscribers on Apple silicon Macs running macOS 15 or later with at least 24GB of unified memory. The system divides each Perplexity Computer task between cloud-based frontier models and a local model running on the Mac. The cloud handles frontier reasoning, web search, and planning; the local model processes private files, sensitive information, and on-device actions. Three local models are available at launch: Gemma 4 E4B, Qwen3.6 35B-A3B, and a Perplexity model post-trained for Perplexity Computer. Perplexity open-sourced the PII classifier used to route workloads between local and cloud. Enterprise admins can set organisation-wide rules for what must remain on-device, what may be masked before cloud use, and what requires user approval per session.
When
Launched Mon 1 September 2026. Announced by Perplexity on perplexity.ai. Reported by 9to5Mac and MarkTechPost on 1 September 2026.
How it shifts discovery
Perplexity's Portable Computer, covered last week, ran every task locally on Nvidia DGX Spark or RTX GPU hardware. Hybrid Compute is architecturally distinct: the cloud orchestrates the full task while the local model handles specific substeps containing sensitive or identifiable data. That distinction matters for enterprise procurement. Where Portable Computer required specialist hardware and full local execution, Hybrid Compute integrates into a cloud-first workflow for the majority of a task while keeping sensitive steps on-device, on hardware many professionals already own. The open-sourced PII classifier gives enterprise procurement and legal teams something auditable: the decision logic for what leaves the device is a published, reviewable component rather than a proprietary black box. For organisations evaluating whether a local AI deployment can satisfy data governance requirements, Hybrid Compute moves the question from architecture to configuration. The routing logic is open, and the enterprise admin layer allows named individuals to enforce organisation-wide rules rather than relying on per-session employee judgement.
Questions to ask
  • Hybrid Compute routes sensitive data to an on-device model based on an open-sourced PII classifier. For teams handling client or regulated data in Perplexity Computer sessions, does the published routing logic satisfy our data governance review, and is there a named owner evaluating the PII classifier against our data classification policy?
  • Enterprise admins can set organisation-wide rules in Hybrid Compute for what must stay on-device. For organisations that have deployed Perplexity Computer without admin-level data handling policies in place, is there an owner accountable for setting those rules before Hybrid Compute adoption spreads informally within the team?
  • Hybrid Compute requires Apple silicon Mac with 24GB unified memory; Portable Computer requires Nvidia DGX Spark or RTX GPU hardware. For organisations evaluating on-device AI agent options, does our current device estate support either path, and does our IT infrastructure roadmap include a view on which hardware tier is the target for local AI workloads?
Sources

Key takeaways

What to walk away with this week

  1. GPT-6 Astra's 99.9 per cent ARC-AGI-3 and 100 per cent ExploitBench scores represent a qualitative step above GPT-5.6 Sol: rebuild your frontier model tier assumptions this week and identify which production workloads justify the $50 output rate before defaulting to Sol for another quarter.

  2. Three of the four leading frontier labs shipped restricted cybersecurity-enhanced model variants in the same week. Access controls, not capability limits, are the industry's settled mechanism for managing sensitive AI applications: update your vendor risk framework to track restricted tiers separately from standard offerings.

  3. Anthropic's Enterprise Frontier Safeguards resolve the zero-data-retention versus misuse-detection conflict that has blocked regulated Claude deployments in financial services, healthcare, and law. If data residency was your blocker, contact Anthropic about EFS eligibility this quarter before the phased rollout determines availability.

  4. Fable 5.1's 75 per cent cache-read cost reduction applies immediately to production agentic pipelines with large reused system prompts or document context. Audit your Fable 5 workloads this week: the saving requires no re-engineering and has been live from 1 September.

  5. Perplexity Hybrid Compute and Anthropic EFS both launched on 1 September with the same underlying premise: sensitive data stays in the customer's environment while the vendor's intelligence still runs across it. Track which AI vendors are resolving data sovereignty at the infrastructure layer and which are leaving it to per-session user judgement.

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.