All issues
Discovery Digest · August 31, 2026

Issue 14. The week enterprise defaults were set and compute moved to the edge

TL;DR

This week, enterprise AI stopped being optional infrastructure and started being built into the default systems organisations already run. Salesforce and Anthropic announced Claudeforce on 26 August, establishing Claude as the reasoning engine across the world's most-used enterprise CRM, while Google launched Ask Gemini inside Google Chat the same day, folding Gemini into the collaboration layer that hundreds of millions of Workspace users work inside daily. OpenAI published the first independently validated results from its Jalapeño chip at Hot Chips 2026, posting a 1.9x efficiency lead over Nvidia Blackwell and accelerating the timeline for frontier API pricing compression. Perplexity launched Portable Computer, the most aggressive attempt yet to move serious AI agent workloads off the cloud and onto local hardware, with a pricing model that charges nothing for work completed on-device. And Anthropic confirmed that Claude Sonnet 5's launch pricing is permanent, cancelling the 50 per cent increase that had been scheduled for 1 September.

Issue 14. The week enterprise defaults were set and compute moved to the edge
01 · Anthropic

Salesforce and Anthropic embed Claude as the default reasoning engine across the CRM stack

What
On 26 August 2026, Salesforce and Anthropic announced Claudeforce, an expanded strategic partnership establishing Claude as the default reasoning model across the Salesforce ecosystem. The partnership launches in two directions simultaneously. Salesforce in Claude is a plugin with 37 prebuilt sales skills that give Claude users direct access to live Salesforce pipeline data, automated workflow updates, and governed commercial actions from inside Claude. In the opposite direction, Claude is being integrated as the reasoning layer inside Salesforce's own CRM, applying Claude's intelligence to revenue data and workflows that sales teams already use daily. Select pilot customers are live now; open beta is expected in September 2026, with additional prebuilt skills to follow in late 2026.
When
Announced jointly by Salesforce and Anthropic on Wed 26 August 2026. Reported by CNBC, VentureBeat, and Digital Commerce 360 on 26 to 27 August 2026.
How it shifts discovery
Salesforce processes the commercial relationships of approximately 150,000 businesses. When Claude becomes the default reasoning model inside that CRM, the AI layer touching sales, pipeline management, and customer data for a substantial share of global enterprise is no longer model-agnostic. That is a structural shift. CNBC coverage explicitly frames the partnership as Salesforce's response to what it calls the SaaSpocalypse concern, the risk that AI agents replace SaaS application interfaces entirely. Claudeforce is Salesforce's answer: embed the agent inside the CRM rather than let the agent replace it. The 37 prebuilt sales skills represent the first native AI sales interface where commercial data and reasoning model are from separate vendors but behave as a single system. For brand and marketing teams whose commercial relationships, pricing data, and customer intelligence live inside Salesforce, the discovery implication is immediate: sales conversations, competitive analysis, and customer intelligence generated inside a Salesforce-Claude session will be shaped by Claude's retrieval and reasoning, with all the source-quality dependencies that implies.
Questions to ask
  • Salesforce in Claude gives AI agents live access to a user's Salesforce pipeline data and the ability to take governed commercial actions. Have we reviewed what data our Salesforce instance exposes through that plugin, and does our data governance policy cover AI agents with access to live CRM records?
  • Claudeforce establishes Claude as the default reasoning engine across the Salesforce ecosystem. For teams whose sales intelligence, forecasting, and competitive analysis originate from Salesforce, how does embedding Claude's reasoning into that system of record change who or what is generating commercial decisions?
  • The open beta launches in September 2026. Which of our Salesforce-dependent workflows have the highest potential to be restructured by a Claude reasoning layer embedded in the CRM, and is there a named owner monitoring the beta before the full rollout begins?
Sources
02 · OpenAI

OpenAI's Jalapeño chip posts a 1.9x efficiency lead over Nvidia Blackwell at Hot Chips 2026

What
OpenAI presented the first public performance results for Jalapeño, its custom AI inference chip developed with Broadcom, at the Hot Chips 2026 conference on 25 August 2026. On the InferenceX benchmark platform, Jalapeño delivered 1.9 times more AI work per watt than Nvidia's GB300 system, with end-to-end latency 2.1 to 4.1 times lower on interactive workloads. The chip carries 216 GB of HBM4 memory, a 700W TDP rating, and ran at 550W or below in sustained testing, compared with the GB300's 1,400W. Independent analysis by semiconductor research firm SemiAnalysis confirmed the 1.9x efficiency lead on 26 August. OpenAI confirmed limited deployment in late 2026, with volume production in 2027.
When
Results presented at Hot Chips 2026 on Mon 25 August 2026. Reported by Bloomberg, The Register, and Tom's Hardware on 25 August 2026. Independent confirmation from SemiAnalysis published 26 August 2026.
How it shifts discovery
Custom silicon has been on both OpenAI and Anthropic roadmaps for the past 18 months. The Jalapeño results at Hot Chips are the moment that planning item becomes a dated commercial reality. Independent confirmation by SemiAnalysis changes the classification of the results from self-published benchmarks to externally validated performance data. A chip that delivers 1.9 times more AI work per watt than Nvidia's current flagship server silicon is not a marginal cost improvement at the scale ChatGPT operates: it is a cost-structure shift. The deployment timeline, limited volumes late 2026, volume production 2027, sets the date at which that advantage starts flowing into OpenAI's inference economics. For practitioners, this changes two planning assumptions that have been stable for three years. First, the direction of frontier AI API pricing is downward and accelerating. Custom silicon at OpenAI, combined with Anthropic's reported Samsung talks, means both major frontier labs are removing Nvidia GPU cost as a hard floor from their infrastructure economics. Second, any strategy that treats today's API pricing as a 2027 planning baseline is working from a model that will not hold. The useful question is not whether prices will fall, but by how much and when.
Questions to ask
  • Jalapeño's 1.9x efficiency lead over Nvidia GB300 was independently confirmed by SemiAnalysis. If OpenAI's inference costs fall materially from 2027 as Jalapeño enters volume production, have we built a scenario for lower ChatGPT and API pricing into our AI cost planning for 2027 and beyond?
  • Both OpenAI and Anthropic are building custom silicon to remove Nvidia GPU cost as an infrastructure floor. For AI workloads we currently price against Nvidia-powered API endpoints, does our 12-month cost model account for the downward pressure custom silicon will apply to inference pricing from late 2026 onwards?
  • OpenAI's Jalapeño results were presented at a public semiconductor conference. Is our organisation tracking custom silicon milestones from frontier labs as part of a standing infrastructure intelligence function, or are results like this arriving as news rather than anticipated data points?
Sources
03 · AI Search

Perplexity launches Portable Computer: a local-first AI agent that charges nothing for on-device work

What
Perplexity launched Portable Computer on 25 August 2026, a local-first AI agent developed in partnership with Nvidia that runs entirely on hardware users already own. The launch supports Nvidia's DGX Spark desktop and Linux machines equipped with Nvidia RTX GPUs. Every task starts on the device by default; the system asks permission before sending any individual step to a cloud frontier model. Work completed locally incurs no billing credits. At launch, users can run Qwen 3.8 27B or PPLX 27B, Perplexity's own post-trained model variant, with Nvidia Nemotron 3.5 Lightning coming soon. Portable Computer is an on-device version of Perplexity's cloud-based Computer agent, which launched in February 2026.
When
Launched Mon 25 August 2026. Announced at perplexity.ai/hub/blog and reported by SiliconANGLE and VentureBeat on 25 August 2026.
How it shifts discovery
Perplexity built its platform on cloud inference where every query incurs a token cost. Portable Computer inverts that model: the default is local, the cost for local work is zero, and cloud access is an explicit opt-in on a task-by-task basis. The privacy and compliance implication is the more immediate signal for enterprise practitioners. An AI agent that starts every task on-device by default, with data leaving the hardware only when the user approves a specific step, represents a qualitatively different compliance posture than any cloud-first agent. For organisations with data residency requirements, legal privilege concerns, or sensitivity around what crosses the enterprise boundary, Portable Computer's architecture directly addresses the governance objections that have blocked AI agent deployments in legal, healthcare, and financial services environments. The partnership with Nvidia positions Portable Computer as the reference deployment for agentic workloads on DGX Spark hardware, extending Perplexity's platform into enterprise edge compute rather than remaining cloud-only. The zero-cost local model also changes the economic argument for organisations that evaluated cloud-based AI agents and found per-query economics prohibitive at the volume they need.
Questions to ask
  • Portable Computer starts every task on-device by default and asks permission before any step goes to a cloud model. For AI agent use cases that have been blocked by data residency or confidentiality requirements, does a local-first architecture with explicit cloud opt-in resolve the governance objection that delayed them?
  • Perplexity charges zero billing credits for work completed locally on Portable Computer. For research and analysis workflows where a 27-billion-parameter model is sufficient for the task, have we assessed the cost case for running those workflows on local hardware versus cloud API calls at current pricing?
  • Portable Computer launches on Nvidia DGX Spark and Linux machines with RTX GPUs. Is there a named owner evaluating the hardware requirements against our existing device estate, and does our AI infrastructure roadmap include edge compute as a planned deployment target alongside cloud infrastructure?
Sources
04 · Google AI

Google replaces Chat's side panel with Ask Gemini, turning Google Chat into a unified AI command line

What
Google launched Ask Gemini in Google Chat on 26 August 2026, replacing the existing Workspace side panel with a unified AI interface that operates across the full Google Workspace suite from a single Chat conversation. Through the interface, users can search across Gmail, Google Drive, and Calendar; generate images; draft content; and schedule meetings without leaving Chat. The launch is rolling out to Workspace customers whose accounts are set to English, with other language support to follow. Through 1 October 2026, Workspace customers have access to elevated usage limits at no additional charge as a promotional adoption period.
When
Launched Tue 26 August 2026. Announced on the Google Workspace Updates blog. Reported by Neowin and Phandroid on 26 August 2026.
How it shifts discovery
Google Chat has competed with Slack and Microsoft Teams for years without a clear capability advantage. Ask Gemini changes the competitive axis from messaging features to task resolution. A Chat interface that can search a Gmail archive, pull a Drive document into context, and draft a reply in the same session without any tab switch is not competing on message threading: it is competing on how much work it completes inside a single conversation. Microsoft placed Copilot inside its unified app in the August 18 rollout covered last week; Google is placing Gemini inside Chat on the same day Claudeforce embedded Claude in Salesforce. The collaboration platform with the most capable AI command line is becoming the primary productivity surface for enterprise teams, and these three moves in a single week are the clearest evidence that competition has shifted to this layer. For practitioners, two operational questions follow. First, work that previously generated URL visits, search queries, and tool switches now resolves inside Chat, leaving no clickstream signal for any source used in that session. Second, the promotional period through 1 October is the lowest-friction window to evaluate which workflows migrate into Chat before elevated limits expire.
Questions to ask
  • Ask Gemini in Chat lets users search Gmail, Drive, and Calendar and complete tasks without leaving the conversation. Which of our team's daily research, drafting, and scheduling workflows currently generate web queries or tool switches that will now resolve inside Chat, and does that shift change what we can measure about how work gets done?
  • Google has placed Gemini inside Chat as a direct capability competitor to Copilot in Teams. If your organisation runs both Workspace and Microsoft 365, how does Ask Gemini change your evaluation of which collaboration platform delivers the richer AI capability for your team's specific workflows?
  • Elevated usage limits for Ask Gemini run through 1 October 2026. Have we identified which workflows to test during this promotional window, and is there a named owner accountable for evaluating Ask Gemini's fit before the limits expire and standard rates apply?
Sources
05 · Anthropic

Anthropic makes Sonnet 5's launch pricing permanent, cancelling the planned 50% increase

What
On 10 August 2026, Anthropic confirmed that Claude Sonnet 5's introductory pricing of $2 per million input tokens and $10 per million output tokens is now the permanent standard price. The standard pricing of $3 per million input tokens and $15 per million output tokens, which had been scheduled to take effect on 1 September 2026 at the close of the stated introductory window, will not apply. Anthropic made the decision public via the Claude AI account on X and updated its developer documentation accordingly. The permanent pricing applies to all API access and metered enterprise usage of Sonnet 5.
When
Announced via the Claude AI account on X on Sun 10 August 2026. Confirmed in Anthropic developer documentation and reported by Enterprise DNA on 10 August 2026.
How it shifts discovery
When Sonnet 5 launched in late June, the introductory pricing was stated to be temporary, with a 50 per cent increase across both input and output tokens scheduled for September 1. Cost models built on that assumption are now wrong in the organisation's favour. The more significant signal is structural: Anthropic cancelling the increase at a moment when the frontier model pricing market is compressing from multiple directions confirms that the direction of travel is set. Any workload cost model that included the September 1 Sonnet 5 price increase, or that used the anticipated increase to justify a migration to a cheaper alternative before 31 August, is built on a planning assumption that no longer exists. Rebuild those models before committing September delivery on any migration. The practical implication extends further: Sonnet 5 at permanent $10 output per million tokens becomes the mid-tier benchmark against which every subsequent Anthropic model pricing decision will be judged. It also reopens the model selection economics for any agentic workload previously considered too expensive to run at Sonnet tier but not worth the cost of a more powerful alternative.
Questions to ask
  • Sonnet 5's pricing is now $2 input and $10 output per million tokens permanently, cancelling the 50 per cent increase planned for 1 September. Have we rebuilt the cost model for any Sonnet 5 workload where the September increase was a factor in a migration or cancellation decision that is currently in progress?
  • Multiple frontier model providers have moved pricing downward this quarter. Has our model selection framework been updated to reflect current pricing rather than the rates in effect at the time it was last reviewed, and does Sonnet 5 at permanent $10 output now become the default selection for mid-complexity agentic workloads?
  • Anthropic made the pricing permanent rather than extending the introductory period, signalling confidence in the rate as a sustainable commercial position. Does that stability signal change our assessment of Claude API pricing predictability compared with providers who have repriced more aggressively or less predictably this year?
Sources

Key takeaways

What to walk away with this week

  1. Salesforce choosing Claude as its default reasoning engine is the strongest signal yet that enterprise AI has moved from experimentation to embedded infrastructure: audit what your Salesforce instance exposes through the Claudeforce plugin before September's open beta opens access to your organisation.

  2. OpenAI's Jalapeño chip posted a 1.9x efficiency lead over Nvidia Blackwell with independent confirmation from SemiAnalysis. Custom silicon at both OpenAI and Anthropic will erode Nvidia's cost floor from late 2026: build a scenario for materially lower frontier API pricing into your 2027 AI cost model now, while the planning assumption is still easy to change.

  3. Perplexity's Portable Computer, zero cost for on-device work with explicit opt-in for cloud steps, directly addresses the governance objections that have stalled AI agent deployments in regulated environments: test whether a local-first architecture resolves the data residency blocker in your highest-value stalled use case this quarter.

  4. Google Chat and Microsoft Teams are now competing on which platform resolves more work through a single AI conversation. The collaboration tool your team already uses is becoming the primary work-completion layer: measure what stays inside it and what still generates web queries, before the shift is complete and baseline data is gone.

  5. Anthropic's decision to make Sonnet 5 pricing permanent rather than raise it confirms frontier AI pricing is compressing structurally. Rebuild any cost model built on the assumption of a September 1 increase, and re-evaluate any migration decision made in anticipation of it before committing September delivery.

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.