All terms

AI Model & Product

Google AI Edge

Also known as: AI Edge, Google on-device AI

Google AI Edge is Google's collection of tools and runtimes for running AI models directly on devices such as phones, browsers and embedded hardware rather than in the cloud. It includes components like LiteRT for on-device inference, MediaPipe tasks for common vision, audio and text jobs, conversion tooling for models built in other frameworks, and APIs for on-device generative models. Developers use it to build features that work offline, respond quickly and keep data local.

What it is

Google AI Edge is an umbrella for Google's on-device machine learning stack, spanning model conversion, optimisation, runtime execution and debugging tools. It covers classic tasks such as image classification, object detection, speech and text handling, as well as running compact generative models on device. It targets Android, iOS, web and embedded targets, with hardware acceleration where the device supports it.

Why it matters

As more AI moves to the device, a growing share of user questions can be answered without a server request, a search query or a page visit. That shifts measurement, because on-device answers leave little trace in analytics, and it raises the value of being the source that on-device systems retrieve from when they do go online. It also affects product teams, since low latency and offline capability change what experiences are possible in apps and stores.

How it works

Developers take a trained model, convert and quantise it for the target runtime, then call it from an app through an inference API, often combining it with local retrieval over app data. Teams profile memory, latency and accuracy trade offs across device tiers, and fall back to cloud models for heavier tasks. Marketing and SEO teams engage indirectly by keeping factual brand information structured, current and retrievable, and by testing how on-device assistants describe their products.

When it applies

Applies when building mobile or embedded AI features, when privacy or connectivity constraints rule out cloud inference, or when assessing how much of your audience journey is being handled on device.

Examples

  • A retail app runs on-device image recognition so shoppers can point a camera at a product and get a match without a network connection.
  • A field service app uses a compact on-device model to summarise engineer notes while working in buildings with no signal.
  • A browser based tool runs a small model locally to redact personal data before anything is uploaded.

How it is measured

  • Proportion of inference requests served on device versus cloud
  • Median and p95 inference latency per device tier
  • Model size and peak memory use after quantisation
  • Offline task completion rate and associated support ticket volume

Related terms in AI Model & Product

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.