gremlin.group / blog/ on-device-ai-for-business

PhonicAIBusinessStrategy

On-Device AI: When You Do Not Need the Cloud

Most AI tools send your data to someone else to process. Some of the most useful jobs can now run on the laptop or phone you already own. Here is when that makes sense, and when it does not.

When most people talk about using AI at work, they mean a cloud service. You type something into a chat window or connect a tool to an API, your data travels to a provider’s servers, a large model processes it there, and the answer comes back. It works well, and for many jobs it is the right choice.

It is no longer the only choice, though. Phones, tablets and laptops now ship with hardware that can run smaller AI models directly. Transcription, summarising, sorting and searching can often happen on the device in your hand, without the data going anywhere.

For a business deciding how to use AI, the difference matters. It affects where your data goes, what you pay, and what happens when the Wi-Fi drops. This post covers the trade-offs in plain terms and finishes with a simple way to decide.

What “on device” actually means

Cloud AI runs on the provider’s computers. You send the input, they run the model, you get the output. The models can be very large, because the provider has racks of specialised hardware to run them on.

On-device AI runs on your own hardware. The model is downloaded to the phone or laptop once, and after that the processing happens locally. The models are smaller, because they have to fit on a consumer device and run without draining the battery.

They suit different jobs.

Privacy

This is usually the strongest argument for on-device processing.

When data is processed in the cloud, it leaves your control for a while. A reputable provider will have contracts, security controls and retention policies, and for most data that is fine. But it is still a third party handling your information. Under UK GDPR, if that data is personal data, the provider is acting as your processor, and you need to be comfortable with the contract, where the data is held and how long it is kept. The ICO publishes guidance on this. None of this is legal advice, but it is worth knowing that cloud AI adds a supplier to the picture.

When the processing happens on the device, the sensitive part never goes to that supplier. A recorded client meeting, a staff disciplinary conversation, a patient note: if the AI step runs locally, there is one less party to vet and one less place the data can leak from.

On-device does not remove every question. The device still needs to be secured, and where the results are stored afterwards still matters. But it narrows the problem considerably.

Cost

Cloud AI is usually billed by use. Every request, every document and every minute of audio has a price. For light use the numbers are small. For a team processing hundreds of documents or hours of recordings a week, they add up, and they scale with how much people use the tool. Success makes the bill bigger.

On-device AI has no per-request charge. Once the model is on the device, running it again costs nothing beyond the electricity. The cost is upfront: the hardware, and whatever you pay for the software.

That makes budgeting simpler. It also changes behaviour. People use a tool more freely when nobody is watching a usage meter.

Offline use and speed

A cloud service needs a connection. On a train, in a client’s basement meeting room or at a site with poor signal, cloud AI either fails or waits. On-device processing carries on regardless once the model is downloaded.

There is also the round trip. Sending data to a server and waiting for a reply adds delay, and that delay varies with the connection and with how busy the service is. Local processing avoids that. For short, frequent tasks, it often feels quicker. For heavy tasks on modest hardware, it can be slower, which leads on to the limits.

The real limits

On-device AI has genuine drawbacks, and it is worth being clear about them.

A model that fits on a phone is less capable than one running in a data centre. For well-defined jobs like transcription or a summary of a meeting, smaller models do well. For open-ended reasoning, long complex documents or tasks that need broad general knowledge, the large cloud models are still ahead.

On-device features usually need recent hardware and a recent operating system. Older laptops and phones may not support them at all. If your team runs on a mix of ageing devices, that is a real constraint.

The model has to get onto the device before it can run. Some are large. Downloading them over mobile data, or at the moment you need them, is a poor experience.

Consistency can suffer too. A cloud service behaves the same for everyone, while on-device results can vary between devices depending on what each one supports.

A worked example: Phonic

Phonic is our first app, a recorder and transcriber for iPhone, iPad, Mac and Apple Watch. It sits firmly on the on-device side, so it is a fair test of the trade-offs above, including the awkward ones.

Recordings are transcribed on the device using Apple’s speech models. With Phonic Pro, speakers are also identified on the device. Phonic does not ask you to create an account with us, and recordings are saved to your own iCloud Drive, or to a library on the device if you do not use iCloud. Details are in the Phonic privacy policy.

The limits show up too. Phonic requires iOS 27, iPadOS 27, macOS 27 or watchOS 27, so devices that cannot run those versions are out. The first transcription in a language downloads Apple’s speech model for that language, and with Phonic Pro the first speaker identification downloads speaker models of about 120 MB. Both are reused afterwards, and our help article on models suggests downloading them on Wi-Fi before your first long meeting. Recordings longer than an hour use clustering alone to tell speakers apart, which can merge similar voices. It is exactly the kind of capability limit worth knowing about before you rely on a tool.

The Insights tab, part of Phonic Pro, produces a summary, key points, action items, decisions, open questions, topics and talk time for each recording. It uses Apple Intelligence when it is available on the device, with a built-in fallback when it is not. That is a sensible pattern for any on-device feature: use the better capability where the hardware supports it, and keep working where it does not.

Cost follows the same logic. Phonic is free to download, and Phonic Pro is a one-off £2.99 in-app purchase. There is no per-minute transcription charge, because there is no server doing the transcribing. More on the app is on the Phonic page.

A simple decision guide

Our line on AI work is that if a spreadsheet would do the job, we will say so. The same thinking applies here. Start with the job, not the technology.

On-device is a good fit when:

  • The data is sensitive and you would rather it did not go to a third party
  • The task is well defined: transcribing, summarising, classifying, searching
  • Usage is high or unpredictable, and per-request billing would be hard to budget
  • People need it to work without a reliable connection
  • Your team already has recent devices

Cloud is a better fit when:

  • The task needs the most capable models: complex reasoning, long documents, broad knowledge
  • Results need to be consistent across everyone, regardless of device
  • The work runs as a central process rather than on individual laptops
  • Volume is low, so per-request costs stay small
  • Your devices are older or mixed

Neither is needed when:

  • The rules are fixed and known. A formula, a filter or a simple script will be cheaper, faster and easier to check than any model.

Many businesses end up with both. Sensitive, routine work stays on the device. Occasional heavy lifting goes to the cloud, with the right contract in place. Louise’s practical guide to AI for business and her look at how small businesses are using AI cover where it tends to pay off.

The practical takeaway

The cloud is not going away, and for hard problems it is still where the best models live. But a lot of everyday work, the transcribing and summarising and sorting, can now stay on the device. You keep the data local and stop paying per request. You give up some capability and need reasonably new hardware.

The right answer depends on the job, the data and the devices your team already uses. If you are weighing up where AI fits in your business and want a straight answer, including “you do not need it”, get in touch.

M Written by Michael, Co-Founder Cloud architect and co-founder of Gremlin Group. Spends most of his time designing AWS infrastructure and writing about cloud architecture, cost optimisation, and DevOps.

Got a process that eats your week?

Tell us about it. If AI is the right fix we will build it, and if a spreadsheet would do, we will say so.