How to run an LLM on your phone with Puma Browser

You do not need a desktop GPU to try on-device AI. Puma Browser runs open LLMs on your phone for private chat and page help.

How to run an LLM on your phone with Puma Browser

Running a large language model on a phone used to sound exotic. Today it is practical for everyday browsing: short chats, page summaries, and follow-up questions without shipping every token to the cloud.

1. Install Puma

Download Puma for iOS or Puma for Android. Open the app and grant storage/network permissions you are comfortable with — local models still need an initial download.

2. Choose a local model

In Puma, pick an on-device model such as Gemma, Qwen, or Ministral. Smaller models respond faster on older phones; larger ones quality-up when you have RAM and patience for the download.

3. Chat without leaving the device

Start a local conversation. Prompts and replies stay on your phone when you use a local model — useful for private drafts, travel notes, or anything you would not paste into a cloud chat box.

4. Summarize the page you are on

Browse to an article, docs page, or long thread. Ask Puma to summarize or explain. Local mode keeps that workflow on-device; if you need a stronger model, switch to OpenAI, Anthropic, or Gemini APIs from the same browser.

5. Go offline (after models are cached)

Once the model weights are on the phone, basic local chat and page help can work without connectivity — handy on planes and spotty networks.

Tips for better on-device results

  • Prefer clear, short prompts over giant paste dumps.
  • Use cloud models for heavy coding or long research; keep local for private everyday tasks.
  • Update Puma when new open models ship — the local catalog improves quickly.

Want the privacy angle? Read Puma as a private local AI browser. Curious about decentralized sites? See browsing ENS, IPFS, and Unstoppable Domains.

— The Puma team