Conversation · Free Included in the free edition

Offline AI Chat That Runs on Your Phone

Full multi-turn conversation with a local LLM running in your phone's own memory. No account, no API key, and identical behaviour in airplane mode.

MyBenAI on Android: a full answer generated on-device at 17.5 tokens per second
Android a full answer generated on-device at 17.5 tokens per second
MyBenAI on iOS: the chat home screen on iPhone
iOS the chat home screen on iPhone

MyBenAI's chat is an ordinary assistant conversation with one unusual property: the model producing the words is running in your phone's own memory, through llama.cpp. There is no request going anywhere. Put the phone in airplane mode and the behaviour is identical, because nothing about it depended on the network in the first place.

What the conversation gives you

Streaming replies as they are generated, a stop button that actually stops decoding rather than hiding the result, regeneration, editable messages, and conversations you can rename, group into folders and search. The model currently loaded is named at the bottom of the composer, next to the tool and attachment buttons, so you always know which one is answering. Under each reply the app reports the speed it managed — the screenshot above shows 17.5 tokens per second from a 2B model on a mid-range Android phone.

What "on-device" costs and what it buys

The honest trade is capability. A 2B or 4B model is not a frontier model, and it will not write your dissertation. What it will do is summarise, rewrite, explain, translate, brainstorm, extract structure from a mess, answer questions about documents you gave it, and drive tools — all without a subscription, an account, or your text becoming somebody's training data.

In exchange you get things a cloud assistant structurally cannot offer: it works on a plane and in a basement, it costs nothing per message, there is no rate limit, no queue and no capacity incident, and the conversation cannot be read by anyone because it never leaves the phone.

Sessions are kept warm, not held hostage

After five idle minutes the app releases the model's context and persists its state, so the phone gets its RAM back rather than holding a gigabyte hostage for a conversation you may not return to. Coming back restores the session in a couple of seconds instead of reloading from scratch. This is also what keeps the operating system from killing the app while it sits in the background.

No account, no key, no setup

There is nothing to sign up for and no API key to paste. On first launch the app measures your phone and downloads a model that fits it. From there you open it and type.