Vision: Ask Questions About Images, On-Device
A local vision-language model reads receipts, whiteboards, error dialogs, forms and handwriting, and answers in the same conversation. The image never leaves your phone.
Attach a photo or a screenshot and ask about it. A vision-language model loaded on your phone reads the image and answers in the same conversation. The picture is processed in your phone's RAM and is never uploaded.
What it reads well
Receipts and invoices. Whiteboards after a meeting. Error dialogs and unfamiliar screens. Forms, letters and printed documents. Handwriting, within reason. Charts and diagrams, at the level of "what is this showing". Labels in a language you do not read.
It is not a specialised OCR engine and it will not transcribe a dense page of small print perfectly. It is very good at "what does this say and what does it mean", which is usually the actual question.
Why local matters more here than anywhere else
Think about what people actually photograph when they want an AI to explain something: a payslip, a prescription, a medical letter, an ID document, a contract, a bank statement, a child's homework. Every one of those is a document you would think twice about uploading to a website — and with most assistants, uploading it is the only way to ask. Here there is nowhere to upload it to.
The models
The catalog carries SmolVLM2, InternVL2, Qwen3-VL 2B and 4B, moondream2, nanoLLaVA and MiniCPM-V, matched to your device tier. A vision model is loaded in place of the text model rather than beside it — see one heavy model at a time — so the swap is explicit and visible, and your conversation is saved before it happens.
Availability
Vision is one of the four things behind the one-time $2.99 unlock in MyBenAI Lite, included in the seven-day trial every new install gets, and included unconditionally in the $2 paid app.
Looking for something else? Every page on this site.