One Heavy Model at a Time
Text, vision and diffusion models are never resident together. Swaps are explicit and visible, and your conversation is saved before the context is released.
The single most important rule in this app: only one heavy model is in memory at a time. Text, vision and diffusion are never co-resident.
Why the rule exists
A phone gives an app a memory allowance and terminates it without warning for exceeding it. A 2B chat model plus a vision model plus a diffusion pipeline is not close to fitting anywhere. An app that loads them optimistically works beautifully in a demo and is killed mid-sentence on a real phone with a browser and a messaging app already open.
What a swap looks like
Explicit and visible. Your conversation is saved first, then the previous context is released, then the next model loads — and the app shows you it is happening rather than appearing frozen. Attaching an image to a chat swaps the text model out for a vision model; asking for a picture swaps in the diffusion pipeline. Nothing is lost across the swap.
Idle release, and warm restarts
After five idle minutes the model context is released and its state persisted, so the phone gets its RAM back. Returning restores the session in about two seconds instead of a full reload. The Whisper context is released after sixty seconds on the same principle.
The budget it swaps within
The RAM ceiling is 45% of total on Android and 50% on iOS, minus 600 MB of headroom — which is how the selector decides what your phone can hold in the first place. Swapping is what lets the app offer more capability than fits, without ever exceeding the budget.
Looking for something else? Every page on this site.