Your AI.
Running on your device.
Cortex runs open-weight language models directly on your Android phone, natively, through llama.cpp. Your conversation never leaves the device.
No spam. One email the day it opens up.
What Cortex is
A local model, not a login screen.
Most AI apps on your phone are a thin client. You type, it uploads, a datacenter answers. Cortex isn't that. It downloads an open-weight model onto your storage and runs it on your own processor.
No account
There is no sign-up, no profile, and no server holding your history. Install it and start typing.
No subscription
The models are open weights from Hugging Face. You download them once; nothing meters your usage.
No connection needed
Once a model is on the device, Cortex answers in airplane mode. Inference makes no network call.
Who builds it
Cortex is an independent project, built by a small team who thought the most interesting thing about small language models was that they had quietly become good enough to run on a phone — and that almost nobody was letting people actually do it. We're building in the open on Instagram, posting the work as it happens.
How it works
Three steps, then the network stops mattering.
- 01
Pick a model
Search Hugging Face from inside the app. Cortex shows you which GGUF files will fit your device's memory.
- 02
Download it once
The weights land in your app storage. This is the only part that touches the network, and you can do it on Wi-Fi and then forget it.
- 03
Talk to it anywhere
llama.cpp loads the model into memory and generates on your CPU. Underground, mid-flight, or with the SIM pulled out.
Where a prompt goes
Model hub
One app. Pick your brain.
Browse model folders by what you want to do, download the GGUF, and give each conversation whichever model suits it.
LLMs
Runs todayGeneral chat & reasoning
Qwen2.5, Llama 3.2, Gemma 2, Phi-3, SmolLM2, Mistral
Coding
Runs todayCode generation & completion
Qwen2.5-Coder and other code-tuned GGUF builds
On the roadmap
Image
Image generation
Needs a diffusion runtime alongside llama.cpp
Video
Video generation
Not feasible on-device at useful quality yet
Speaking
Speech & voice synthesis
Dictation works today; synthesis is not wired in
Cortex generates through llama.cpp, which executes text GGUF models. The roadmap folders appear in the app too, marked unsupported until the runtimes for them exist on-device — we'd rather show you the roadmap than pretend.
Under the hood
The spec sheet, without the marketing.
- Engine
- llama.cpp
- native, compiled in
- Weights
- GGUF
- quantized, your choice of level
- Architecture
- arm64-v8a
- 64-bit ARM devices
- App size
- 35 MB
- models downloaded separately
- Android
- 8.0+
- Oreo through 15
- Inference egress
- None
- no network call to answer you
One honest caveat. Downloading a model and connecting an MCP server both use the network — that's unavoidable. The claim we make is narrower and exact: generating a reply never makes a network call.
Test before launch
Help get Cortex onto the Play Store.
Google runs new apps through a round of closed testing before they can go live. If you want Cortex in your hands early, this is the fastest way to get it.
- 01
Send your Play Store email
It has to be the address your Play Store is signed in with — that's what the tester list is checked against.
- 02
We add you to the closed test
You'll get an opt-in link by email. Open it on your phone, accept, and Cortex appears in the Play Store for you.
- 03
Install and use it normally
Updates arrive through Play like any other app. Tell us what breaks — that's the whole point.
Become a beta tester
Add the address your Play Store is signed in with and we'll send you the opt-in link.
Use the address on your Play Store account.
Your address is used only to add you to the tester list. Nothing else, and you can leave the test any time from the Play Store.
Shipping soon
Be first.
Cortex v1.0.5 is built and running on real devices. Register now and you'll get the download the day it opens up.
No spam. One email the day it opens up.