Cortiq Mobile Runs Neural Networks on Your Phone in the CMF Format
Published: 2026-10-02 · Author: AI Release · @ai_release1
⚡ The Gist in 5 Seconds - Cortiq Mobile is an open-source app for Android and iOS (not yet on the App Store, still under review) that runs models in the CMF format directly on your phone. - Three offline scenarios: chatting with document reading, translating emails and messages, and delivering ready-made decisions for programs instead of text replies. - Benchmarks were taken on TestFlight build 1.3.0 (47): lfm2.5-230m-q4tp produced 95 tokens in 0.6 seconds on an iPhone — 162.1 tok/s. The app is available on Google Play, and the source code for version 1.3.0 is at infosave2007/cmfmobile. ### 🔍 What Was Found The secret to its autonomy lies in the file format. Models are packaged in CMF (Cortiq Model Format): a single file containing not only the weights, but also the architecture, tokenizer, task descriptions, and chat template. The file describes itself, so the app knows before loading how many layers the model has, what its context is, and how much memory it needs — and displays this in the library. In the catalog, each file shows its purpose — Chat or Decisions. There are three ways to get a model: a ready-made.cmf from the catalog in one tap (with resuming from the last byte if the connection drops), a repository from Hugging Face in safetensors — the app converts it to CMF right on the phone — or GGUF via the desktop utility cortiq import-gguf. The same format is read by the Cortiq engine on computers. Benchmarks on iPhone 16 (A18): the lfm2.5-230m-q4tp model — 230 million parameters, 14 layers, a 126 MiB file, running on CPU with five threads, no GPU needed. Generation parameters: temperature 0.10, top_p 0.95, max_tokens 128. The chat keeps each conversation as a separate thread, with responses streaming in along with timing, token counts, and tok/s; documents in txt, md, code, and JSON are supported up to 512 KiB. Translation doesn't require a chat model: the catalog includes Hy-MT2 with 1.8 billion parameters, in two files — q4tp at 950.5 MB (perplexity 14.02) and q1t at 694.9 MB (perplexity 27.88). All seven test translations came back complete with finish_reason: stop, JSON survived translation without loss, and the engine delivered 14.4–21.8 tok/s (17.75 on average). The "Server" tab, protected by a token and QR code, exposes an OpenAI-compatible API: /v1/chat/completions, /v1/completions, /v1/models, while the decisions model responds via /v1/decide with a structure like {"accepted": true, "choice": "card_arrival", "confidence": 0.998} or an honest accepted: false. ### 💡 Why It Matters The practical benefit is working where there's no network or it's limited. On a plane, the app answers questions, translates emails, and parses text files. At home, the same phone provides a laptop with a local API, and the decisions model returns a structured result to a program instead of a paragraph from which the code would later have to extract meaning with regexes. For code, "I don't know" isn't an error but a clean signal to hand the task to a human. Meanwhile, the data never leaves the device. ### 🧩 Context The author emphasizes that phone performance in some cases exceeds PC performance, so small LLMs can quite reasonably be used for work. One model in CMF works across all devices: both on the phone and on the computer. All the source da
⚡ The Gist in 5 Seconds - Cortiq Mobile is an open-source app for Android and iOS (not yet on the App Store, still under review) that runs models in the CMF format directly on your phone.
- Three offline scenarios: chatting with document reading, translating emails and messages, and delivering ready-made decisions for programs instead of text replies.
- Benchmarks were taken on TestFlight build 1.3.0 (47): lfm2.5-230m-q4tp produced 95 tokens in 0.6 seconds on an iPhone — 162.1 tok/s.
The app is available on Google Play, and the source code for version 1.3.0 is at infosave2007/cmfmobile.
🔍 What Was Found The secret to its autonomy lies in the file format.
Models are packaged in CMF (Cortiq Model Format): a single file containing not only the weights, but also the architecture, tokenizer, task descriptions, and chat template.
The file describes itself, so the app knows before loading how many layers the model has, what its context is, and how much memory it needs — and displays this in the library.
In the catalog, each file shows its purpose — Chat or Decisions.