RU

Local Neural Networks Without Python and CUDA: Ollivo on Vulkan Outpaces CUDA

Published: 2026-09-27 · Author: AI Release · @ai_release1
Local Neural Networks Without Python and CUDA: Ollivo on Vulkan Outpaces CUDA

⚡ The gist in 5 seconds - Ollivo is a Windows app for running local neural networks like a regular program: install it, pick a model, start typing. The internet is only needed to download the model. - An early version 0.3 is available, and the code is open source. The llama.cpp, whisper.cpp, and ComfyUI engines install and repair themselves, and the app updates itself. - Limitation: no images in the window yet — that's the next major release. The author hasn't yet tested Vulkan performance on RTX cards. ### 🔍 What was found On September 27, author WufCorp published a case study on Habr about developing Ollivo. The test machine is a GTX 1080 with 8 GB of VRAM and 32 GB of RAM. On Qwen2.5 3B, he compared llama.cpp builds: CUDA 12 delivered 52 tok/s with a footprint of 242 + 373 MB, CUDA 13 came in at 143 + 403 MB and "crashes," while Vulkan was 30 MB and 81 tok/s. Vulkan generates one and a half times faster than CUDA, is twenty times smaller, and doesn't require CUDA libraries. However, it reads long prompts more slowly, so Vulkan is the default for chat, while the CUDA build is enabled based on the GPU generation. A built-in "traffic light" estimates speed even before the model is downloaded. For a GGUF file on disk, the header is read: architecture, number of layers, compression, context size. For models from the catalog, the estimate is rougher: 700 MB is added to the weights plus one-eighth of their size. Speed is calculated from video memory bandwidth: roughly 60% divided by the model size. For Llama 3.1 8B Q4 on a GTX 1080, that comes out to about 37 tok/s — close to reality. For MoE models, speed is calculated based on the active portion of the weights, otherwise a 35-billion-parameter model would appear four times slower. If the result is less than three tokens per second, the app warns you before downloading. ### 💡 Why it matters Ollivo demonstrates in practice that CUDA isn't required for local neural networks. The Vulkan build of llama.cpp runs on older Pascal cards where CUDA 13 won't launch, and it turns out to be faster. For the average user, this removes the main barrier: no need to install Python, deal with CUDA, or choose between Q4_K_M and Q5_K_S. The app automatically finds models already downloaded in LM Studio, Ollama, or ComfyUI and doesn't copy them — saving tens of gigabytes. The window doesn't contain the word "tokens": speed is shown in words per second and compared to reading speed, and conversation memory is measured in pages. Instead of "Q4_K_M," it says "the usual choice: nearly identical to the original, but half the size." ### 🧩 Context The author took on the project because he himself went through Python, CUDA, torch builds specific to his video card, and a virtual environment that breaks when you rename a folder. The second reason: the people around him shouldn't have to know the words "quantization," "context," and "sampler." The third: everything is scattered across different programs — text in LM Studio or Ollama, images in ComfyUI, speech recognition separately. The project's main rule: if something can't be explained in one clear sentence, it doesn't make it into simple mode. The shell is written in Tauri 2 and React, the core is in Rust; the installer is about 10 MB instead of the ~150 MB typical for Electron. Python is only needed for th

🔗 Read on habr.com

🤖 AI summary
локальныеllama.cppVulkan
📖
Read the guide on this topic
Read →
← PreviousAndroid Student Management: a reference repository for app architectureNext →SLM: Where Small Language Models Come In Handy and What to Consider

Source: habr.com · post in Telegram