curl -fsSL https://ollama.com/install.sh | sh; the status is checked via systemctl status ollama. The daemon listens on 127.0.0.1:11434 by default. The CLI resembles Docker: pull, run, rm, and ps commands are available. The model is downloaded from Hugging Face via ollama pull hf.co/ggml-org/gemma-4-26B-A4B-it-GGUF:Q4_0. To test the API, use curl http://localhost:11434/api/tags and a POST to /api/chat with "stream": false to get a single JSON response. llama.cpp requires building from source. Install git, cmake, build-essential, libssl-dev, verify nvcc is present (nvcc --version), and install nvidia-cuda-toolkit if needed. After cloning the repository, run cmake -B build -DGGML_CUDA=ON and build with -j$(nproc). GPU availability is checked via llama-server --list-devices. The model is loaded with llama-cli -hf ggml-org/gemma-4-26B-A4B-it-GGUF:Q4_0. The server starts with the parameters --n-gpu-layers all, --ctx-size 4096, --parallel 1, host 0.0.0.0, port 8080. The API is OpenAI-compatible and accepts requests at http://localhost:8080/v1/chat/completions. ### 💡 Why it matters The choice between Ollama and llama.cpp often turns into a religious war in the comments. This article is useful because it offers not abstract arguments but concrete steps: how to install, which command to run, how to test the API. A reader can deploy both options in 10 minutes on the same machine with the same model and decide for themselves what suits them better — Ollama, which hides complexity, or llama.cpp, which is flexible but requires manual building. This approach reduces reliance on other people's opinions and provides the facts needed to make your own choice. ### 🧩 Context Ollama launched in 2023 as a layer on top of llama.cpp, got its own engine for multimodal models (working with GGML) in 2025, and in 2026 returned to llama.cpp for GGUF support. The Gemma 4 model is a fully open (Apache 2.0) development from Google that appeared this spring. The test uses the MoE variant 26B A4B: 26 billion parameters, but only ~4 billion active per token, making the model more resource-efficient. Inside this architecture there are several experts, and only a subset is engaged for each request, striking a balance between response quality and resource consumption.Source: habr.com · post in Telegram