RU

Running LLMs locally on Windows: WSL2, Docker, CUDA and vLLM — a stack breakdown

Published: 2026-10-06 · Author: AI Release · @ai_release1
Running LLMs locally on Windows: WSL2, Docker, CUDA and vLLM — a stack breakdown

⚡ The gist in 5 seconds - The point: deploying the Qwen3-0.6B LLM on Windows 11 via WSL2, Docker, NVIDIA Container Toolkit and vLLM, with a full breakdown of every layer of the stack. - Where it works: on a regular NVIDIA GPU with 4 GB of VRAM — the author tested on a GPU around the RTX 3050 level. - Requirements: WSL2, an NVIDIA driver compatible with WSL2, at least 16 GB of RAM and 30–50 GB of free disk space. ### 🔍 What was found A translation of Harshit Kumar's guide, published on Habr, describes the full stack for running Qwen3-0.6B locally on Windows 11. The chain includes WSL2, Docker Engine, NVIDIA Container Toolkit and vLLM. The author points out that most tutorials give you just one or two commands that work until the first error. After that, it's unclear at which level the problem occurred: Windows, WSL2, the NVIDIA driver, Docker, CUDA, or the inference server itself. That's why the guide proposes verifying the infrastructure bottom-up. The test configuration is a regular NVIDIA GPU with 4 GB of VRAM, roughly RTX 3050 level. On such a card you can realistically run a 0.6-billion-parameter model, but you need to account for memory spent on weights, KV cache, the CUDA context and runtime buffers. The steps start with installing WSL2 and Ubuntu via wsl --install and checking the GPU with nvidia-smi. If nvidia-smi doesn't work, there's no point moving on to Docker — the problem is lower in the stack. After updating Ubuntu, Docker Engine is installed from the official repository, with a GPG key added. Then the NVIDIA Container Toolkit is installed, which configures the Docker runtime and gives containers access to the GPU. The key check is running the nvidia/cuda:12.8.1-base-ubuntu24.04 CUDA container with nvidia-smi inside. Only after this check passes is the vllm/vllm-openai:latest image pulled, with vllm serve set as the entrypoint. The author separately explains the difference between the Docker image (Python, PyTorch, CUDA libraries, vLLM — several gigabytes) and the model cache (~/.cache/huggingface with weights, tokenizer and config). It's convenient to keep them separate: when recreating a container, you won't have to download the model again. ### 💡 Why it matters The practical value of the guide isn't just getting answers from a model, but understanding how the entire stack is built. This approach lets you diagnose errors at any level yourself, swap models, and gradually move toward production-like infrastructure. Instead of a black box of one command, you get a transparent chain where every component can be checked separately. If something breaks, you know exactly where to look. ### 🧩 Context The piece is a translation of an article by Harshit Kumar, published on Habr. It falls under the topics of machine learning, DevOps, GPGPU and Linux. The author starts from the observation that most instructions for running LLMs locally boil down to one or two commands. That approach works until the first error, after which it becomes unclear at which level it occurred. To fix this, the guide deploys Qwen3-0.6B on Windows 11 and explains the role of each component: WSL2, th

🔗 Read on habr.com

🤖 AI summary
LLMWindowsWSL2DockerCUDAvLLM
📖
Read the guide on this topic
Read →
← PreviousWindows 11 26H2: 3 GB less RAM at idle and faster app launchesNext →Gigawatts, custom chips and multi-cloud: how OpenAI, Anthropic, Google and xAI are fighting over the platform

Source: habr.com · post in Telegram