RU

LLM Harness: How to Turn a Bare Model into an AI Service

Published: 2026-09-30 · Author: AI Release · @ai_release1
LLM Harness: How to Turn a Bare Model into an AI Service

⚡ The gist in 5 seconds - Downloading and running an LLM on a GPU is only the first step: the model knows nothing about the user, the context, or access rules. - A working service needs a harness: an interface/API, RAG with corporate data, tool integrations, limits, and monitoring. - The catch: out of 33 enterprise AI pilots, on average only four make it to production — 88% remain experiments. ### 🔍 What was found The article was published on Habr on September 30 at 12:00 by natlysky in the Selectel company blog. Reading time is 7 minutes, and the piece has racked up 12K views. The author explains: an LLM downloaded from Hugging Face and deployed on a GPU server technically already works — it receives text and generates a response. But it doesn't know who sent the request, whether the user is allowed to see the requested data, where to save the result, or whether an action can be performed. Even conversation history isn't built-in "memory": the application must store messages and pass them back in the context with each new request. What goes into the harness? The set of components depends on the task, but several layers appear almost always. These are an interface or API (Open WebUI provides a ready-made web interface, with Ollama deployed alongside for running models locally), context and corporate data via RAG (Open WebUI has built-in RAG, RAGFlow specializes in this), integrations and actions via tool calling and Structured Outputs with JSON Schema, plus low-code tools like n8n for connecting to CRM, ERP, email, and chats. A separate layer covers limits, cost, and fault tolerance: capping steps and tokens, caching, distributing tasks across models, and a fallback model. The article also mentions the Selectel AI router: access to 300+ models from a single panel, a unified API key, end-to-end analytics, limits, and quotas. ### 💡 Why it matters The model is the engine; the harness is the whole car. Without a harness that decides who to trust, what to allow, and when to hand a decision to a human, an agent can get stuck in a loop on a failed request, forget about access restrictions, and allow unauthorized interference with the system's logic. Research confirms: out of 33 enterprise AI pilots, on average only four reach production, and the model is rarely the bottleneck. This means choosing an LLM is just one stage of building an AI service, and the bulk of the work lies in the harness. The practical benefit is being able to estimate the scope of work upfront and understand which components will be needed before investing in deployment. ### 🧩 Context Between a working LLM and a finished service lies a large layer around which authentication, data handling, integrations, logging, and monitoring emerge. This is the point where an experiment with a model starts turning into a product. The article emphasizes that a bare model cannot be called a business service: the user needs an interface, the application needs an API, and the model needs context and corporate data. The author presents "without harness" and "with harness" diagrams, showing that it is precisely at the interface level that the LLM

🔗 Read on habr.com

🤖 AI summary
llmai-инфраструктураharness
📖
Read the guide on this topic
Read →
← PreviousWindows 11: which background services slow down your PC and what to disableNext →Jupiter: How a Non-Programmer Built Dictation for Windows Using ChatGPT

Source: habr.com · post in Telegram