RU

Stitching Together Truncated LLM Responses: How We Complete Them

Published: 2026-09-25 · Author: AI Release · @ai_release1
Stitching Together Truncated LLM Responses: How We Complete Them

⚡ The gist in 5 seconds - The point: in a Russian-language service running multiple language models in a single chat, a model cut off its response mid-word due to the max_tokens cap, and restoring coherence required building a continuation-and-stitching mechanism. - Where it's available: the case study was published on Habr on September 24, reading time — 3 minutes, views — 8K. - Limitation: no more than four continuations, so a single question doesn't turn into an endless back-and-forth; the cutoffs themselves haven't gone away. ### 🔍 What was found Every request to a model has a max_tokens cap. In the service it was set to 1200 tokens: in English that's roughly 900 words, but Cyrillic consumes more tokens, so 1200 tokens covered about fifteen hundred characters. When a response hits the cap, the model returns finish_reason: "length" (max_tokens with some providers), and that's the only sign of a truncation. Raising the cap to 2400 tokens didn't solve the problem: a response can still turn out longer, and every extra token increases the reserved amount. The server sends the client a truncated flag if finish_reason indicated the cap was hit. The client itself sends the model a service request to continue from the same spot, attaching the tail of what's already been written: the server trims messages to 6000 characters, so the last 5800 are used. Continuations are limited to four. The hardest part is the stitching: the first version simply appended the continuation, producing artifacts like "102. ai102. ai and neural networks" or "Best serviBest services." The result was a function with three rules: remove the repeated tail of the written text (with a word-boundary check for short matches up to 12 characters), replace the truncated line with the new one if the model started it over, and add a line break if the model started with the next list item, table, or heading. The function was tested on 13 cases from real truncations plus invented edge cases; the same function runs on the server so the response sits as a single entry in chat history. ### 💡 Why it matters Without such a mechanism, any long generation can end up useless: the user gets a fragment instead of a hundred-item list, and the chat history holds several chunks instead of one coherent answer. Checking finish_reason on every response is cheaper than dealing with user screenshots later. The mechanism is already in use in the service, and while truncations haven't disappeared, the user gets a complete answer and the log holds one tidy entry. ### 🧩 Context It all started with a screenshot from a partner: the model began answering with a hundred-item list and cut off mid-word. The admin panel showed the same fragment — no error, no warning. After that, the cause was traced to the max_tokens cap, which was raised to 2400, and continuation was added. Money had to be handled separately: each continuation is a separate paid request, the amount is reserved before the call, charged based on actual usage, and the reservation is refunded on error. Models with mandatory reasoning spend part of the cap on reasoning, so they hit the cutoff earlier than one might

🔗 Read on habr.com

🤖 AI summary
LLMfinish_reasonчат-бот
📖
Read the guide on this topic
Read →
← Previousmesshub — an open-source notification aggregator for Linux and WindowsNext →Top categories on Google Play (Sep 2026) - AppBrain

Source: habr.com · post in Telegram