RU

YandexGPT got stuck in a loop for 37 minutes and burned nearly 300,000 tokens — without a single useful answer

Published: 2026-09-28 · Author: AI Release · @ai_release1
YandexGPT got stuck in a loop for 37 minutes and burned nearly 300,000 tokens — without a single useful answer

We test models not on pretty demos, but in real agentic scenarios. YandexGPT was next in line. Instead of a normal result, we got 37 minutes of an endless loop and a count approaching 300,000 tokens. ## Key facts - The task: the agent was supposed to check a website's indexing (open a page → look at it → return a conclusion). - The model got stuck in a loop: tool call → context rebuild → another call. - 295 consecutive repeated tool calls with no progress. - 296 context rebuilds (compaction). - 37 minutes of pure looping. - ~155,000 input + ~142,000 output tokens burned. - Result: zero useful output. We stopped it manually. - Across the entire series of tests with YandexGPT: 225k input + 161k output + 5.6 million cached-reading tokens — and not a single task was completed. ## What happened We connected YandexGPT as a regular agent via the OpenAI-compatible API. The task was simple and routine. The model started calling tools — and didn't stop. Each time it re-read the entire context, made another call, and re-read again. It couldn't get out of this state on its own. Without a manual abort, it would have kept spinning until it hit account limits. ## Why this is expensive In agentic work, billing is per step. Even if the model generates a couple of words and calls the tool again, you're charged for: - input tokens (the re-read context), - cache-reading tokens. In our case, each loop iteration cost about 15,000 tokens of context reading. Over 37 minutes, nearly 300,000 tokens accumulated — all of it wasted. ## Context: what models actually cost The topic of "price per task, not per million tokens" is now openly discussed in 2026. Habr has reviews where the price for the same result varies across LLMs by tens of times. In open comparisons of Russian and open models (for example, kuk/simple-evals-ru), the picture is: - yandexgpt-5-pro / yandexgpt-4-pro — about $12 per million tokens - DeepSeek V3 — about $0.80 - Qwen 2.5-72B — about $0.29 Meanwhile, on a number of benchmarks, open models deliver comparable quality. A separate topic is agent "tokenomics": how much a model burns on the same job. The spread between models in the same price category can reach 2.6x. Our case is extreme but telling: the model had no internal stop mechanism for useless repetitions. ## What to do about it - Set a hard step limit for agents (after N tool calls — stop). - Watch for billing anomalies: a sharp rise in costs without progress almost always means a loop. - Don't run expensive agentic scenarios without a cost ceiling. - Look not only at the price per million tokens, but also at how much the model actually burns per task. - Cache tokens are the most invisible "devourer" with constant context re-reading. We left Yande (the original text was cut off at this point.)

🤖 AI summary
#YandexGPT#LLM#агенты#токены#биллинг#сравнение#хабр#github
📖
Read the guide on this topic
Read →
← PreviousAnthropic buried prompt injections, but Claude Code Auto Mode breaks throughNext →solvi: the model proposes, code decides — verifiable business decisions in Python

Source: yandex.cloud · post in Telegram