RU

Local 27B model vs Claude: investigating a production incident

Published: 2026-10-06 · Author: AI Release · @ai_release1
Local 27B model vs Claude: investigating a production incident

⚡ The gist in 5 seconds - A local 27B model handled the investigation of a real incident, but it took about 40 runs. - The experiment is reproducible on a rig with two RTX 5090s and a real bug. - Limitation: server-side WebSearch in Claude Code was unavailable to local models; regular internet access was only via curl/WebFetch. ### 🔍 What was found The production incident unfolded in a Java service that calls an external service via Feign on top of Apache HttpClient 5. When the external service degraded, the leased metric maxed out at 200 out of 200, even though there were only about a dozen actually active calls. New requests failed with timeouts, and a restart temporarily restored everything to normal. The cause turned out to be a rare race condition in httpcore5 5.3.6: a thread waiting for a connection from the pool did not cancel the operation after a timeout, and the issued connection remained in the leased state with no owner. The bug was fixed in httpcore5 5.4.3. The situation was provoked by configuration: the limiter allowed 300 concurrent operations with a pool of 200 connections. A single such event is almost unnoticeable, but then positive feedback kicks in: the effective pool size shrinks, the queue grows, timeouts multiply, the probability of hitting the race again increases — and gradually the entire pool can be lost. The reference run by Claude Opus 5.5 completed all five investigation stages in about 20 minutes. The task was then run through a local 27B model on a rig: 2× RTX 5090, Ryzen 9 9950X3D, 128 GB RAM, Linux, NVMe. Conditions were identical: a clean clone of the repository, a fresh Maven repository without new versions of httpcore5 and httpclient5, the same prompt, no memory or history. The local models were connected to Claude Code via ANTHROPIC_BASE_URL and a small proxy. The agents could read code, write tests, and run Maven, but were not allowed to modify production code. As it later turned out, not everyone understood the restriction. The five stages included finding inconsistent limits, testing the logger hypothesis, identifying the race condition, reproducing the leak (leased > 0 after requests completed), and finding the fix upstream. ### 💡 Why it matters Running locally costs nothing, but you pay with time and number of attempts: about 40 runs. Anyone with two RTX 5090s and a real bug can repeat the test. An additional filter was a deliberately incorrect hypothesis about a custom Feign logger left in the prompt: with logLevel = NONE the logger is never called, and the model was supposed to reject this version. It was precisely this hypothesis that became one of the most interesting filters of the experiment. Also, server-side WebSearch in Claude Code was unavailable to local models, making the search for the upstream fix somewhat harder. ### 🧩 Context The experiment grew out of a real incident analysis: from Grafana charts, the author suspected a connection leak, asked Claude to compose a prompt for another agent, and got a full investigation. Then came the desire to check whether a local model could cope. One version of the motivation — a wish to put two RTX 5090s to work and warm up the room [text cut off]

🔗 Read on habr.com

🤖 AI summary
ИИJavaClaude
📖
Read the guide on this topic
Read →
← PreviousELMA365: three iterations to a useful AI for testingNext →ChatGPT signs fake New Yorker cartoons with real artists' names

Source: habr.com · post in Telegram