Claude Code via an API Gateway: Requirements for the Agent Loop
Published: 2026-10-07 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - Claude Code can work not only with Anthropic: via the ANTHROPIC_BASE_URL variable, requests can be routed to a third-party Anthropic-compatible API. - Using its own Caila gateway as an example, the Just AI team broke down the requirements for endpoints, streaming, and automatic model discovery. - Compatibility with the Anthropic API doesn't guarantee that any model behind the gateway will correctly handle the agent loop. ### 🔍 What was found In an article on Habr, the Just AI team explained how to connect Claude Code to an API gateway. Claude Code calls three endpoints: POST /v1/messages — required, nothing works without it; POST /v1/messages/count_tokens — optional, without it /context shows a character-based estimate instead of an exact count; GET /v1/models?limit=1000 — optional, without it /model shows only the built-in list. In Caila, the token-counting endpoint is implemented, so /context through the gateway shows exact numbers. Streaming must not be buffered: Claude Code reads the response as it arrives, and if the gateway assembles the entire response first, the agent will appear frozen. The client tracks SSE events and terminates the stream after a long wait, so the gateway must pass through SSE pings during lengthy generation. Automatic model discovery is enabled via CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1. There are two limitations: the request must complete within a three-second timeout, and redirects are treated as an error — the client doesn't pass authorization credentials to another host. Additionally, Claude Code filters the retrieved list of models. ### 💡 Why it matters A gateway handles access, keys, limits, and spending. Provider keys stay behind the gateway, while the developer is issued a separate platform key. You can cap the budget before the first run and see spending broken down by keys, models, and requests. This is especially relevant for users in Russia: paying Anthropic directly requires a foreign card and a VPN, and as of June 27, 2026, OpenRouter no longer serves Russian accounts. The client-side part of the setup also applies to other proxies — for example, a self-hosted LiteLLM instance or a corporate API gateway. ### 🧩 Context In a year and a half, Claude Code has gone from a niche hobby to one of the main tools in AI-assisted development. According to the JetBrains Developer Ecosystem Survey 2026 (over 15,000 developers surveyed, May–July 2026), 90% of respondents use AI coding agents at least once a week, and 68% use them daily. 39% use Claude Code, and 31% call it their primary tool. Anthropic analyzed roughly 400,000 Claude Code sessions over six months: the second-largest user group is business and finance, and among non-technical professions, the share of managers, salespeople, and lawyers is growing fastest.
⚡ The gist in 5 seconds - Claude Code can work not only with Anthropic: via the ANTHROPIC_BASE_URL variable, requests can be routed to a third-party Anthropic-compatible API.
- Using its own Caila gateway as an example, the Just AI team broke down the requirements for endpoints, streaming, and automatic model discovery.
- Compatibility with the Anthropic API doesn't guarantee that any model behind the gateway will correctly handle the agent loop.
🔍 What was found In an article on Habr, the Just AI team explained how to connect Claude Code to an API gateway.
Claude Code calls three endpoints: POST /v1/messages — required, nothing works without it; POST /v1/messages/count_tokens — optional, without it /context shows a character-based estimate instead of an exact count; GET /v1/models?limit=1000 — optional, without it /model shows only the built-in list.
In Caila, the token-counting endpoint is implemented, so /context through the gateway shows exact numbers.
Streaming must not be buffered: Claude Code reads the response as it arrives, and if the gateway assembles the entire response first, the agent will appear frozen.
The client tracks SSE events and terminates the stream after a long wait, so the gateway must pass through SSE pings during lengthy generation.