Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions docs-site/src/content/docs/guides/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -276,6 +276,16 @@ free-experimentation model.
| Cloudflare AI Gateway | `https://gateway.ai.cloudflare.com/v1/{account-id}/{gateway}/anthropic` |
| …and more | opencode zen, Vercel AI Gateway, Venice, NanoGPT, Synthetic, Qianfan, Alibaba, Parallel, ZenMux, LiteLLM |

**OpenCode Zen** (`opencode-zen`) and the keyless **OpenCode Free** preset share
`https://opencode.ai/zen/v1`. Free models on that gateway often hit a short-window burst
limit around 15–20 requests/minute (community-measured; OpenCode does not publish RPM).
Zen may return generic rate-limit 429 responses without `Retry-After` / `X-RateLimit-*`
headers. That is separate from the keyless desktop quota OpenCode advertises
(~200 Big Pickle/free-model requests per 5 hours on `opencode-free`). When Zen omits
`Retry-After` on such a 429, opencodex adds provider guidance to the client error and a
synthetic `Retry-After`; an upstream `Retry-After` still takes precedence. Same-key
wait-and-retry remains opt-in via [`retryOn429`](/reference/configuration/).

Comment thread
coderabbitai[bot] marked this conversation as resolved.
Most use the `openai-chat` adapter with a bearer key; a few that expose only an Anthropic-compatible
endpoint (e.g. **Xiaomi MiMo**) use the `anthropic` adapter (`x-api-key`).
Volcengine Agent Plan uses its native Responses endpoint through `openai-responses`.
Expand Down
3 changes: 3 additions & 0 deletions docs-site/src/content/docs/ja/guides/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -203,6 +203,9 @@ Cline IDE/CLI のみで API からは使えません。`minimax/minimax-m2.5`
| Cloudflare AI Gateway | `https://gateway.ai.cloudflare.com/v1/{account-id}/{gateway}/anthropic` |
| …その他多数 | opencode zen、Vercel AI Gateway、Venice、NanoGPT、Synthetic、Qianfan、Alibaba、Parallel、ZenMux、LiteLLM |

**OpenCode Zen**(`opencode-zen`)とキー不要の **OpenCode Free** プリセットは
`https://opencode.ai/zen/v1` を共有します。このゲートウェイ上の無料モデルは、しばしばおおよそ毎分 15–20 リクエストの短時間レート制限に当たります(コミュニティ計測。OpenCode は RPM を公表しません)。Zen は `Retry-After` / `X-RateLimit-*` ヘッダーなしの汎用 429 を返すことがあります。これはキー不要デスクトップ枠(`opencode-free` で Big Pickle/無料モデル約 200 回 / 5 時間)とは別です。Zen がそのような 429 で `Retry-After` を省略した場合、opencodex はクライアント向けエラーに案内を足し、合成 `Retry-After` を付けます(上流の `Retry-After` があればそれが優先されます)。同一キーの待機再試行は [`retryOn429`](/ja/reference/configuration/) でオプトインします。

大半は bearer キーと共に `openai-chat` アダプターを使い、Anthropic 互換エンドポイントのみを公開する一部
(例: **Xiaomi MiMo**)は `anthropic` アダプター(`x-api-key`)を使います。
Volcengine Agent Plan は `openai-responses` アダプターでネイティブ Responses エンドポイントを使用します。
Expand Down
3 changes: 3 additions & 0 deletions docs-site/src/content/docs/ko/guides/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -203,6 +203,9 @@ Cline IDE/CLI에서만 제공되며 API로는 사용할 수 없습니다. `minim
| Cloudflare AI Gateway | `https://gateway.ai.cloudflare.com/v1/{account-id}/{gateway}/anthropic` |
| …그 외 다수 | opencode zen, Vercel AI Gateway, Venice, NanoGPT, Synthetic, Qianfan, Alibaba, Parallel, ZenMux, LiteLLM |

**OpenCode Zen**(`opencode-zen`)과 키 없는 **OpenCode Free** 프리셋은
`https://opencode.ai/zen/v1`을 공유합니다. 그 게이트웨이의 무료 모델은 종종 분당 약 15–20회 요청의 짧은 창 속도 제한에 걸립니다(커뮤니티 측정; OpenCode는 RPM을 공개하지 않음). Zen은 `Retry-After` / `X-RateLimit-*` 헤더 없는 일반 429를 반환할 수 있습니다. 이는 키 없는 데스크톱 할당량(`opencode-free`에서 약 5시간당 Big Pickle/무료 모델 200회)과 별개입니다. Zen이 그런 429에서 `Retry-After`를 생략하면 opencodex는 클라이언트 오류에 안내를 더하고 합성 `Retry-After`를 붙입니다(업스트림 `Retry-After`가 있으면 그것이 우선). 동일 키 대기 재시도는 [`retryOn429`](/ko/reference/configuration/)로 선택합니다.

대부분은 bearer 키와 함께 `openai-chat` 어댑터를 사용하며, Anthropic 호환 엔드포인트만 노출하는 일부
(예: **Xiaomi MiMo**)는 `anthropic` 어댑터(`x-api-key`)를 사용합니다.
Volcengine Agent Plan은 `openai-responses` 어댑터로 네이티브 Responses 엔드포인트를 사용합니다.
Expand Down
9 changes: 9 additions & 0 deletions docs-site/src/content/docs/ru/guides/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -214,6 +214,15 @@ opencodex поставляется с 76 встроенными пресетам
| Cloudflare AI Gateway | `https://gateway.ai.cloudflare.com/v1/{account-id}/{gateway}/anthropic` |
| …и другие | opencode zen, Vercel AI Gateway, Venice, NanoGPT, Synthetic, Qianfan, Alibaba, Parallel, ZenMux, LiteLLM |

**OpenCode Zen** (`opencode-zen`) и бесключевой пресет **OpenCode Free** используют один
`https://opencode.ai/zen/v1`. Бесплатные модели на этом шлюзе часто упираются в короткое окно
примерно 15–20 запросов в минуту (оценка сообщества; OpenCode не публикует RPM).
Zen может отвечать общими 429 без заголовков `Retry-After` / `X-RateLimit-*`. Это отдельно от
бесключевой десктопной квоты (~200 запросов Big Pickle/бесплатных моделей за 5 часов на
`opencode-free`). Когда Zen опускает `Retry-After` на таком 429, opencodex добавляет пояснение
в ошибку клиента и синтетический `Retry-After`; при наличии upstream `Retry-After` он имеет
приоритет. Повтор с тем же ключом по-прежнему включается через [`retryOn429`](/ru/reference/configuration/).

Большинство использует адаптер `openai-chat` с bearer-ключом; немногие провайдеры, предоставляющие
только Anthropic-совместимую конечную точку (например, **Xiaomi MiMo**), используют адаптер
`anthropic` (`x-api-key`).
Expand Down
3 changes: 3 additions & 0 deletions docs-site/src/content/docs/zh-cn/guides/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -191,6 +191,9 @@ Cline IDE/CLI 中提供,不能通过 API 使用;`minimax/minimax-m2.5` 是
| Cloudflare AI Gateway | `https://gateway.ai.cloudflare.com/v1/{account-id}/{gateway}/anthropic` |
| ……以及更多 | opencode zen、Vercel AI Gateway、Venice、NanoGPT、Synthetic、Qianfan、Alibaba、Parallel、ZenMux、LiteLLM |

**OpenCode Zen**(`opencode-zen`)与免密钥的 **OpenCode Free** 预设共用
`https://opencode.ai/zen/v1`。该网关上的免费模型常会触发约每分钟 15–20 次请求的短窗口限流(社区观测;OpenCode 未公布 RPM)。Zen 可能返回不带 `Retry-After` / `X-RateLimit-*` 的通用 429。这与免密钥桌面配额(`opencode-free` 上约每 5 小时 200 次 Big Pickle/免费模型请求)是分开的。当这类 429 省略 `Retry-After` 时,opencodex 会在客户端错误中补充说明并附带合成的 `Retry-After`;若上游已提供 `Retry-After`,则仍以它为准。同密钥等待重试仍可通过 [`retryOn429`](/zh-cn/reference/configuration/) 选择开启。

大多数使用带 bearer 密钥的 `openai-chat` adapter;少数仅暴露 Anthropic 兼容端点的提供商(例如 **Xiaomi MiMo**)使用 `anthropic` adapter(`x-api-key`)。
火山方舟 Agent Plan 通过 `openai-responses` adapter 使用原生 Responses 端点。

Expand Down
102 changes: 102 additions & 0 deletions src/providers/opencode-zen-rate-limit.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,102 @@
/**
* OpenCode Zen short-window rate-limit guidance (#1145 / OCX-56).
*
* OpenCode's keyed and keyless Zen chat endpoints share `https://opencode.ai/zen/v1`.
* Free-model traffic can hit a short-window burst ceiling around 15–20 RPM
* (community-measured). Zen often answers with opaque `429 Rate limit exceeded`
* bodies and may omit `Retry-After` / `X-RateLimit-*`; when those headers are
* present they still take precedence. Distinct from the keyless desktop
* ~200 requests / 5h quota documented on `opencode-free`.
*/
import { validateClientRetryAfterHeader } from "../lib/retry-after";
import { registryEntryForProviderDestination } from "./registry";

const OPENCODE_ZEN_PROVIDER_IDS = new Set(["opencode-zen", "opencode-free"]);

/** Observed free-model burst ceiling on Zen (not an official OpenCode figure). */
export const OPENCODE_ZEN_OBSERVED_RPM_HINT = "roughly 15-20 requests per minute";

/**
* Synthetic client backoff when Zen omits Retry-After after a rate-limit 429.
* Longer than the generic 2s default so Codex-shaped clients do not immediately
* re-hammer a ~15-20 RPM window.
*/
export const OPENCODE_ZEN_SYNTHETIC_RETRY_AFTER_SEC = 15;

const ENRICHMENT_MARKER = "15-20 requests per minute";

export function isOpenCodeZenRateLimitProvider(opts: {
providerName?: string;
baseUrl?: string;
adapter?: string;
}): boolean {
const name = opts.providerName?.trim();
if (name && OPENCODE_ZEN_PROVIDER_IDS.has(name)) return true;
const baseUrl = opts.baseUrl?.trim();
if (!baseUrl) return false;
const entry = registryEntryForProviderDestination({
baseUrl,
adapter: opts.adapter?.trim() || "openai-chat",
authMode: "key",
});
return entry !== undefined && OPENCODE_ZEN_PROVIDER_IDS.has(entry.id);
}

/**
* Same-key `retryOn429` only applies on key-authenticated HTTP paths — not
* keyless `opencode-free` traffic and not custom `runTurn` transports.
*/
export function supportsOpenCodeZenRetryOn429Guidance(opts: {
authMode?: string;
hasApiKey?: boolean;
supportsHttpSameKeyRetry?: boolean;
}): boolean {
if (opts.supportsHttpSameKeyRetry === false) return false;
if (opts.authMode !== undefined && opts.authMode !== "key") return false;
return opts.hasApiKey === true;
}

/**
* Append actionable Zen rate-limit context to a generic upstream 429 message and
* embed a parseable `try again in Ns` hint so {@link resolveClientRetryAfter}
* surfaces a useful Retry-After when the gateway sent none.
*/
export function enrichOpenCodeZenRateLimitMessage(
message: string,
opts: {
status: number;
providerName?: string;
baseUrl?: string;
adapter?: string;
authMode?: string;
hasApiKey?: boolean;
/** Upstream Retry-After header; when valid, skip the synthetic 15s text hint. */
upstreamRetryAfter?: string | null;
/** False for custom `runTurn` transports outside the HTTP retry loop. */
supportsHttpSameKeyRetry?: boolean;
now?: number;
},
): string {
if (opts.status !== 429) return message;
if (!isOpenCodeZenRateLimitProvider(opts)) return message;
if (!/rate\s*limit/i.test(message)) return message;
if (message.includes(ENRICHMENT_MARKER)) return message;

const upstreamRetry = validateClientRetryAfterHeader(
opts.upstreamRetryAfter,
opts.now ?? Date.now(),
);
const retryHint = upstreamRetry || /try again in \d/i.test(message)
? ""
: ` Try again in ${OPENCODE_ZEN_SYNTHETIC_RETRY_AFTER_SEC}s.`;
const paceHint = supportsOpenCodeZenRetryOn429Guidance(opts)
? " Slow the request pace, or set providers.opencode-zen.retryOn429 for same-key backoff."
: " Slow the request pace.";
return (
`${message}`
+ ` OpenCode Zen free-model traffic is often limited to ${OPENCODE_ZEN_OBSERVED_RPM_HINT}`
+ ` (observed; OpenCode does not publish this RPM, and may omit rate-limit headers).`
+ `${retryHint}`
+ paceHint
);
}
3 changes: 2 additions & 1 deletion src/providers/registry.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2037,6 +2037,7 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [
// continuations, or the gateway answers HTTP 400 (issues #950/#994). Mirror the DeepSeek
// reasoning + thinking metadata so `opencode-zen/deepseek-v4-flash-free` — and the other
// Zen DeepSeek thinking models — never serialize a bare tool-call turn.
note: "Keyed OpenCode Zen gateway. Free models on this tier are often short-window rate-limited at roughly 15-20 requests/minute (community-measured; OpenCode does not publish RPM). Zen may return generic 429s without Retry-After / X-RateLimit headers; when Retry-After is omitted, opencodex adds a synthetic backoff hint (upstream Retry-After still wins). Distinct from the keyless opencode-free desktop quota (~200 Big Pickle/free-model requests per 5 hours). Docs: https://opencode.ai/docs/zen/. Free-model prompts may be retained for training — do not send confidential material.",
modelReasoningEfforts: Object.fromEntries(
[...DEEPSEEK_THINKING_MODELS, ...OPENCODE_FREE_DEEPSEEK_MODELS].map(id => [id, deepseekThinkingEffortsFor(id)]),
),
Expand All @@ -2056,7 +2057,7 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [
keyOptional: true,
featured: true,
liveModels: true,
note: "No key needed — public desktop tier. OpenCode currently advertises about 200 Big Pickle/free-model requests per 5 hours. Free models are discovered live from Zen. Data use: per OpenCode's Zen docs (https://opencode.ai/docs/zen/), prompts sent to free models may be retained and used for training/improvement — do not send confidential material through this provider.",
note: "No key needed — public desktop tier. OpenCode currently advertises about 200 Big Pickle/free-model requests per 5 hours. The same Zen gateway can also short-window rate-limit free models at roughly 15-20 requests/minute, and may return generic 429s without Retry-After (opencodex synthesizes backoff only when that header is omitted). Free models are discovered live from Zen. Data use: per OpenCode's Zen docs (https://opencode.ai/docs/zen/), prompts sent to free models may be retained and used for training/improvement — do not send confidential material through this provider.",
dashboardUrl: "https://opencode.ai",
staticHeaders: {
"x-opencode-client": "desktop",
Expand Down
20 changes: 18 additions & 2 deletions src/server/responses/core.ts
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@ import {
import { isInjectionDebugEnabled } from "../../lib/debug-settings";
import { injectionDebugLog } from "../../lib/injection-debug-log";
import { resolveClientRetryAfter } from "../../lib/retry-after";
import { enrichOpenCodeZenRateLimitMessage } from "../../providers/opencode-zen-rate-limit";
import { modelInList, namespacedToolName } from "../../types";
import type { AdapterEvent, OcxConfig, OcxParsedRequest, OcxProviderConfig, OcxProviderContinuationState, OcxUsage } from "../../types";
import {
Expand Down Expand Up @@ -3154,11 +3155,26 @@ async function handleResponsesInner(
}
// Upstreams occasionally echo request details in error bodies — scrub token-shaped
// material before it reaches the client-facing error surface.
const message = `Provider error ${upstreamResponse.status}: ${redactSecretString(errorText.slice(0, 500))}`;
const upstreamRetryAfter = upstreamResponse.headers.get("retry-after");
const message = enrichOpenCodeZenRateLimitMessage(
`Provider error ${upstreamResponse.status}: ${redactSecretString(errorText.slice(0, 500))}`,
{
status: upstreamResponse.status,
providerName: route.providerName,
baseUrl: route.provider.baseUrl,
adapter: route.provider.adapter,
authMode: route.provider.authMode,
hasApiKey: Boolean(route.provider.apiKey?.trim()),
upstreamRetryAfter,
// This recovery path is the HTTP Responses wire; custom runTurn transports
// never reach enrichOpenCodeZenRateLimitMessage here.
supportsHttpSameKeyRetry: true,
},
);
const retryAfter = resolveClientRetryAfter({
status: upstreamResponse.status,
message,
upstreamRetryAfter: upstreamResponse.headers.get("retry-after"),
upstreamRetryAfter,
});
return formatErrorResponse(upstreamResponse.status, "upstream_error", message, {
...(retryAfter !== undefined ? { retryAfter } : {}),
Expand Down
Loading
Loading