fix(translation): reduce Anthropic token fallback - #282
Merged
Conversation
Signed-off-by: nachiketb <nachiketb@nvidia.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (4)
WalkthroughAnthropic request encoding now defaults omitted ChangesAnthropic request translation
Estimated code review effort: 1 (Trivial) | ~5 minutes Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
Comment |
elyasmnvidian
approved these changes
Aug 4, 2026
grahamking
approved these changes
Aug 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Reduce the implicit Anthropic
max_tokensfallback from 128,000 to 64,000 when an OpenAI-format request omits bothmax_tokensandmax_completion_tokens.Why
Omitting an output limit is valid for OpenAI clients, while Anthropic Messages requires
max_tokens. Switchyard therefore needs to supply a fallback during cross-format translation.The existing fallback used 128,000, which is the model's absolute output ceiling rather than a conservative default. Anthropic counts both adaptive reasoning and visible response tokens against this limit. Combined with
reasoning_effort=max, the implicit 128K budget can allow requests to run beyond the 600-second upstream deadline and return HTTP 500.64K matches Claude Code's default for recognized modern Opus models and still leaves substantial room for agentic and coding output without automatically granting every uncapped request the maximum possible budget.
Tracks SWITCH-1194.
How
What to review
Validation
Summary by CodeRabbit