🤖 fix: clamp GPT-6 Sol/Luna headless tool loops to reasoning none and correct their Codex context cap - #4259
Conversation
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Automations to automatically generate PRs for you. |
|
@codex review |
This comment has been minimized.
This comment has been minimized.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 01ca0537be
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
🛡️ Codex Security Review · Automatically triggered
Here are some automated security review suggestions for this pull request.
Reviewed commit: 01ca0537be
ℹ️ About Codex security reviews in GitHub
This is an experimental Codex feature. Security reviews are triggered when:
- You comment "@codex security review"
- A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review
Once complete, Codex will leave suggestions, or a comment if no findings are found.
|
@codex review |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cb5ac4b075
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
|
Review-gate bookkeeping: the Codex Review Summary card (node |
782cca6 to
d24d776
Compare
Launch UI evidence — Sol and LunaExact source:
Limits: this was an unsent scratch draft, not a persistent-workspace send. No live API, Codex or gateway call. Negative Pro-route gating is covered by targeted automated tests, not these UI recordings. A supplementary phone Sol-reload clip showed only its earlier state and was excluded; the phone Sol reload has screenshot/snapshot evidence instead. The two accepted recordings below were decoded and inspected. Desktop recordingdesktop.webmPhone recordingmobile.webmGenerated with |
Final handoff — blocked for merge, not a release-ready verdictBoth GPT-6 Sol and Luna are implemented at
The PR body preserves the full post-cap sequence, including the earlier Generated with |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 40e4b2a6b8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bfc0906fa6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a4c7c15635
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
🛡️ Codex Security Review · Automatically triggered
Here are some automated security review suggestions for this pull request.
Reviewed commit: a4c7c15635
ℹ️ About Codex security reviews in GitHub
This is an experimental Codex feature. Security reviews are triggered when:
- You comment "@codex security review"
- A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review
Once complete, Codex will leave suggestions, or a comment if no findings are found.
|
@codex review |
|
Codex Review: Didn't find any major issues. Swish! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
c9a0f57 to
76bff93
Compare
|
@codex review |
|
Codex Review: Didn't find any major issues. Nice work! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
## Summary Bumps `packages/mux-compat/package.json` (`version` and the `@coder/xum` dependency pin) from 0.29.0 to 0.30.0 to match the root package after the `release: v0.30.0` commit (`81b0b744`), which bumped only the root `package.json`. ## Background `src/common/compat/productIdentity.test.ts` › "keeps the published mux forwarding package version-locked to @coder/xum" asserts `legacyPackageJson.version === packageJson.version`. Since the release commit it fails on `main` ([run 35783501719](https://github.com/coder/xum/actions/runs/35783501719)) and on every PR's merge commit (e.g. coder#4259 [run 35785546631](https://github.com/coder/xum/actions/runs/35785546631), coder#4350), so `Required` is red repo-wide and the merge queue cannot admit anything. Same fix as coder#4048 for v0.28.3; release commits since (e.g. v0.28.5, coder#4175) bumped both files together. `@coder/xum@0.30.0` is already published (Publish to NPM run 35783501943), so the new pin resolves. ## Validation `bun test src/common/compat/productIdentity.test.ts` on this branch: 8 pass (the version-lock case fails on `main` without it). --- _Generated with `xum` • Model: `coder:anthropic/claude-fable-5-1` • Thinking: `xhigh` • Cost: `$174.60`_ <!-- mux-attribution: model=coder:anthropic/claude-fable-5-1 thinking=xhigh costs=174.60 -->
… correct their Codex context cap Stacked on #4350, which makes a requested reasoning_effort "none" reach the wire for GPT-6 Sol/Luna. Headless tool loops (Dream consolidation, memory harvest, refine, sidebar status) call streamText with tools and no provider options, so nothing requests "none" and Chat Completions defaults to medium, which rejects function calling. A model-boundary middleware fills in "none" for tool-bearing Chat Completions requests that asked for no effort, wrapped outside #4350's body-patching wrapper so the injected value survives the SDK; explicit caller efforts are left alone. The Coder openai-compat path resolves scoped "Treat as" aliases through resolveModelForMetadata before clamping. The assumed 372K Codex OAuth cap for Sol/Luna becomes the 272K published in the pinned Codex catalog. _Generated with `xum` • Model: `coder:anthropic/claude-fable-5-1` • Thinking: `xhigh` • Cost: `$174.60`_ <!-- mux-attribution: model=coder:anthropic/claude-fable-5-1 thinking=xhigh costs=174.60 -->
76bff93 to
bde7c0f
Compare
|
@codex review |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bde7c0fe03
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The Astra comment still claimed the catalog lists the same window for Astra and Sol, contradicting the 272K Sol/Luna entries below. State that Astra keeps its pre-existing cap (tracked in #4347) independently. _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `xhigh` • Cost: `$281.40`_ <!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=xhigh costs=281.40 -->
|
@codex review |
|
Codex Review: Didn't find any major issues. Swish! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
…cap (coder#4684) ## Summary Reconciles two OpenAI catalog entries with their first-party sources: GPT-5.6 Sol now prices at OpenAI's current promotional rates, and GPT-6 Astra's Codex OAuth context cap drops from 372K to the catalog's 272K. ## Background Both discrepancies were found while validating coder#4259 and tracked in coder#4347. **Pricing source.** [OpenAI API pricing](https://developers.openai.com/api/docs/pricing), fetched 2026-09-26. GPT-5.6 Sol (per 1M tokens): | | Input | Cached input | Cache writes | Output | |---|---|---|---|---| | Short context (≤272K) | $4.00 | $0.40 | $5.00 | $20.00 | | Long context (>272K) | $8.00 | $0.80 | $10.00 | $30.00 | The page says "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026." The previous entry ($5/$30/$0.50/$6.25; long context $10/$45/$1/$12.50) was the standard rate. The multipliers (2x input, 1.5x output, 1.25x cache writes, 272K boundary) are unchanged. **Catalog source.** [openai/codex@04fc75ad `models.json`](https://github.com/openai/codex/blob/04fc75adbe67a612a1cb0fc469533f24b24fa499/codex-rs/models-manager/models.json), the pin the repo already cites: `gpt-6-astra` has `context_window: 272000` and `max_context_window: 872000`. openai/codex `main` (checked 2026-09-26) has the same values. ## Implementation - `models-extra.ts`: `GPT_56_SOL_STATS` (shared by `gpt-5.6-sol` and the bare `gpt-5.6` alias) now uses the promo rates. - Time-limited pricing follows the most recent repo precedent (Gemini 3.7/3.8 Flash intro rates): encode the rate actually billed, with a dated `TODO(2026-11-21)` that records the standard rates and the source. No new mechanism. The older precedent (Claude Sonnet 5, DeepSeek V4 Pro) listed the standard rate instead. I picked the billed rate because the issue asks for it, the promotion has no fixed end date, and the pricing page no longer shows a standard rate for Sol. - `codexOAuth.ts`: `gpt-6-astra` 372K → 272K, with the pin cited next to Sol/Luna. Public API windows (1.05M) and user-mapped models are unchanged: the cap applies only when a request actually routes through Codex OAuth. ## Validation Tests were changed first and failed on `main`: - `displayUsage.test.ts` › "applies %s pricing only above the 272K boundary" now includes `openai:gpt-5.6-sol` and `openai:gpt-5.6`. It checks 272,000 vs 272,001 prompt tokens and now also checks cache writes (1.25x the active input rate), so the boundary is shown to move input, cache-read, cache-write and output rates together. - `codexOAuth.test.ts` (Astra override) and `contextLimit.test.ts` › "caps GPT-6 Astra on the OAuth route but keeps the API window for API-key auth" (272K on OAuth, 1.05M with API-key auth). - `modelStats.test.ts` base-rate row for Sol updated. Its long-context assertions are multipliers of the base rates, so they checked the new values unchanged. <details> <summary>Pre-fix failures</summary> ``` Expected: 272000 Received: 372000 (fail) codexOAuth model gating > allows gpt-6-astra through Codex OAuth with a 272000 context cap Expected: 0.668 Received: 0.8350000000000001 (fail) createDisplayUsage > tiered long-context pricing > applies openai:gpt-5.6-sol pricing only above the 272K boundary [1.00ms] (fail) createDisplayUsage > tiered long-context pricing > applies openai:gpt-5.6 pricing only above the 272K boundary (fail) getEffectiveContextLimit > caps GPT-6 Astra on the OAuth route but keeps the API window for API-key auth 60 pass 4 fail ``` </details> After the fix, `bun test src/common` passes (2380 tests) and `make static-check` passes. ## Risks Low. Sol cost estimates and goal budgets drop by 20–33% until the promotion ends. The dated TODO marks when to restore them. OAuth-routed Astra now starts limit-driven compaction at 272K instead of 372K. ## Deferred - The GPT-5.6 family Codex caps (372K since coder#3730) also differ from the catalog (272K). Out of scope here: tracked in coder#4683. Fixes coder#4347 --- _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high`_ <!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high -->
…log (coder#4710) ## Summary The Codex OAuth context cap for the GPT-5.6 family (`gpt-5.6`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`) goes from 372K to 272K, matching the first-party Codex catalog. API-key requests keep the public 1.05M window. ## Background coder#4683 asked where coder#3730's 372K came from before changing it. - coder#3730 (merged 2026-07-15) says it copied the `openai/codex` catalog, which "lists `context_window: 372000` for Sol, Terra, and Luna". That PR has no discussion of observed Codex behavior and no deliberate override. - Catalog history (`codex-rs/models-manager/models.json`): - [`c38c30ef`](https://github.com/openai/codex/blob/c38c30ef514dc74aa6ea59748f280c235b1bca2a/codex-rs/models-manager/models.json) (2026-07-14): Sol/Terra/Luna `context_window: 372000`, `max_context_window: 372000`. - [`2eee483e`](openai/codex@2eee483) (openai/codex#39102, "Raise the GPT-5.6 maximum context window", 2026-08-17): `context_window: 272000`, `max_context_window: 872000`. - The pin the repo already cites ([`04fc75ad`](https://github.com/openai/codex/blob/04fc75adbe67a612a1cb0fc469533f24b24fa499/codex-rs/models-manager/models.json)) and `main` (checked 2026-09-26) both still have 272000 / 872000. So 372K came only from an older catalog. This PR uses the default `context_window`, the same convention as GPT-5.5 and the GPT-6 Astra/Sol/Luna entries (coder#4259, coder#4684). ## Validation The routing and meter expectations were updated first and failed on `main`: `codexOAuth.test.ts` (GPT-5.6 family overrides), `contextLimit.test.ts` (OAuth-routed effective limit per tier), and `tokenMeterUtils.test.ts` (OAuth meter max and percentage for Sol). The API-key → 1.05M assertions in those files are unchanged and still pass. <details> <summary>Pre-fix failures</summary> ``` Expected: 272000 Received: 372000 (fail) codexOAuth model gating > allows the GPT-5.6 family through Codex OAuth without requiring it (fail) calculateTokenMeterData > uses the Codex OAuth cap for GPT-5.6 token meter percentages (fail) getEffectiveContextLimit > caps the GPT-5.6 family at each tier's Codex OAuth context window 36 pass 3 fail ``` </details> ## Risks Low. Only direct OpenAI requests routed through Codex OAuth are affected. They now start limit-driven compaction at 272K instead of 372K, which is the documented default window. API-key, gateway and Coder routes are unchanged. Fixes coder#4683 --- _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high`_ <!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high -->



Summary
Follow-up to #4350 (landed as
7358fa80), which makes a requestedreasoning_effort: "none"reach the wire for GPT-6 Sol/Luna. This PR closes the three gaps left after #4348 and #4350: headless tool loops that request no effort at all, Coder-scoped aliases onopenai-compatinstances, and the assumed 372K Codex OAuth context cap (272K in the pinned Codex catalog).Background
reasoning_effort: "none"; an omitted effort defaults tomedium.buildProviderOptionsclamps agent turns (🤖 feat: add GPT-6 Sol and Luna model support #4348) and 🤖 fix: send GPT-6 Sol/Luna reasoning effort none and repair main CI #4350 makes that clamp serialize, but Dream consolidation, memory harvest, refine and sidebar-status generation callstreamTextwith tools and no provider options, so nothing requestsnoneand their tool calls are rejected on Chat Completions routes (directwireFormat: chatCompletions, Coderopenai-compatinstances). Raised by Codex on the earlier revision of this PR; the Coder alias case is its round-6 finding.codex-rs/models-manager/models.json—gpt-6-sol/gpt-6-lunacontext_window: 272000.maincopies Astra's 372K with an "assumed" comment.Implementation
clampGpt6ChatCompletionsToolReasoning(providerModelFactory.ts): atransformParamsmiddleware that setsreasoningEffort: "none"when the request carries tools and no effort was requested, for Sol/Luna capability identities on Chat Completions. Explicit caller efforts are preserved (this keeps 🤖 fix: send GPT-6 Sol/Luna reasoning effort none and repair main CI #4350's "requestedhighstayshigh" contract intact). It wraps outsidecreateOpenAIModelWithPreservedOptionsso the injectednoneis seen by that wrapper's body patch instead of being stripped by the SDK.resolveModelForMetadata(coder:<instance>/<model>)so scoped "Treat as" aliases are clamped too.CODEX_OAUTH_CONTEXT_WINDOW_OVERRIDES:gpt-6-sol/gpt-6-luna372K → 272K with the pin cited inline; Astra and the GPT-5.6 family are unchanged (🤖 fix: reconcile GPT-5.6 Sol promotional pricing and Astra Codex context cap #4347).Dropped from earlier revisions: catalog/aliases/pricing/labels/migration/docs (landed in #4348), the
@ai-sdk/openaiBun patch (BunpatchedDependenciesnever reach npm installs of@coder/xum; #4350's runtime rewrite is the right fix), and thehasToolsplumbing throughbuildProviderOptions.Validation
Repo-pinned Bun 1.3.5, on
main+ this commit with a freshly reinstalled, unpatched@ai-sdk/openai(so a green body assertion proves the composition with #4350's wrapper).providerModelFactory.test.ts› "clamps %s tool requests … at the model boundary" (Sol raw id, mappedteam-luna, Astra control, Responses control) and › "clamps %s tool requests through a Coder openai-compat instance" (coder:chat-proxy/team-lunamapped,coder:chat-proxy/gpt-6-solraw,team-astracontrol). Red proof: with the clamp disabled, exactly the four Sol/Luna Chat Completions cases fail (Expected "none" / Received undefined); 🤖 fix: send GPT-6 Sol/Luna reasoning effort none and repair main CI #4350's 11 wire cases are unaffected. Reverting the Coder capability argument to the raworiginModelIdfails exactly the mappedteam-lunacase.codexOAuth.test.tsasserts Astra (372K) and Sol/Luna (272K) separately.contextLimit,tokenMeterUtils,thinking,turnRequestBuilder,providerOptionssuites pass.make static-checkgreen.Risks
toolsarray, and no requested effort. Responses, tool-free requests, explicit efforts, Astra, GPT-5.6 and non-OpenAI providers are untouched.Readiness and follow-ups
mainafter 🤖 fix: send GPT-6 Sol/Luna reasoning effort none and repair main CI #4350 landed (7358fa80). The residual diff is unchanged (+183/−21).mainunblock forTest / Unit(release commit leftpackages/mux-compatat 0.29.0): 🤖 fix: version-lock mux compat package to v0.30.0 #4354 standalone, also folded into 🤖 fix: send GPT-6 Sol/Luna reasoning effort none and repair main CI #4350; whichever becomes empty is closed.Preserved review and delivery record — fifteen Codex executions including this head, one code-advisor pass
01ca053252bba9/cb5ac4b, replied and resolvedcb5ac4b64169a3, replied and resolved64169a340e4b2a6ce20ecd0, replied and resolvedbfc0906f(supersedingce20ecd0)a4c7c156, replied and resolveda4c7c156@codex review(code + security)2c4803c0(residual onmainafter #4348)@codex review(code + security)c9a0f572(stacked on #4350; rebased ontomainwithout changes)@codex review(code + security)76bff932(rebased ontomainafter #4350 landed)@codex review(code + security)compaction1MRetrycontent-filter failure (#4398), fixed by #4401bde7c0fe(rebased ontomainafter #4401 landed; diff unchanged)@codex review(code + security)7deeea6e, replied and resolved7deeea6e(comment-only fix)@codex review(code + security)ca425b49at 20:50ZPost-cap work, retained in order:
f04892f→fcf5d0fbudget-comment accuracy;782cca6→04bad77maintainer-authorized provisional estimate;d24d776rebase integration;870259aofficial Sol launch + authorized Luna expansion;40e4b2a6verified Codex defaults, labels, route/pricing regressions, docs;ce20ecd0tool-awarenoneclamp + SDK patch;bfc0906fNix hash refresh;a4c7c156model-boundary clamp;2c4803c0residual onmain(SDK patch + clamp + Coder alias + Codex cap). Other rebase mappings:01ca053→c47ba9c,252bba9→5363ca0,cb5ac4b→a866d96,64169a3→c0d5def. The cap did not reset on rebase, handoff, scope expansion or re-scope. Backups of superseded heads:a4c7c15635f3a865baa4b937618da17459b48209,2c4803c0c9481924ad682f969542a7a6db71e9e7.Historical stop reason: implementation and validation were complete for both models, but the final head lacked review coverage under the exhausted cap. The later direct conditional merge instruction reopened delivery for one automatic final-head pair; its findings were then fixed under the user's explicit "get it merged" authorization through the normal gated merge path. When #4348 and then #4350 landed the same work first, this PR was reduced to the residual fixes above rather than closed.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:xhigh• Cost:$281.40