Skip to content

🤖 fix: clamp GPT-6 Sol/Luna headless tool loops to reasoning none and correct their Codex context cap - #4259

Merged
ThomasK33 merged 2 commits into
mainfrom
gpt-6-sol-prep
Sep 23, 2026
Merged

ThomasK33 merged 2 commits into
mainfrom
gpt-6-sol-prep

Conversation

@ThomasK33

@ThomasK33 ThomasK33 commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

Summary

Follow-up to #4350 (landed as 7358fa80), which makes a requested reasoning_effort: "none" reach the wire for GPT-6 Sol/Luna. This PR closes the three gaps left after #4348 and #4350: headless tool loops that request no effort at all, Coder-scoped aliases on openai-compat instances, and the assumed 372K Codex OAuth context cap (272K in the pinned Codex catalog).

Background

Implementation

  • clampGpt6ChatCompletionsToolReasoning (providerModelFactory.ts): a transformParams middleware that sets reasoningEffort: "none" when the request carries tools and no effort was requested, for Sol/Luna capability identities on Chat Completions. Explicit caller efforts are preserved (this keeps 🤖 fix: send GPT-6 Sol/Luna reasoning effort none and repair main CI #4350's "requested high stays high" contract intact). It wraps outside createOpenAIModelWithPreservedOptions so the injected none is seen by that wrapper's body patch instead of being stripped by the SDK.
  • Applied on the direct OpenAI path (Chat Completions wire only) and the Coder gateway Chat path, where the capability identity is resolveModelForMetadata(coder:<instance>/<model>) so scoped "Treat as" aliases are clamped too.
  • CODEX_OAUTH_CONTEXT_WINDOW_OVERRIDES: gpt-6-sol / gpt-6-luna 372K → 272K with the pin cited inline; Astra and the GPT-5.6 family are unchanged (🤖 fix: reconcile GPT-5.6 Sol promotional pricing and Astra Codex context cap #4347).

Dropped from earlier revisions: catalog/aliases/pricing/labels/migration/docs (landed in #4348), the @ai-sdk/openai Bun patch (Bun patchedDependencies never reach npm installs of @coder/xum; #4350's runtime rewrite is the right fix), and the hasTools plumbing through buildProviderOptions.

Validation

Repo-pinned Bun 1.3.5, on main + this commit with a freshly reinstalled, unpatched @ai-sdk/openai (so a green body assertion proves the composition with #4350's wrapper).

  • providerModelFactory.test.ts › "clamps %s tool requests … at the model boundary" (Sol raw id, mapped team-luna, Astra control, Responses control) and › "clamps %s tool requests through a Coder openai-compat instance" (coder:chat-proxy/team-luna mapped, coder:chat-proxy/gpt-6-sol raw, team-astra control). Red proof: with the clamp disabled, exactly the four Sol/Luna Chat Completions cases fail (Expected "none" / Received undefined); 🤖 fix: send GPT-6 Sol/Luna reasoning effort none and repair main CI #4350's 11 wire cases are unaffected. Reverting the Coder capability argument to the raw originModelId fails exactly the mapped team-luna case.
  • codexOAuth.test.ts asserts Astra (372K) and Sol/Luna (272K) separately. contextLimit, tokenMeterUtils, thinking, turnRequestBuilder, providerOptions suites pass. make static-check green.

Risks

  • The middleware fires only for Sol/Luna capability identities, Chat Completions wire, a non-empty tools array, and no requested effort. Responses, tool-free requests, explicit efforts, Astra, GPT-5.6 and non-OpenAI providers are untouched.
  • Lower Codex OAuth cap (272K) starts limit-driven compaction earlier for OAuth-routed Sol/Luna; API-key routes keep the public limits.

Readiness and follow-ups

Preserved review and delivery record — fifteen Codex executions including this head, one code-advisor pass
Round Reviewed head (before rebase / re-scope) Executions Findings / disposition
1 01ca053 Codex code + security Hidden-model compaction exposure and zero pricing bypassing CLI budget; fixed in 252bba9 / cb5ac4b, replied and resolved
2 cb5ac4b Codex code + security Migration-marker flip and hiding prior opt-ins; fixed in 64169a3, replied and resolved
3 64169a3 Codex code + security Clean completion on both loops
4, renewed authorization 40e4b2a6 Automatic Codex code + security on undraft Security clean; code P2 Chat Completions default-effort issue → fixed in ce20ecd0, replied and resolved
5, explicit "get it merged" authorization bfc0906f (superseding ce20ecd0) Automatic Codex code review on undraft P2: headless tool loops bypass provider options → fixed in a4c7c156, replied and resolved
6 a4c7c156 Explicit @codex review (code + security) Code P2: Coder aliases not resolved before clamping → fixed, replied and resolved. Security P2: priority-tier budget accounting → pre-existing/generic, #4352, replied and resolved
7 2c4803c0 (residual on main after #4348) Explicit @codex review (code + security) Both clean ("Didn't find any major issues" / "No security issues")
8 c9a0f572 (stacked on #4350; rebased onto main without changes) Explicit @codex review (code + security) Both clean ("Didn't find any major issues" 21:57Z / "No security issues" 21:59Z, 👍)
9 76bff932 (rebased onto main after #4350 landed) Explicit @codex review (code + security) Both clean; merge blocked by the repo-wide compaction1MRetry content-filter failure (#4398), fixed by #4401
10 bde7c0fe (rebased onto main after #4401 landed; diff unchanged) Explicit @codex review (code + security) Security clean. Code P3: stale Astra comment claimed the catalog couples Astra and Sol → fixed in 7deeea6e, replied and resolved
11 7deeea6e (comment-only fix) Explicit @codex review (code + security) Both clean ("Didn't find any major issues" 20:13Z / "No security issues" 20:17Z, 👍). Merged via queue as ca425b49 at 20:50Z
Advisory — One independent advisor Recommended scope reduction and accuracy fixes before official release. Scope reduction realised by #4348 and #4350 landing first; this PR is the residual.

Post-cap work, retained in order: f04892f→fcf5d0f budget-comment accuracy; 782cca6→04bad77 maintainer-authorized provisional estimate; d24d776 rebase integration; 870259a official Sol launch + authorized Luna expansion; 40e4b2a6 verified Codex defaults, labels, route/pricing regressions, docs; ce20ecd0 tool-aware none clamp + SDK patch; bfc0906f Nix hash refresh; a4c7c156 model-boundary clamp; 2c4803c0 residual on main (SDK patch + clamp + Coder alias + Codex cap). Other rebase mappings: 01ca053→c47ba9c, 252bba9→5363ca0, cb5ac4b→a866d96, 64169a3→c0d5def. The cap did not reset on rebase, handoff, scope expansion or re-scope. Backups of superseded heads: a4c7c15635f3a865baa4b937618da17459b48209, 2c4803c0c9481924ad682f969542a7a6db71e9e7.

Historical stop reason: implementation and validation were complete for both models, but the final head lacked review coverage under the exhausted cap. The later direct conditional merge instruction reopened delivery for one automatic final-head pair; its findings were then fixed under the user's explicit "get it merged" authorization through the normal gated merge path. When #4348 and then #4350 landed the same work first, this PR was reduced to the residual fixes above rather than closed.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: xhigh • Cost: $281.40

@mintlify

mintlify Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
Mux 🟢 Ready View Preview Sep 23, 2026, 8:01 PM

💡 Tip: Enable Automations to automatically generate PRs for you.

@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

This comment has been minimized.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 01ca0537be

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/common/constants/knownModels.ts Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛡️ Codex Security Review · Automatically triggered

Here are some automated security review suggestions for this pull request.

Reviewed commit: 01ca0537be

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Comment thread src/common/utils/tokens/models-extra.ts Outdated
@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: cb5ac4b075

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cb5ac4b075

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/config/index.ts Outdated
Comment thread src/node/config/index.ts
@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 64169a3fa5

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@ThomasK33

Copy link
Copy Markdown
Member Author

Review-gate bookkeeping: the Codex Review Summary card (node IC_kwDOPxxmWM8AAAABUnaNuQ, body sha256[0:16] d1158341d515a122) was minimized as RESOLVED after verifying: all 4 findings across rounds 1–2 are fixed with inline replies, all review threads are resolved, and both the Code Review and Security Review rows show Completed on the current head 64169a3 with no new findings. The Codex Comments CI job counts the mutable summary card as unresolved even after settlement (known checker limitation); minimizing only that card is its supported resolution. This will be re-assessed if the head or review state changes.

@ThomasK33

ThomasK33 commented Sep 22, 2026 •

Copy link
Copy Markdown
Member Author

Launch UI evidence — Sol and Luna

Exact source: 40e4b2a6b8d116d75fcda76b156ea9d5a5a0520c. Functional UI checks only, not another code/security review or provider-availability test.

  1. At 1440×1000 and 390×844, both GPT-6 models are visible in the normal selector. Selections persist after reload. GPT-5.6 Sol and Luna remain selectable.
  2. Phone controls distinguish Sol 6 / Luna 6 from Sol / Luna. Both new models show all six reasoning levels and Pro after explicitly lowering the fixture's minimum effort to Off. Max and Pro controls were exercised without sending messages.
  3. Genuine app, fresh HOME/XUM_ROOT, fake non-secret key and non-serving loopback endpoint. Backend/Vite each had exactly nine allowlisted environment names and zero extras. Source stayed clean; owned browser/server processes were stopped.

Limits: this was an unsent scratch draft, not a persistent-workspace send. No live API, Codex or gateway call. Negative Pro-route gating is covered by targeted automated tests, not these UI recordings. A supplementary phone Sol-reload clip showed only its earlier state and was excluded; the phone Sol reload has screenshot/snapshot evidence instead. The two accepted recordings below were decoded and inspected.

Screenshots

Both released GPT-6 models in the normal selector

GPT-6 Sol reasoning controls at 390px

GPT-6 Luna reasoning controls at 390px

Desktop recording

desktop.webm

Phone recording

mobile.webm

Generated with xum • Model: coder:openai/gpt-6-astra • Thinking: xhigh • Cost: $75.88

@ThomasK33

Copy link
Copy Markdown
Member Author

Final handoff — blocked for merge, not a release-ready verdict

Both GPT-6 Sol and Luna are implemented at 40e4b2a6b8d116d75fcda76b156ea9d5a5a0520c.

  1. Local make static-check and 855 tests / 18 files passed. Exact-head CI run 35771257885 and Visual Regression Testing passed. Optional Pixel / Review remains pending.
  2. Desktop/390px fixture UI validation passed; screenshots and recordings are uploaded and hash-verified. No real inference was sent. Official release is verified independently; gateway/account access is not.
  3. Stopping at the authorized review cap: three paired rounds = six code/security executions, plus the existing one advisor pass. All four threads are resolved, but the last verdict covers 64169a3, not this launch head. wait_pr_codex.sh --once still returns 10 (pending). The green CI Codex Comments check is not fresh approval. No further review/advisor was requested or review-board state changed.
  4. PR remains draft/unmerged. Current-head review coverage requires new maintainer direction; no gate is waived. The historical advisor recommendation (scope reduction before release) remains recorded, not relabeled as approval. Pre-existing GPT-5.6 pricing and Astra-cap discrepancies are tracked separately in 🤖 fix: reconcile GPT-5.6 Sol promotional pricing and Astra Codex context cap #4347.

The PR body preserves the full post-cap sequence, including the earlier f04892f and 782cca6 work. This is the final validation/readiness record for the expanded handoff.


Generated with xum • Model: coder:openai/gpt-6-astra • Thinking: xhigh • Cost: $75.88

@ThomasK33
ThomasK33 marked this pull request as ready for review September 22, 2026 19:38

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 40e4b2a6b8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/common/types/thinking.ts Outdated
@ThomasK33
ThomasK33 marked this pull request as draft September 22, 2026 19:50
@ThomasK33
ThomasK33 marked this pull request as ready for review September 22, 2026 20:26

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bfc0906fa6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/common/utils/ai/providerOptions.ts Outdated
@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a4c7c15635

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/providerModelFactory.ts Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛡️ Codex Security Review · Automatically triggered

Here are some automated security review suggestions for this pull request.

Reviewed commit: a4c7c15635

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Comment thread src/common/utils/tokens/models-extra.ts Outdated
@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: c9a0f57238

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: c9a0f57238

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Base automatically changed from gpt6-reasoning-none-followup to main September 23, 2026 14:37
@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Nice work!

Reviewed commit: 76bff932fd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 76bff932fd

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 23, 2026
## Summary

Bumps `packages/mux-compat/package.json` (`version` and the `@coder/xum`
dependency pin) from 0.29.0 to 0.30.0 to match the root package after
the `release: v0.30.0` commit (`81b0b744`), which bumped only the root
`package.json`.

## Background

`src/common/compat/productIdentity.test.ts` › "keeps the published mux
forwarding package version-locked to @coder/xum" asserts
`legacyPackageJson.version === packageJson.version`. Since the release
commit it fails on `main` ([run
35783501719](https://github.com/coder/xum/actions/runs/35783501719)) and
on every PR's merge commit (e.g. coder#4259 [run
35785546631](https://github.com/coder/xum/actions/runs/35785546631),
coder#4350), so `Required` is red repo-wide and the merge queue cannot admit
anything. Same fix as coder#4048 for v0.28.3; release commits since (e.g.
v0.28.5, coder#4175) bumped both files together. `@coder/xum@0.30.0` is
already published (Publish to NPM run 35783501943), so the new pin
resolves.

## Validation

`bun test src/common/compat/productIdentity.test.ts` on this branch: 8
pass (the version-lock case fails on `main` without it).

---

_Generated with `xum` • Model: `coder:anthropic/claude-fable-5-1` •
Thinking: `xhigh` • Cost: `$174.60`_

<!-- mux-attribution: model=coder:anthropic/claude-fable-5-1
thinking=xhigh costs=174.60 -->
… correct their Codex context cap

Stacked on #4350, which makes a requested reasoning_effort "none" reach
the wire for GPT-6 Sol/Luna. Headless tool loops (Dream consolidation,
memory harvest, refine, sidebar status) call streamText with tools and
no provider options, so nothing requests "none" and Chat Completions
defaults to medium, which rejects function calling. A model-boundary
middleware fills in "none" for tool-bearing Chat Completions requests
that asked for no effort, wrapped outside #4350's body-patching wrapper
so the injected value survives the SDK; explicit caller efforts are left
alone. The Coder openai-compat path resolves scoped "Treat as" aliases
through resolveModelForMetadata before clamping. The assumed 372K Codex
OAuth cap for Sol/Luna becomes the 272K published in the pinned Codex
catalog.

_Generated with `xum` • Model: `coder:anthropic/claude-fable-5-1` • Thinking: `xhigh` • Cost: `$174.60`_

<!-- mux-attribution: model=coder:anthropic/claude-fable-5-1 thinking=xhigh costs=174.60 -->
@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: bde7c0fe03

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bde7c0fe03

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/common/constants/codexOAuth.ts
The Astra comment still claimed the catalog lists the same window for
Astra and Sol, contradicting the 272K Sol/Luna entries below. State
that Astra keeps its pre-existing cap (tracked in #4347) independently.

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `xhigh` • Cost: `$281.40`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=xhigh costs=281.40 -->
@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: 7deeea6e94

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 7deeea6e94

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@ThomasK33
ThomasK33 added this pull request to the merge queue Sep 23, 2026
Merged via the queue into main with commit ca425b4 Sep 23, 2026
19 of 20 checks passed
@ThomasK33
ThomasK33 deleted the gpt-6-sol-prep branch September 23, 2026 20:50
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 26, 2026
…cap (coder#4684)

## Summary

Reconciles two OpenAI catalog entries with their first-party sources:
GPT-5.6 Sol now prices at OpenAI's current promotional rates, and GPT-6
Astra's Codex OAuth context cap drops from 372K to the catalog's 272K.

## Background

Both discrepancies were found while validating coder#4259 and tracked in
coder#4347.

**Pricing source.** [OpenAI API
pricing](https://developers.openai.com/api/docs/pricing), fetched
2026-09-26. GPT-5.6 Sol (per 1M tokens):

| | Input | Cached input | Cache writes | Output |
|---|---|---|---|---|
| Short context (≤272K) | $4.00 | $0.40 | $5.00 | $20.00 |
| Long context (>272K) | $8.00 | $0.80 | $10.00 | $30.00 |

The page says "GPT-5.6 Sol's promotional pricing is available at least
through November 21, 2026." The previous entry ($5/$30/$0.50/$6.25; long
context $10/$45/$1/$12.50) was the standard rate. The multipliers (2x
input, 1.5x output, 1.25x cache writes, 272K boundary) are unchanged.

**Catalog source.** [openai/codex@04fc75ad
`models.json`](https://github.com/openai/codex/blob/04fc75adbe67a612a1cb0fc469533f24b24fa499/codex-rs/models-manager/models.json),
the pin the repo already cites: `gpt-6-astra` has `context_window:
272000` and `max_context_window: 872000`. openai/codex `main` (checked
2026-09-26) has the same values.

## Implementation

- `models-extra.ts`: `GPT_56_SOL_STATS` (shared by `gpt-5.6-sol` and the
bare `gpt-5.6` alias) now uses the promo rates.
- Time-limited pricing follows the most recent repo precedent (Gemini
3.7/3.8 Flash intro rates): encode the rate actually billed, with a
dated `TODO(2026-11-21)` that records the standard rates and the source.
No new mechanism. The older precedent (Claude Sonnet 5, DeepSeek V4 Pro)
listed the standard rate instead. I picked the billed rate because the
issue asks for it, the promotion has no fixed end date, and the pricing
page no longer shows a standard rate for Sol.
- `codexOAuth.ts`: `gpt-6-astra` 372K → 272K, with the pin cited next to
Sol/Luna. Public API windows (1.05M) and user-mapped models are
unchanged: the cap applies only when a request actually routes through
Codex OAuth.

## Validation

Tests were changed first and failed on `main`:

- `displayUsage.test.ts` › "applies %s pricing only above the 272K
boundary" now includes `openai:gpt-5.6-sol` and `openai:gpt-5.6`. It
checks 272,000 vs 272,001 prompt tokens and now also checks cache writes
(1.25x the active input rate), so the boundary is shown to move input,
cache-read, cache-write and output rates together.
- `codexOAuth.test.ts` (Astra override) and `contextLimit.test.ts` ›
"caps GPT-6 Astra on the OAuth route but keeps the API window for
API-key auth" (272K on OAuth, 1.05M with API-key auth).
- `modelStats.test.ts` base-rate row for Sol updated. Its long-context
assertions are multipliers of the base rates, so they checked the new
values unchanged.

<details>
<summary>Pre-fix failures</summary>

```
Expected: 272000
Received: 372000
(fail) codexOAuth model gating > allows gpt-6-astra through Codex OAuth with a 272000 context cap
Expected: 0.668
Received: 0.8350000000000001
(fail) createDisplayUsage > tiered long-context pricing > applies openai:gpt-5.6-sol pricing only above the 272K boundary [1.00ms]
(fail) createDisplayUsage > tiered long-context pricing > applies openai:gpt-5.6 pricing only above the 272K boundary
(fail) getEffectiveContextLimit > caps GPT-6 Astra on the OAuth route but keeps the API window for API-key auth
 60 pass
 4 fail
```

</details>

After the fix, `bun test src/common` passes (2380 tests) and `make
static-check` passes.

## Risks

Low. Sol cost estimates and goal budgets drop by 20–33% until the
promotion ends. The dated TODO marks when to restore them. OAuth-routed
Astra now starts limit-driven compaction at 272K instead of 372K.

## Deferred

- The GPT-5.6 family Codex caps (372K since coder#3730) also differ from the
catalog (272K). Out of scope here: tracked in coder#4683.

Fixes coder#4347

---

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking:
`high`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high -->
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 27, 2026
…log (coder#4710)

## Summary

The Codex OAuth context cap for the GPT-5.6 family (`gpt-5.6`,
`gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`) goes from 372K to 272K,
matching the first-party Codex catalog. API-key requests keep the public
1.05M window.

## Background

coder#4683 asked where coder#3730's 372K came from before changing it.

- coder#3730 (merged 2026-07-15) says it copied the `openai/codex` catalog,
which "lists `context_window: 372000` for Sol, Terra, and Luna". That PR
has no discussion of observed Codex behavior and no deliberate override.
- Catalog history (`codex-rs/models-manager/models.json`):
-
[`c38c30ef`](https://github.com/openai/codex/blob/c38c30ef514dc74aa6ea59748f280c235b1bca2a/codex-rs/models-manager/models.json)
(2026-07-14): Sol/Terra/Luna `context_window: 372000`,
`max_context_window: 372000`.
-
[`2eee483e`](openai/codex@2eee483)
(openai/codex#39102, "Raise the GPT-5.6 maximum context window",
2026-08-17): `context_window: 272000`, `max_context_window: 872000`.
- The pin the repo already cites
([`04fc75ad`](https://github.com/openai/codex/blob/04fc75adbe67a612a1cb0fc469533f24b24fa499/codex-rs/models-manager/models.json))
and `main` (checked 2026-09-26) both still have 272000 / 872000.

So 372K came only from an older catalog. This PR uses the default
`context_window`, the same convention as GPT-5.5 and the GPT-6
Astra/Sol/Luna entries (coder#4259, coder#4684).

## Validation

The routing and meter expectations were updated first and failed on
`main`: `codexOAuth.test.ts` (GPT-5.6 family overrides),
`contextLimit.test.ts` (OAuth-routed effective limit per tier), and
`tokenMeterUtils.test.ts` (OAuth meter max and percentage for Sol). The
API-key → 1.05M assertions in those files are unchanged and still pass.

<details>
<summary>Pre-fix failures</summary>

```
Expected: 272000
Received: 372000
(fail) codexOAuth model gating > allows the GPT-5.6 family through Codex OAuth without requiring it
(fail) calculateTokenMeterData > uses the Codex OAuth cap for GPT-5.6 token meter percentages
(fail) getEffectiveContextLimit > caps the GPT-5.6 family at each tier's Codex OAuth context window
 36 pass
 3 fail
```

</details>

## Risks

Low. Only direct OpenAI requests routed through Codex OAuth are
affected. They now start limit-driven compaction at 272K instead of
372K, which is the documented default window. API-key, gateway and Coder
routes are unchanged.

Fixes coder#4683

---

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking:
`high`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant