Skip to content

🤖 feat: add GPT-6 Sol and Luna model support - #4348

Merged
ammario merged 4 commits into
mainfrom
model-catalog-xra9
Sep 22, 2026
Merged

ammario merged 4 commits into
mainfrom
model-catalog-xra9

Conversation

@ammar-agent

@ammar-agent ammar-agent commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add GPT-6 Sol and GPT-6 Luna, released September 22, to the curated model catalog. Promote gpt/sol and luna to the new tiers while preserving older model strings. Today's Anthropic release, Claude Opus 5.5, is already supported on main by #3993; this PR completes the missing OpenAI additions rather than duplicating that work.

Implementation

  • Add published pricing, cache rates, the 272K long-context pricing threshold, 922K maximum input, and 128K output limits.
  • Support none through native max reasoning and Pro mode on supported Responses API routes, without broadening Astra's always-on reasoning rule or matching unannounced variants.
  • Send explicit none reasoning for Sol/Luna on opt-in Chat Completions routes so agent function calls remain supported; preserve selected reasoning on Responses. Cover canonical IDs, dated IDs, and mapped aliases.
  • Allow Sol/Luna through Codex OAuth, assuming Codex serves them like Astra, with Astra's 372K Codex context cap. The workspace-naming fallback moves to GPT-6 Luna. Explicit-cache opt-in remains GPT-5.6-only.
  • Keep GPT-5.6 Sol/Luna tokenizer overrides so older model strings still count tokens correctly.
  • Update model documentation and generated built-in help.

Sources: OpenAI release notes, Sol, Luna, reasoning modes, Opus 5.5.

Validation

  • Targeted model, pricing, reasoning, context, auth-gating, and Codex OAuth routing tests passed, along with the full local bun test src run.
  • Reasoning selector and thinking-persistence UI integration tests passed (persistence rerun in isolation after a shared-process Jest global-setter recursion).
  • make static-check and make static-check-full passed, including documentation links.
  • No live provider API calls were made.

Risks

The short OpenAI aliases now select the newer tiers. Existing fully-qualified model IDs remain usable. Codex OAuth availability and the 372K cap for Sol/Luna are assumptions, not yet confirmed by a published Codex catalog. If either is wrong, OAuth-only users of gpt/sol/luna get request errors or early compaction until the allowlist is adjusted.


Generated with xum • Model: openai:gpt-6-astra • Thinking: high • Cost: $24.87

Complete September 22 model coverage alongside the Opus 5.5 support already on main.

---

_Generated with `xum` • Model: `openai:gpt-6-astra` • Thinking: `high` • Cost: `$12.61`_

<!-- mux-attribution: model=openai:gpt-6-astra thinking=high costs=12.61 -->
@mintlify

mintlify Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
Mux 🟢 Ready View Preview Sep 22, 2026, 8:06 PM

💡 Tip: Enable Automations to automatically generate PRs for you.

@ammar-agent

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-22T20:56:04.017565Z 2e46d26 Manual request
🔒 Security Review ✅ Completed 2026-09-22T20:53:47.453908Z 2e46d26 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 23bdfc2909

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 23bdfc2909

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/common/utils/tokens/models-extra.ts
Disable Sol/Luna reasoning on Chat Completions while retaining the selected effort on Responses.

---

_Generated with `xum` • Model: `openai:gpt-6-astra` • Thinking: `high` • Cost: `$12.61`_

<!-- mux-attribution: model=openai:gpt-6-astra thinking=high costs=12.61 -->
@ammar-agent

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 6a070448c3

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 🚀

Reviewed commit: 6a070448c3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

GPT alias now tracks GPT-6 Sol, which is not yet OAuth-allowlisted.
@ammar-agent

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: fcf07b311d

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: fcf07b311d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Assume Codex serves the new tiers like Astra (372K cap); restore GPT-6 Luna as the naming fallback.

---

_Generated with `xum` • Model: `openai:gpt-6-astra` • Thinking: `high` • Cost: `$24.87`_

<!-- mux-attribution: model=openai:gpt-6-astra thinking=high costs=24.87 -->
@ammar-agent

Copy link
Copy Markdown
Collaborator Author

@codex review

@ammario
ammario merged commit bd66643 into main Sep 22, 2026
18 of 20 checks passed
@ammario
ammario deleted the model-catalog-xra9 branch September 22, 2026 20:51
@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 2e46d26ef2

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2e46d26ef2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +121 to +122
"gpt-6-sol",
"gpt-6-luna",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep unverified GPT-6 tiers off Codex OAuth

When either new model is selected and Codex OAuth is connected—including API-key users whose default auth remains oauth—these entries make providerModelFactory.ts route the request to the ChatGPT Codex endpoint. The current Codex catalog does not advertise these IDs, and the adjacent comment explicitly labels their availability as assumed, so the promoted gpt/sol and luna choices can return unsupported-model errors instead of using a configured API key; OAuth-only users also see unusable entries. Keep them API-key-only until Codex endpoint support is verified.

Useful? React with 👍 / 👎.

ThomasK33 added a commit that referenced this pull request Sep 22, 2026
…their Codex context cap

#4348 added GPT-6 Sol/Luna and clamps their Chat Completions reasoning
effort to "none" in buildProviderOptions, but @ai-sdk/openai@4.0.71
strips "none" for every gpt-6 model, so the clamp never reached the
request body and tool-bearing turns were rejected upstream. This adds a
Bun patch that accepts "none" for exact Sol/Luna ids only, enforces the
same clamp at the model boundary for headless tool loops (including
Coder openai-compat aliases resolved through their scoped mapping), and
replaces the assumed 372K Codex OAuth cap with the 272K published in the
pinned Codex catalog. Wire-level regressions cover both.

_Generated with `xum` • Model: `coder:anthropic/claude-fable-5-1` • Thinking: `xhigh` • Cost: `$174.60`_

<!-- mux-attribution: model=coder:anthropic/claude-fable-5-1 thinking=xhigh costs=174.60 -->
ThomasK33 added a commit that referenced this pull request Sep 23, 2026
Merge-queue admission failed after #4348 renamed GPT_56_LUNA to GPT_6_LUNA
and moved the `gpt` alias to gpt-6-sol. Swap the renamed key in the explicit
picker-order lists and derive the gateway catalog entries for the routed
`gpt` built-in from KNOWN_MODELS, so the gating assertions keep testing
routing instead of one historical model ID. No production files change.

---

_Generated with [`xum`](https://github.com/coder/xum) • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$237.64`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high costs=237.64 -->
ThomasK33 added a commit that referenced this pull request Sep 23, 2026
Merge-queue admission failed after #4348 renamed GPT_56_LUNA to GPT_6_LUNA
and moved the `gpt` alias to gpt-6-sol. Swap the renamed key in the explicit
picker-order lists and derive the gateway catalog entries for the routed
`gpt` built-in from KNOWN_MODELS, so the gating assertions keep testing
routing instead of one historical model ID. No production files change.

---

_Generated with [`xum`](https://github.com/coder/xum) • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$237.64`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high costs=237.64 -->
ammario pushed a commit that referenced this pull request Sep 23, 2026
…4350)

## Summary

Follow-up to #4348. GPT-6 Sol/Luna now actually send `reasoning_effort:
"none"`. The PR also fixes two post-merge `Test / Unit` failures on
`main`: a unit test file that crashed Bun, and the v0.30.0 version-lock
contract.

## Background

- **SDK drops `none`:** `@ai-sdk/openai@4.0.71`, including the latest
4.0.72, allows only `low`/`medium`/`high`/`xhigh`/`max` for every
`gpt-6-*` ID and silently strips any other effort from Chat Completions
and Responses requests. That list is correct for Astra, but Sol and Luna
accept `none`. OpenAI documents that an omitted effort defaults to
`medium` and that Chat Completions function calling requires `none`. So
after #4348:
- Chat Completions agent turns on `gpt`, `sol` and `luna` had their tool
calls rejected.
  - On Responses, choosing "off" silently ran at medium reasoning.
- #4348's tests only checked the `buildProviderOptions` output, not the
serialized request, so they missed this.
- **CI:** the `Test / Unit` job in `main` run
[35782976621](https://github.com/coder/xum/actions/runs/35782976621)
exited with code 133. Bun hit an allocator panic (`pas panic:
deallocation did fail`) just as `mcpIconDecodeClient.test.ts` started,
with no failing assertions. That file loads the native `sharp` addon
into the shared coverage process. The crash is not caused by #4348's
model changes.
- **CI (release):** `release: v0.30.0` (81b0b74) bumped `package.json`
but not `packages/mux-compat`. `productIdentity.test.ts`, which requires
the legacy `mux` forwarding package to stay version-locked to
`@coder/xum`, therefore fails on `main`.

## Implementation

- Generalizes `createOpenAIModelWithServiceTier` into
`createOpenAIModelWithPreservedOptions`. For Sol/Luna wire IDs
(including dated IDs) with effort `none`, it removes the effort from the
SDK options and writes it into the serialized body after the SDK's
capability checks. This reuses the per-call fetch-wrapper pattern
already used for service tiers.
- The wrapper covers native OpenAI (including Codex OAuth, whose tier
behavior is unchanged), custom `openai-responses` providers and the
Coder gateway OpenAI route.
- Moves `mcpIconDecodeClient.test.ts` into `isolated_unit_tests`,
following the existing pattern, so it runs in its own process with the
signal-exit retry.
- Bumps `packages/mux-compat` (version and its `@coder/xum` dependency)
to 0.30.0.

## Validation

- A new factory-level wire test covers Sol, Luna and dated Sol over both
wire formats, with a tool attached, and asserts that `none` is sent. Two
controls check that Sol `high` passes through and that Astra `none` is
still dropped. With the fix disabled, the 6 `none` cases fail.

## Risks

Low. Only requests to Sol/Luna wire IDs with effort `none` are
rewritten. Other models, and Sol/Luna at other efforts, keep the
existing path.

**Deferred (pre-existing):** on Responses, the SDK treats opaque aliases
(for example `team-model` → `openai:gpt-6-sol`) as non-reasoning models
and drops every reasoning effort, not just `none`. This already affected
aliases of any reasoning model before #4348. On Chat Completions, which
is where `none` matters for tool calls, those aliases already serialize
`none`, and a test now covers that.

<details>
<summary>Review history</summary>

The first revision fixed this with a `bun patch` of `@ai-sdk/openai`.
Codex pointed out that `patchedDependencies` does not apply to npm
installs of the published package, so the fix now lives in Xum runtime
code and the patch has been removed.

</details>

---

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking:
`high` • Cost: `$26.31`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high
costs=26.31 -->
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 23, 2026
…#4346)

## Summary

Add `models_list`, a read-only tool that returns the visible models
selectable under the current configuration, with their accepted aliases
and thinking levels. It shares the composer picker's selection pipeline,
so agents can discover a valid `task.model` instead of guessing.

The catalog is advisory. It does not probe providers or grant permission
to override the model. `task.model` still says to omit it unless the
user requests a specific model, and send-time validation remains
authoritative.

## Implementation

1. Extract the existing picker filters into a shared module. Preserve
routing, credentials, authoritative catalogs, OAuth, policy,
hidden-model filtering, and UI loading behavior.
2. Register the zero-argument tool through normal tool assembly. Read
current backend configuration on each invocation. Return only model IDs,
aliases, and thinking levels; return fixed error text on failure.
3. Keep the extraction and tool addition in separate commits. Follow
them with test-only repairs needed to make the required full gate pass.

## Validation

- Phase A: 42 existing hook tests passed before and after extraction;
the broader parity gate passed 2,191 tests.
- Feature coverage: 399 targeted tests passed. Tests cover selector
branches, fresh configuration reads, task input compatibility and
forwarding for both task kinds, agent policy exposure, and secret-safe
results.
- Final head `1eb6069124` (rebased onto `main` `80efaa2ac5` after
merge-queue admission failed on coder#4348's GPT-6 catalog rename): `make
static-check` passed. Full `make test-unit` passed in 1,173 seconds:
20,089 source tests passed, 9 skipped, and 41 Storybook/DOM tests
passed. The follow-up commit only updates `selectableModels.test.ts`
fixtures (renamed `GPT_6_LUNA` key; gateway catalog entries derived from
the moving `gpt` alias).
- Local full tests used pinned Bun 1.3.5 and process-only
`GIT_TEMPLATE_DIR=/usr/share/git-core/templates` to avoid host-specific
Git templates. No test-runner flags, timeout increases, or skipped
coverage were added.
- Remote Coder UAT passed on `31d4e64ab` (the recorded/tested SHA; not a
rerun on the final head): picker/tool parity; tool availability in exec,
plan, explore and sub-agents; fixture-backed launches for both task
kinds; disabled-provider negative checks; and desktop/mobile rendering.
Requests went to controlled loopback fixtures, not real providers.
- The mobile recording shows the actual expanded tool card at 390×844
with the sidebar drawer closed, including both model entries after
scrolling the result pane. The turn was submitted through the production
CLI; this does not validate the mobile Send button. Three earlier
evidence attempts were retained; the first round-two video remains
disqualified because its launcher did not meet the environment-allowlist
requirement. The separately authorized final attempt met the original
criterion without source changes.
- [Remote UAT
chat](https://dogfood.cdr.dev/agents/ce4d68f9-139c-4a91-9e91-a2c4e20f29b8).
Screenshots, sequentially decoded recording frames, request transcripts,
and artifact hashes were independently inspected.

<details>
<summary>Test-gate repairs and plan corrections</summary>

The full run exposed test isolation leaks. The separate repair commit
restores Dialog/Tooltip module mocks, uses the shared DOM teardown
harness, and targets filesystem error injection at the intended path.
Ordered reproductions for the DOM and module-mock failures went from red
to green; the final full run passed.

CI then exposed two moving-model fixture assumptions after the base
branch promoted the Opus alias to 5.5. Reproducing the exact CI merge
tree produced both failures. The final test-only commit derives the
gateway allowlist from model metadata and verifies enrichment against
the authoritative thinking policy, while keeping fixed-model capability
expectations explicit. All 270 selector, AI-service, and thinking-policy
tests pass on both the feature tree and repaired CI merge tree. Only two
test files differ from the UAT-tested commit; all other tracked files,
including product code and dependencies, are identical. Following the
user's instruction to continue (standing delegation), the existing UAT
evidence was accepted as carried forward to this test-only commit on the
basis of the verified two-test-file delta and green final-head
validation. The evidence retains its actual tested SHA (`31d4e64ab`); no
extra recording was made.

The retained Storybook budget had fallen behind the current inventory.
Its separate repair sets zero-headroom limits at 115 enabled files and
605 estimated snapshots, without removing coverage. An audited inventory
and disposable extra-file/extra-export mutations verify both guards
still fail on growth. These are estimates, not executed capture counts.

Three plan assumptions were corrected against existing behavior: the
hook-doc generator omits zero-input tools, so only the `task.model` hint
changes; custom IDs are trimmed before selection; internal colons are
valid model-ID characters. The generator and task normalization behavior
are unchanged.

</details>

## Risks

The shared selection code is on the composer hot path. The extraction
preserves its behavior and has explicit parity coverage. Backend
discovery reads persisted hidden-model preferences, so it can briefly
differ from the UI immediately after a preference change. The catalog is
not a promise that a provider will accept a request.

## Mobile UAT evidence

![Expanded models_list tool card at 390
pixels](https://github.com/user-attachments/assets/940de44f-5c88-4c75-8b85-fc97cf1776a4)

Continuous mobile recording: expand the card and scroll both model
entries.


https://github.com/user-attachments/assets/8aa66906-454d-42c2-bd5c-ede1f19ea2a9

---

<details>
<summary>📋 Implementation Plan</summary>

# Plan: `models_list` tool — let agents discover valid `model` values
for `task`

## Goal

Add a read-only agent tool, `models_list`, that returns the models
**selectable under the user's current configuration** (the same set the
chat composer's model picker offers), each with the aliases and thinking
levels the `task` tool accepts. Agents can then pass a valid
`model`/`thinking` to `task` (both `kind: "subagent"` and `kind:
"workspace"`) instead of guessing. The list is computed by one shared
function used by both the UI picker and the tool, so the two cannot
drift.

The catalog is **advisory**: it reflects configuration (credentials
present, routing, policy, user-hidden models), not a guarantee that a
provider will accept a request. Send-time validation in
`providerModelFactory` stays authoritative. It is also a *visible*
catalog: user-hidden models are omitted even though `task.model` would
still accept them.

## Evidence (verified in repo)

- No backend code computes "selectable models" today. The only
implementation is the React hook
`src/browser/hooks/useModelsFromSettings.ts` (`useMemo` ~L347–418):
dedupe(custom models of enabled providers, `KNOWN_MODELS`) → minus
`hiddenModels` → authoritative-catalog + route availability
(`isModelAvailable`) → OpenAI Codex-OAuth gating on the direct route →
policy on the active route. The composer picker
(`src/browser/features/ChatInput/index.tsx:2941–2959` → `ModelSelector`)
renders `useModelsFromSettings().models` as-is (no sorting, no "recent"
entries); Settings → Models shows inventory tables, not this list.
**Parity reference = `useModelsFromSettings().models`.**
- `task.model` is validated syntactically only, for both kinds:
`parseTaskAiOverrides` (`src/node/services/tools/task.ts:120–147`,
called at L419 before the `kind` branch) → `normalizeModelInput`
(`src/common/utils/ai/normalizeModelInput.ts`) → `MODEL_ABBREVIATIONS`
alias table (`src/common/constants/knownModels.ts:307`) →
`normalizeSelectedModel` (preserves explicit gateway prefixes) →
`isValidModelFormat`. A well-formed but unavailable model passes
creation and fails at first send (`api_key_not_found`,
`model_not_available`, `policy_denied`, …), surfacing as a failed `task`
call or an `interrupted` task.
- Reusable helpers already in `src/common`:
`resolveRoute`/`isModelAvailable`
(`src/common/routing/resolve.ts:280,305`),
`isProviderModelAccessibleFromAuthoritativeCatalog`/`isGatewayModelAccessibleFromAuthoritativeCatalog`
(`src/common/utils/providers/gatewayModelCatalog.ts`),
`isCodexOauthAllowedModel`/`isCodexOauthRequiredModel`
(`src/common/constants/codexOAuth.ts`),
`getThinkingPolicyForModel(model, providersConfig?: ProvidersConfigMap |
null): readonly ThinkingLevel[]`
(`src/common/utils/thinking/policy.ts:108`, honors `mappedToModel`),
`normalizeModelInput`, `KNOWN_MODELS`, `DEFAULT_HIDDEN_MODELS`,
`MODEL_ABBREVIATIONS`.
- Browser-only helpers the shared pipeline needs: `parseModelString`,
`isModelAllowedByPolicy`, `isGatewayModelAccessibleForUi` in
`src/browser/utils/policyUi.ts` (imports are all `@/common`; 10
importers incl. `src/node/services/providerService.test.ts`), plus
private helpers in the hook (`getCustomModels`, `getSuggestedModels`,
`filterHiddenModels`, `dedupeKeepFirst`, `resolvesToDirectOpenAI`,
`isModelAllowedByPolicyOnActiveRoute`; none imported elsewhere).
- UI inputs and their backend sources: `providersConfig` =
`api.providers.getConfig()` → `providerService.getConfig()`
(`ProvidersConfigMap`); `routePriority`/`routeOverrides` =
`config.getConfig()` with UI defaults `["direct"]` / `{}`
(`src/browser/hooks/useRouting.ts:19–21,82`); `hiddenModels` = persisted
browser state seeded from backend model prefs on mount
(`src/browser/contexts/WorkspaceContext.tsx:683–705`), default
`DEFAULT_HIDDEN_MODELS`; `effectivePolicy` = `usePolicy()` when
enforced.
- Backend has the same inputs where tools are assembled:
`TurnRequestBuilderDependencies`
(`src/node/services/turnRequestBuilder.ts:585–600`) has `config:
Config`, `providerService: ProviderService`, `policyService?:
PolicyService`. `Config.loadConfigOrDefault()`
(`src/node/config/index.ts:1456`) returns `ProjectsConfig`
(`src/common/types/project.ts:79`; fields `hiddenModels?`,
`routePriority?`, `routeOverrides?`) from a warm in-memory snapshot; it
is the accessor every backend read uses (there is no separate
`getConfig()`), and its one-time migration persistence is pre-existing
behavior shared by all reads — the tool adds no write path.
- **Single production wiring point**: `TurnRequestBuilder` builds
`ToolConfiguration` at L2247 and calls `getToolsForModel` at L2518. CLI
`xum run`, ACP, workflows, and the PTC bridge all flow through it
(`src/cli/run.ts:651–684`, `src/node/acp/agent.ts:121`,
`src/node/services/workflowContinuation.ts:28`,
`src/node/services/ptc/toolBridge.ts:43` receives already-assembled
tools). The only other `ToolConfiguration` construction
(`src/node/services/refinement/refineService.ts:1476–1488`) builds
skill-write tools only and never calls `getToolsForModel`.
- Tool plumbing: schema in `TOOL_DEFINITIONS`
(`src/common/utils/tools/toolDefinitions.ts`, optional `resultSchema`),
arg/result types in `src/common/types/tools.ts`, implementation
`src/node/services/tools/<name>.ts` (`ToolFactory = (config:
ToolConfiguration) => Tool`), registration in `nonRuntimeTools`
(`src/common/utils/tools/tools.ts` ~L871), gating in
`getAvailableTools()` `baseTools` (~L3770). Built-in agents
`exec.md`/`plan.md`/`explore.md` use `tools.add: [".*"]` with explicit
removes, so a new tool reaches exec/plan/explore and sub-agents without
agent-definition edits (`SUBAGENT_HARD_DENY` only denies
`ask_user_question`; depth denial covers `task*` only). UI falls back to
`GenericToolCall`; icon via `TOOL_NAME_TO_ICON`
(`src/browser/features/Tools/Shared/ToolPrimitives.tsx:244`).
`docs/hooks/tools.mdx` is regenerated from `TOOL_DEFINITIONS` by `bun
scripts/gen_docs.ts` (`make fmt`). No test asserts a total tool count or
schema token budget.
- Template: `agent_skill_list`
(`src/node/services/tools/agent_skill_list.ts:158`): `{ success: true, …
} | { success: false, error }`, `.strict()` schema, `.nullish()`
optionals.
- Existing test patterns for real assembly/gating:
`src/common/utils/tools/toolDefinitions.test.ts:683–753`
(`getAvailableTools`), `src/common/utils/tools/tools.test.ts`
(`getToolsForModel`), `src/node/services/toolAssembly.test.ts`
(`applyToolPolicyAndExperiments` + `resolveToolPolicyForAgent` for
built-in agents).

## Design decisions

1. **Name**: `models_list` (plain snake_case like `task_list`,
`agent_skill_list`; `mux_` prefix is reserved for config/agent-store
tools).
2. **Input**: `z.object({}).strict()` — no parameters, no
`includeHidden` option.
3. **Output** (`ModelsListToolResultSchema`, also attached as
`resultSchema`). The domain type lives in the shared module; the schema
only mirrors it (type-only dependency, no cycle):
   ```ts
   // src/common/utils/ai/selectableModels.ts
   export interface AvailableModel {
model: string; // normalized selection ID accepted by task.model
(explicit gateway prefixes preserved)
aliases: string[]; // MODEL_ABBREVIATIONS keys that normalize to `model`
(e.g. ["sonnet"]); [] for custom models
thinkingLevels: ThinkingLevel[]; // Xum's thinking policy for this model
(getThinkingPolicyForModel), not a provider capability probe
   }
   // src/common/utils/tools/toolDefinitions.ts
   const AvailableModelSchema = z.object({
model: z.string(), aliases: z.array(z.string()), thinkingLevels:
z.array(ThinkingLevelSchema), // ThinkingLevelSchema exists in
src/common/types/thinking.ts
   }).strict() satisfies z.ZodType<AvailableModel>;
   const ModelsListToolResultSchema = z.discriminatedUnion("success", [
z.object({ success: z.literal(true), models:
z.array(AvailableModelSchema) }).strict(),
z.object({ success: z.literal(false), error: z.string() }).strict(),
   ]);
   ```
`getThinkingPolicyForModel` returns a readonly array → spread into a
mutable array when building the entry. Nothing else (no display names,
context windows, pricing, `defaultModel`, credentials, provider config).
4. **Semantics = composer-picker parity, defined as "valid, normalized,
selectable picker entries".** `computeSelectableModels` reproduces the
hook pipeline exactly (including the `providersConfig == null` "still
loading → skip availability filters" branch, used only by the UI) and
returns raw picker strings. The enrichment step then, for each raw entry
`raw` with `selectable = new Set(rawList)`:
- `n = normalizeModelInput(raw).model`; if `null` (e.g. a persisted
custom ID with a second colon) → **omit** (persisted custom IDs are
input data, not programmer invariants);
- if `n !== raw` (normalization changed the identity —
`normalizeModelInput` trims and canonicalizes; explicit gateway
selections such as `openrouter:…` are preserved by
`normalizeSelectedModel`, so in practice this is e.g. a persisted custom
entry with surrounding whitespace) → emit `n` **only if
`selectable.has(n)`**, i.e. the emitted identity itself passed routing,
credential, catalog, policy and hidden-model checks; otherwise omit
rather than advertise an unchecked replacement;
- dedupe after normalization (two raw entries normalizing to the same ID
collapse into one);
- `aliases` = `MODEL_ABBREVIATIONS` keys with
`normalizeModelInput(alias).model === n`.
Omissions are reported through an optional `onSkipped?: (raw: string,
reason: "malformed" | "unchecked_identity") => void` callback so the
Node boundary can `log.debug`; the shared module itself imports no
logger.
5. **Single source of truth.** The hook and the tool both call
`computeSelectableModels`; the hook's own code shrinks to input
plumbing.
6. **Service access**: one optional closure on `ToolConfiguration`,
`listAvailableModels?: () => AvailableModel[]`, built in
`TurnRequestBuilder` at the existing construction site. The closure
reads `providerService.getConfig()`, `config.loadConfigOrDefault()`, and
the *effective policy* (`policyService?.isEnforced() ?
policyService.getEffectivePolicy() : null`) **at each invocation** — no
cache, no new service class, no `PolicyService` passed into shared code.
Absent closure (test contexts, `refineService`) → `{ success: false,
error: "Model catalog unavailable in this context" }` (mirrors `task`
without `taskService`). A throwing closure → fixed message `{ success:
false, error: "Failed to compute the model catalog" }` with the
exception logged via `log.error` on the Node side — exception text
(which could echo configuration) never reaches the model.
7. **Backend config is the source even for SSH/remote workspaces**
(models are resolved on the Xum host, not in the workspace runtime); the
tool never probes providers.
8. **No agent-definition changes**; no `ptcExcluded`; generic CLI
formatter and UI renderer; one icon line.
9. **Prompt cost**: description ≤ ~60 words; it is added to every
request's tool schema.
10. **Discovery vs. override permission are separate.** `models_list`
may be called whenever the agent needs to know which models exist (e.g.
the user asks "which models are available?"). The *override* restriction
stays where it is: the `task.model` description keeps its "omit unless
the user explicitly instructed a specific model" sentence verbatim and
gains only the hint *"Use `models_list` to see valid values."*

## Delivery: one change set, two gated phases

Work lands on this branch as two commits (Phase A, Phase B). Publication
(PR, `gh stack` split) only when requested; the phases are independently
reviewable if a split is wanted later.

### Phase A — pure refactor, zero behavior change: shared
selectable-models pipeline

1. Create `src/common/utils/policy/modelPolicy.ts` with
`parseModelString`, `isModelAllowedByPolicy`,
`isGatewayModelAccessibleForUi` moved verbatim from
`src/browser/utils/policyUi.ts`; `policyUi.ts` re-exports them so its 10
importers and `policyUi.test.ts` are untouched
(`getAllowedProvidersForUi` stays in the browser file).
→ verify: `make typecheck`; `bun test
src/browser/utils/policyUi.test.ts`.
2. Move `DEFAULT_ROUTE_PRIORITY` (`["direct"]`) from `useRouting.ts` to
`src/common/routing/resolve.ts` (reuse an existing direct-route constant
there if one exists) so UI and backend default identically;
`useRouting.ts` imports it.
3. Create `src/common/utils/ai/selectableModels.ts`:
   ```ts
   export interface SelectableModelsInput {
providersConfig: ProvidersConfigMap | null; // null = still loading (UI
only): skip availability filters
     hiddenModels: string[];
effectivePolicy: EffectivePolicy | null; // null = policy not enforced
routePriority: string[]; // matches resolveRoute/isModelAvailable
signatures (not readonly)
     routeOverrides: Record<string, string>;
   }
export function computeSelectableModels(input: SelectableModelsInput):
string[];
   ```
Move `BUILT_IN_MODELS`, `getCustomModels`, `getSuggestedModels`,
`filterHiddenModels`, `dedupeKeepFirst`, `resolvesToDirectOpenAI`,
`isModelAllowedByPolicyOnActiveRoute`, and the `isConfigured` /
`isGatewayModelAccessible` / `isAuthoritativeProviderModelAccessible`
predicates from the hook into this file **without logic changes**
(predicates become plain closures over
`providersConfig`/`effectivePolicy`). Export
`getSuggestedModels`/`filterHiddenModels` only if the hook still needs
them (`getAllCustomModels`/`providerHiddenModels` bucket stays in the
hook).
4. `useModelsFromSettings`: replace the `useMemo` body with
`computeSelectableModels({ providersConfig: config, hiddenModels,
effectivePolicy, routePriority, routeOverrides })`; keep
`providerHiddenModels` / `ensureModelInSettings` logic in place, reusing
the moved predicates where it needs them.
5. Guard: `rg '@/browser' src/common` stays empty (lint
`no-restricted-imports` boundary also enforces this).

Gate A (must pass before Phase B): `make typecheck && make lint && bun
test src/browser/hooks/useModelsFromSettings.test.ts src/browser/utils
src/common/utils/ai src/common/routing` — the 40 existing hook tests run
**before** (baseline = this branch's pre-change HEAD `9a6a2e2bf`,
recorded before the first edit) and **after** the extraction with
identical results.

Tests added in Phase A — `src/common/utils/ai/selectableModels.test.ts`,
each with an explicit expected list (never "compare two calls"):
- custom models: included only from enabled providers; `mux-gateway` and
`github-copilot` entries skipped; a custom entry equal to a built-in ID
deduped (custom first); object entries (`{ id, mappedToModel }`) surface
by `id`.
- hidden: model in `hiddenModels` excluded; empty list is a no-op.
- route availability: built-in model of an unconfigured/disabled
provider excluded; same model included when a configured gateway
(openrouter / coder) can route it; `routeOverrides` pinning a model to
an unavailable gateway falls back per `resolveRoute` (assert the actual
`isModelAvailable` outcome); explicit gateway-prefixed IDs
(`coder:openai/gpt-…`) survive.
- authoritative catalogs: coder `discoveredModels` present/absent and
`removedModels` exclusion; github-copilot catalog gating of a built-in
model.
- Codex OAuth on the direct OpenAI route: key+oauth, key-only,
oauth-only, neither × required/allowed model; a gateway-routed OpenAI
model is not gated.
- policy: model denied under canonical identity but allowed via its
active gateway route is included; denied on the active route excluded;
`effectivePolicy: null` skips filtering.
- `providersConfig: null`: returns the hidden-filtered suggested list
(loading semantics), still policy-filtered when a policy is present.
(Move equivalent cases out of `useModelsFromSettings.test.ts` only when
they test pure pipeline logic; leave cases that exercise React state,
persistence, or `ensureModelInSettings`.)

### Phase B — the tool

1. `src/common/utils/tools/toolDefinitions.ts`
- Add `AvailableModelSchema`, `ModelsListToolResultSchema` (Design §3)
near the other result schemas (import `ThinkingLevelSchema` from
`src/common/types/thinking.ts`; `import type { AvailableModel }` from
the shared module).
- Add `models_list: { description, schema: z.object({}).strict(),
resultSchema: ModelsListToolResultSchema }`. Description: *"List models
selectable under the current configuration, with aliases and thinking
levels. Hidden models are omitted. This is a configuration snapshot, not
a provider availability probe. Use returned IDs when a model override is
requested; otherwise leave `task.model` unset."*
   - Add `"models_list"` to `baseTools` in `getAvailableTools()`.
- Append *"Use `models_list` to see valid values."* to the `task`
`model` description.
2. `src/common/types/tools.ts`: `ModelsListToolArgs`,
`ModelsListToolResult` (z.infer); re-export `type AvailableModel` from
the shared module for tool consumers.
3. `src/common/utils/ai/selectableModels.ts`: add the pure enrichment
step (Design §4)
   ```ts
   export function listAvailableModels(
input: SelectableModelsInput & { providersConfig: ProvidersConfigMap },
onSkipped?: (raw: string, reason: "malformed" | "unchecked_identity") =>
void
   ): AvailableModel[]
   ```
= `computeSelectableModels` → normalize / recheck / dedupe per §4 →
`aliases` → `thinkingLevels = [...getThinkingPolicyForModel(model,
providersConfig)]`. `assert(thinkingLevels.length > 0)` (the policy
helper always returns a non-empty fallback; this documents the
assumption). No logger import; no runtime import of tool-definition
modules.
4. `src/common/utils/tools/tools.ts`: add `listAvailableModels?: () =>
AvailableModel[]` to `ToolConfiguration` with a doc comment (`import
type`); register `models_list: createModelsListTool(config)` in
`nonRuntimeTools`.
5. `src/node/services/turnRequestBuilder.ts` (~L2247): wire the closure:
   ```ts
   listAvailableModels: () => {
     const appConfig = this.dependencies.config.loadConfigOrDefault();
     const policy = this.dependencies.policyService;
     return listAvailableModels(
       {
         providersConfig: this.dependencies.providerService.getConfig(),
hiddenModels: appConfig.hiddenModels ?? [...DEFAULT_HIDDEN_MODELS],
routePriority: appConfig.routePriority ?? [...DEFAULT_ROUTE_PRIORITY],
         routeOverrides: appConfig.routeOverrides ?? {},
effectivePolicy: policy?.isEnforced() ? policy.getEffectivePolicy() :
null,
       },
(raw, reason) => log.debug(`[models_list] skipped ${raw}: ${reason}`)
     );
   },
   ```
6. `src/node/services/tools/models_list.ts` (new):
`createModelsListTool: ToolFactory` — `config.listAvailableModels ==
null` → `{ success: false, error: "Model catalog unavailable in this
context" }`; else call it inside try/catch; on throw `log.error(...)`
and return the fixed `{ success: false, error: "Failed to compute the
model catalog" }`; otherwise `{ success: true, models }` (an initialized
configuration with nothing selectable yields `models: []`, still
`success: true`).
7. `src/browser/features/Tools/Shared/ToolPrimitives.tsx`: `models_list:
<lucide icon already imported there, e.g. Cpu/Boxes>` in
`TOOL_NAME_TO_ICON`.
8. `make fmt` → commit the regenerated `docs/hooks/tools.mdx` row.

Tests added in Phase B:
- `src/common/utils/ai/selectableModels.test.ts` (extend):
`listAvailableModels` with a fixture `ProvidersConfigMap` (anthropic
configured, openai unconfigured, openrouter configured with custom entry
`anthropic/claude-sonnet-5`, custom `fixture` provider with
`["fixture-echo", "fixture-echo " (trailing space), { id:
"fixture-mapped", mappedToModel: "anthropic:claude-opus-4-7" },
"bad:id:colon", "lonely "]`, one hidden built-in) → explicit expected
entries: aliases only on the built-in they normalize to (`sonnet` →
`anthropic:claude-sonnet-5`), `[]` on custom models; malformed
`fixture:bad:id:colon` omitted (`onSkipped` called with `"malformed"`)
while the rest survive; **changed-identity regression** (an input
demonstrably changed by `normalizeModelInput`): `fixture:fixture-echo `
normalizes to `fixture:fixture-echo`, which is independently selectable
→ one deduped entry; `fixture:lonely ` normalizes to `fixture:lonely`,
which is not in the selectable set → omitted with
`"unchecked_identity"`; **explicit gateway selection preserved**:
`openrouter:anthropic/claude-sonnet-5` is emitted unchanged (not
collapsed into the canonical ID) — assert the actual
`normalizeModelInput` output rather than assuming; `fixture-mapped` gets
the mapped target's thinking policy; a model whose policy excludes `off`
vs one on the default policy; every entry passes
`AvailableModelSchema.strict().parse`. Do not change task normalization
to make a fixture pass; adjust expectations to observed behavior.
- `src/node/services/tools/models_list.test.ts`: (a) no closure → `{
success: false, error: "Model catalog unavailable in this context" }`;
(b) closure result returned and `ModelsListToolResultSchema.parse`
succeeds; (c) closure throws an error whose message contains a fixture
secret → result is the fixed failure message and does not contain the
secret; (d) closure returns `[]` → `{ success: true, models: [] }`; (e)
**task-input compatibility**: for every returned entry,
`parseTaskAiOverrides({ model, thinking })` accepts `model`, each
`alias`, and each `thinkingLevel`; (f) serialized success result
contains none of the fixture's secret strings (`apiKey`, `baseUrl`,
header values).
- **Production closure test** in `src/node/services/turnRequestBuilder`
coverage (pattern: `aiService.test.ts` `getToolsForModelSpy` /
`workspaceService.multiProject.test.ts` `capturedToolConfig`): capture
the real `ToolConfiguration`, assert `listAvailableModels` is present,
call it with fake `providerService.getConfig()`,
`config.loadConfigOrDefault()`, and `policyService` whose return values
are **mutated between invocations** (provider configured→disabled,
`routePriority` gateway added, a model added to `hiddenModels`, policy
enforced→denying) and assert each subsequent call returns the explicit
expected list — the tool is instantiated once, the closure recomputes.
- **Task-handler forwarding test** in
`src/node/services/tools/task.test.ts` (existing `createTaskTool`
harness with stubbed `taskService` / `workspaceTurnManager`): call the
public `task` handler with a `models_list`-shaped entry (`model` =
canonical ID, then its alias) and a listed `thinking` level for `kind:
"subagent"` and `kind: "workspace"` (`mode: "new"`), asserting the stubs
receive the normalized `modelString` and resolved `thinkingLevel`. No
new full-stack harness.
- Assembly/gating: in `src/node/services/toolAssembly.test.ts` (existing
`resolveToolPolicyForAgent` pattern) assert `models_list` survives the
exec, plan, and explore policies and the sub-agent hard-deny; extend
`toolDefinitions.test.ts:683–753` if it enumerates exposed base tools.

Gate B: new tests green; `make static-check`; full `make test-unit` (the
repo's required gate — no touched-suite substitute; if the known Bun
1.3.5 Wasm SIGSEGV appears, rerun with the documented invocation-scoped
`BUN_JSC_useWasmIPInt=0` and report it); then the hands-on dogfood
below.

## Files touched (product code)

| File | Change | Net new LoC (moved code excluded) |
|---|---|---|
| `src/common/utils/policy/modelPolicy.ts` (new) +
`src/browser/utils/policyUi.ts` | move 3 helpers, re-export | +3 |
| `src/common/routing/resolve.ts` + `src/browser/hooks/useRouting.ts` |
shared `DEFAULT_ROUTE_PRIORITY` | +1 |
| `src/common/utils/ai/selectableModels.ts` (new) |
`computeSelectableModels` (moved, not counted) + `AvailableModel` +
`listAvailableModels` with normalize/recheck/dedupe/onSkipped (new) |
+60 |
| `src/browser/hooks/useModelsFromSettings.ts` | replace moved pipeline
with one call (moved deletions not counted) | +6 |
| `src/common/utils/tools/toolDefinitions.ts` | schemas, tool entry,
baseTools, task hint | +32 |
| `src/common/types/tools.ts` | types | +6 |
| `src/common/utils/tools/tools.ts` |
`ToolConfiguration.listAvailableModels`, registration | +6 |
| `src/node/services/turnRequestBuilder.ts` | closure + skip logging |
+18 |
| `src/node/services/tools/models_list.ts` (new) | tool with fixed error
messages | +35 |
| `src/browser/features/Tools/Shared/ToolPrimitives.tsx` | icon | +1 |
| `docs/hooks/tools.mdx` | generated | n/a |

**Estimated net new product LoC: ≈ +170 (range 150–200)**, excluding
moved code, tests (≈ +350), and generated docs.

## Acceptance criteria

1. `models_list` is exposed through the real assembly path for exec,
plan, explore, and sub-agents (test + dogfood), and returns `{ success:
true, models }` with `model`, `aliases`, `thinkingLevels` per entry; `{
success: false, error }` where no catalog closure exists.
2. For the same backend state, the tool's `model` set equals the
**valid, normalized, selectable** entries of the settled composer picker
(`computeSelectableModels` is the shared implementation; malformed
picker entries and changed identities that are not themselves selectable
are the only omissions; dogfood compares the rendered picker with the
tool output).
3. **Task-input compatibility**: every returned `model`, alias, and
thinking level is accepted by `parseTaskAiOverrides` (unit test), and
launching `task` with a returned model succeeds against controlled
loopback fixtures for both `kind: "subagent"` and `kind: "workspace"`
(dogfood). Real-provider acceptance is not claimed.
4. Hidden models, direct-only models of unconfigured/disabled providers,
authoritative-catalog-removed gateway models, and policy-denied models
are absent; a gateway-routable model appears when its gateway is
configured. Malformed persisted custom IDs are skipped, not fatal.
5. Neither success nor failure output contains credentials or provider
configuration: strict schemas bound the shape, string fields carry only
model IDs / aliases / level names, and failure messages are fixed
constants (tests c, f).
6. `useModelsFromSettings.test.ts` results are identical before/after
Phase A; `make static-check` and full `make test-unit` pass on the final
commit; `docs/hooks/tools.mdx` gained exactly the `models_list` row.

## Dogfooding (deterministic, loopback fixtures by default)

No real provider key is used unless the user explicitly authorizes it.
Everything runs against an owned temp `XUM_ROOT` and a loopback
OpenAI-compatible SSE fixture; every request must reach only that
fixture.

1. **Isolated environment.** Read the `dev-server-sandbox` skill. Start
the dev server under an **allowlisted** environment (`env -i PATH=…
HOME=… XUM_ROOT=<temp> …` plus only the variables the sandbox script
needs) so no inherited `ANTHROPIC_*`/`OPENAI_*`/Coder gateway credential
can make a real provider "configured". Seed `providers.jsonc` with two
custom `openai-compatible` providers: `fixture` (`baseUrl` = loopback
fixture, `apiKey: "fixture"`, `models: ["fixture-driver",
"fixture-echo", "fixture-hidden"]`) and `fixture-off` (same `baseUrl`,
`isEnabled: false`, `models: ["never"]`); `config.json` `hiddenModels:
["fixture:fixture-hidden"]`, default model `fixture:fixture-driver`.
**Before sending anything**, verify the effective snapshot with `xum api
providers get-config` (only `fixture` is `isConfigured && isEnabled`; no
built-in provider configured) and `xum api config get-config`
(routePriority `["direct"]`). Keep fixture scripts, logs, and artifacts
**outside** the checkout.
2. **Fixture script (state-driven, not history-substring).** The fixture
dispatches on the *latest* message of each request (role +
tool-call/tool-result state) and on the requested model:
- parent, latest user message = discovery prompt → tool call
`models_list {}`; latest = tool result of `models_list` → text turn
echoing the result JSON, then stop.
- parent, latest user message = spawn prompt → tool call `task {
agentId: "explore", title: "Echo", prompt: "reply ok", model:
"fixture:fixture-echo", thinking: "low" }` (second variant: `kind:
"workspace", workspace: { mode: "new" }`); latest = tool result of
`task` → text summary, stop.
- child / workspace turn (model `fixture-echo`, latest user message
contains "reply ok") → text `ok` (sub-agents auto-report after the first
turn; the workspace turn completes on the text).
- parent, latest user message = negative prompt → tool call `task { …,
model: "fixture-off:never" }`; latest = tool result (error) → text
summary, stop.
- anything else → a fixed text `unexpected request`, logged with the
request body, so a loop is visible immediately.
The fixture logs every request (`model`, latest message kind) to a file
that is checked at the end: only expected shapes, no repeats.
3. Open the app with `agent-browser` under an owned `--session <name>`
(never `close --all`); `agent-browser record start` (ffmpeg present)
before the steps below.
4. Open the composer model picker (settled, no search text);
`agent-browser snapshot -i` to capture the rendered entries; screenshot
→ `picker.png`. Expected: `fixture:fixture-driver`,
`fixture:fixture-echo`; no built-ins, no `fixture-off` model, no hidden
model.
5. Send the discovery prompt (*"Call the models_list tool and paste the
JSON result verbatim, then stop."*). Expand the tool card; snapshot +
screenshot → `models_list.png`. Expected set equals step 4; `aliases:
[]`; `thinkingLevels` = default policy.
6. Send the spawn prompt (sub-agent) → child starts, replies `ok`,
reports → `task_subagent.png`. Send the workspace-kind spawn prompt →
new workspace turn completes with `ok` → `task_workspace.png`. Send the
negative prompt → `task` fails at launch with `provider_disabled` for
`fixture-off:never` → `task_negative.png` (a model the catalog correctly
omitted; no real provider is involved).
7. **390 px check (required)**: resize the viewport to 390×844, close
the sidebar drawer (tap backdrop), expand the tool card →
`models_list_mobile.png`; assert via `agent-browser eval` that
`document.documentElement.scrollWidth <= window.innerWidth` (long IDs
wrap/truncate, no right-edge overflow).
8. Stop the recording → `dogfood.webm`; inspect frames with `ffmpeg -vf
fps=1` to confirm steps 4–7 are visible; confirm from the fixture log
that every request hit the fixture and matched an expected shape;
`attach_file` the PNGs and video in the report.
9. Cleanup: stop the dev server and fixture processes you started,
`agent-browser close --session <name>`, confirm no owned PIDs/ports
remain, delete the temp `XUM_ROOT` only if it holds nothing needed for
the report.

## Risks and accepted trade-offs

- **Refactor of a hot UI path** (Phase A): mitigated by moving code
without logic changes, baseline-vs-after runs of the 40 hook tests, and
explicit-expectation tests for the pure function.
- **Hidden-models source**: UI reads persisted browser state seeded from
`config.json`; backend reads `config.json`. Both are kept in sync by
`updateModelPreferences`; a transient mismatch right after hide/unhide
is acceptable.
- **Advisory list**: `task.model` still accepts any well-formed string;
validating `task.model` against the catalog is out of scope.
- **Config read side effects**: `loadConfigOrDefault()` may persist
one-time migrations, but that is existing behavior of every backend
read; the tool introduces no new writes.
- **Real-launch coverage is hands-on, not a new automated real-stack
suite**: adding a fixture-backed end-to-end test harness is heavier than
the feature; unit tests cover input compatibility, task-handler
forwarding for both kinds, and production wiring; dogfood covers actual
launches.
- **Changed-identity omission**: a picker entry whose normalized form
differs and is not itself selectable (e.g. a whitespace-padded custom ID
with no clean counterpart) is omitted rather than emitted in either
form. Conservative reading of parity; the picker still shows the raw
entry. Acceptable because `task` would rewrite the raw form anyway and
such entries indicate a configuration typo.
- **Token cost**: one more tool schema in every request; kept minimal.

## Out of scope / follow-ups

- `defaultModel`, display names, context windows, or cost data in the
output; a dedicated tool card or CLI formatter.
- Rejecting unavailable `task.model` values at task creation.
- Exposing the catalog over oRPC/CLI (`xum api models list`).

</details>

---

_Generated with [`xum`](https://github.com/coder/xum) • Model:
`coder:openai/gpt-6-astra` • Thinking: `high` • Cost: `$171.28`_

<!-- mux-attribution: model=coder:openai/gpt-6-astra thinking=high
costs=171.28 -->

---------

Signed-off-by: Thomas Kosiewski <tk@coder.com>
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Sep 24, 2026
… correct their Codex context cap (coder#4259)

## Summary

Follow-up to coder#4350 (landed as `7358fa80`), which makes a *requested*
`reasoning_effort: "none"` reach the wire for GPT-6 Sol/Luna. This PR
closes the three gaps left after coder#4348 and coder#4350: headless tool loops
that request no effort at all, Coder-scoped aliases on `openai-compat`
instances, and the assumed 372K Codex OAuth context cap (272K in the
pinned Codex catalog).

## Background

- Sol/Luna Chat Completions accepts function calling only with
`reasoning_effort: "none"`; an omitted effort defaults to `medium`.
`buildProviderOptions` clamps agent turns (coder#4348) and coder#4350 makes that
clamp serialize, but Dream consolidation, memory harvest, refine and
sidebar-status generation call `streamText` with tools and **no provider
options**, so nothing requests `none` and their tool calls are rejected
on Chat Completions routes (direct `wireFormat: chatCompletions`, Coder
`openai-compat` instances). Raised by Codex on the earlier revision of
this PR; the Coder alias case is its round-6 finding.
- Codex catalog pin: [openai/codex@04fc75ad
`codex-rs/models-manager/models.json`](https://github.com/openai/codex/blob/04fc75adbe67a612a1cb0fc469533f24b24fa499/codex-rs/models-manager/models.json)
— `gpt-6-sol` / `gpt-6-luna` `context_window: 272000`. `main` copies
Astra's 372K with an "assumed" comment.

## Implementation

- `clampGpt6ChatCompletionsToolReasoning` (providerModelFactory.ts): a
`transformParams` middleware that sets `reasoningEffort: "none"` when
the request carries tools **and no effort was requested**, for Sol/Luna
capability identities on Chat Completions. Explicit caller efforts are
preserved (this keeps coder#4350's "requested `high` stays `high`" contract
intact). It wraps *outside* `createOpenAIModelWithPreservedOptions` so
the injected `none` is seen by that wrapper's body patch instead of
being stripped by the SDK.
- Applied on the direct OpenAI path (Chat Completions wire only) and the
Coder gateway Chat path, where the capability identity is
`resolveModelForMetadata(coder:<instance>/<model>)` so scoped "Treat as"
aliases are clamped too.
- `CODEX_OAUTH_CONTEXT_WINDOW_OVERRIDES`: `gpt-6-sol` / `gpt-6-luna`
372K → 272K with the pin cited inline; Astra and the GPT-5.6 family are
unchanged (coder#4347).

Dropped from earlier revisions:
catalog/aliases/pricing/labels/migration/docs (landed in coder#4348), the
`@ai-sdk/openai` Bun patch (Bun `patchedDependencies` never reach npm
installs of `@coder/xum`; coder#4350's runtime rewrite is the right fix), and
the `hasTools` plumbing through `buildProviderOptions`.

## Validation

Repo-pinned Bun 1.3.5, on `main` + this commit with a freshly
reinstalled, unpatched `@ai-sdk/openai` (so a green body assertion
proves the composition with coder#4350's wrapper).

- `providerModelFactory.test.ts` › "clamps %s tool requests … at the
model boundary" (Sol raw id, mapped `team-luna`, Astra control,
Responses control) and › "clamps %s tool requests through a Coder
openai-compat instance" (`coder:chat-proxy/team-luna` mapped,
`coder:chat-proxy/gpt-6-sol` raw, `team-astra` control). Red proof: with
the clamp disabled, exactly the four Sol/Luna Chat Completions cases
fail (`Expected "none" / Received undefined`); coder#4350's 11 wire cases are
unaffected. Reverting the Coder capability argument to the raw
`originModelId` fails exactly the mapped `team-luna` case.
- `codexOAuth.test.ts` asserts Astra (372K) and Sol/Luna (272K)
separately. `contextLimit`, `tokenMeterUtils`, `thinking`,
`turnRequestBuilder`, `providerOptions` suites pass. `make static-check`
green.

## Risks

- The middleware fires only for Sol/Luna capability identities, Chat
Completions wire, a non-empty `tools` array, and no requested effort.
Responses, tool-free requests, explicit efforts, Astra, GPT-5.6 and
non-OpenAI providers are untouched.
- Lower Codex OAuth cap (272K) starts limit-driven compaction earlier
for OAuth-routed Sol/Luna; API-key routes keep the public limits.

## Readiness and follow-ups

- Originally stacked on coder#4350; rebased onto `main` after coder#4350 landed
(`7358fa80`). The residual diff is unchanged (+183/−21).
- `main` unblock for `Test / Unit` (release commit left
`packages/mux-compat` at 0.29.0): coder#4354 standalone, also folded into
coder#4350; whichever becomes empty is closed.
- Round-6 security finding (budget accounting ignores priority-tier
pricing) is pre-existing/generic → coder#4352. coder#4347 tracks GPT-5.6 Sol
promotional pricing and Astra's Codex cap.

<details>
<summary>Preserved review and delivery record — fifteen Codex executions
including this head, one code-advisor pass</summary>

| Round | Reviewed head (before rebase / re-scope) | Executions |
Findings / disposition |
| --- | --- | --- | --- |
| 1 | `01ca053` | Codex code + security | Hidden-model compaction
exposure and zero pricing bypassing CLI budget; fixed in `252bba9` /
`cb5ac4b`, replied and resolved |
| 2 | `cb5ac4b` | Codex code + security | Migration-marker flip and
hiding prior opt-ins; fixed in `64169a3`, replied and resolved |
| 3 | `64169a3` | Codex code + security | Clean completion on both loops
|
| 4, renewed authorization | `40e4b2a6` | Automatic Codex code +
security on undraft | Security clean; code P2 Chat Completions
default-effort issue → fixed in `ce20ecd0`, replied and resolved |
| 5, explicit "get it merged" authorization | `bfc0906f` (superseding
`ce20ecd0`) | Automatic Codex code review on undraft | P2: headless tool
loops bypass provider options → fixed in `a4c7c156`, replied and
resolved |
| 6 | `a4c7c156` | Explicit `@codex review` (code + security) | Code P2:
Coder aliases not resolved before clamping → fixed, replied and
resolved. Security P2: priority-tier budget accounting →
pre-existing/generic, coder#4352, replied and resolved |
| 7 | `2c4803c0` (residual on `main` after coder#4348) | Explicit `@codex
review` (code + security) | Both clean ("Didn't find any major issues" /
"No security issues") |
| 8 | `c9a0f572` (stacked on coder#4350; rebased onto `main` without changes)
| Explicit `@codex review` (code + security) | Both clean ("Didn't find
any major issues" 21:57Z / "No security issues" 21:59Z, 👍) |
| 9 | `76bff932` (rebased onto `main` after coder#4350 landed) | Explicit
`@codex review` (code + security) | Both clean; merge blocked by the
repo-wide `compaction1MRetry` content-filter failure (coder#4398), fixed by
coder#4401 |
| 10 | `bde7c0fe` (rebased onto `main` after coder#4401 landed; diff
unchanged) | Explicit `@codex review` (code + security) | Security
clean. Code P3: stale Astra comment claimed the catalog couples Astra
and Sol → fixed in `7deeea6e`, replied and resolved |
| 11 | this head (`7deeea6e`, comment-only fix) | Explicit `@codex
review` (code + security) | Pending |
| Advisory | — | One independent advisor | Recommended scope reduction
and accuracy fixes before official release. Scope reduction realised by
coder#4348 and coder#4350 landing first; this PR is the residual. |

Post-cap work, retained in order: `f04892f`→`fcf5d0f` budget-comment
accuracy; `782cca6`→`04bad77` maintainer-authorized provisional
estimate; `d24d776` rebase integration; `870259a` official Sol launch +
authorized Luna expansion; `40e4b2a6` verified Codex defaults, labels,
route/pricing regressions, docs; `ce20ecd0` tool-aware `none` clamp +
SDK patch; `bfc0906f` Nix hash refresh; `a4c7c156` model-boundary clamp;
`2c4803c0` residual on `main` (SDK patch + clamp + Coder alias + Codex
cap). Other rebase mappings: `01ca053`→`c47ba9c`, `252bba9`→`5363ca0`,
`cb5ac4b`→`a866d96`, `64169a3`→`c0d5def`. The cap did not reset on
rebase, handoff, scope expansion or re-scope. Backups of superseded
heads: `a4c7c15635f3a865baa4b937618da17459b48209`,
`2c4803c0c9481924ad682f969542a7a6db71e9e7`.

Historical stop reason: implementation and validation were complete for
both models, but the final head lacked review coverage under the
exhausted cap. The later direct conditional merge instruction reopened
delivery for one automatic final-head pair; its findings were then fixed
under the user's explicit "get it merged" authorization through the
normal gated merge path. When coder#4348 and then coder#4350 landed the same work
first, this PR was reduced to the residual fixes above rather than
closed.

</details>

---

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking:
`xhigh` • Cost: `$281.40`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=xhigh
costs=281.40 -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants