Skip to content

feat(web): add an unauthenticated GET /healthz probe to the backend - #2585

Merged
cliffhall merged 3 commits into
v2/mainfrom
v2/feat/2438-web-healthz
Oct 5, 2026
Merged

cliffhall merged 3 commits into
v2/mainfrom
v2/feat/2438-web-healthz

Conversation

@cliffhall

Copy link
Copy Markdown
Member

Closes #2438

What

The web backend now answers GET /healthz (and HEAD) with 200, a fixed {"status":"ok"} body and Cache-Control: no-store. Orchestrators (Docker, Kubernetes, process managers) can use it to check the backend is up without going through a real proxy or connect flow.

  • clients/web/server/health.ts (new) holds the path, the body, the headers, a request matcher and the response builder. It is unit tested and gated at 100% on all four dimensions.
  • clients/web/server/server.ts (prod): app.get(HEALTH_PATH, …) is registered before the static and SPA fallbacks, so the probe is never answered with index.html.
  • clients/web/server/vite-hono-plugin.ts (dev): the existing middleware answers the same route, so npm run dev and prod agree. I checked this by hand against a live vite dev server: GET returns 200 with {"status":"ok"}, HEAD returns 200 with no body, and /api/config still returns 401.
  • Docs: a new Health check section and a health.ts entry in clients/web/README.md, plus one sentence in docs/docker.md pointing external orchestrators at the route.

Decisions (made deliberately, per the issue)

  • Unauthenticated, at /healthz rather than /api/health. Every /api/* route sits behind the x-mcp-remote-auth bearer check and the origin allow-list (AGENTS.md, "Web backend auth token"). A probe has no way to learn a token that is generated fresh on each start. A top-level path needs no exception in the auth middleware. /api/health would have needed one carved in.
  • It discloses nothing beyond "up". No version, uptime, connected servers, config or storage state. Anything that can reach the port can read the route, including a DNS-rebinding page, since the origin check only covers /api. A version string helps fingerprinting. "Is it running" is already answerable from GET /, so this adds no new information. The tests also check that the response never carries the API token, which GET / embeds.
  • Liveness and readiness are the same answer. Both backends start listening only after the sandbox listener, the app-origin listener and the API app are built, so any response at all means the backend is ready.
  • The image's own HEALTHCHECK still probes /. Switching scripts/docker-healthcheck.mjs to /healthz would be a natural follow-up. I left it out to keep this diff to the endpoint the issue asks for.

Tests

clients/web/src/test/integration/server/health.test.ts covers:

  • the matcher: methods, query/fragment, trailing slash, sub-paths, a missing url
  • the response builder: GET, HEAD, the default method, the frozen body
  • a real startHonoServer: 200 with only the fixed body and no token leak, HEAD with no body, a disallowed Origin is still answered, /api/config still returns 401

Gate

npm run format was clean. npm run local:gate did not finish as a single run. The machine-wide gate lease was ~18–20 gates deep from parallel agent sessions, and two queued runs reached the lease's 45-minute wait cap without starting. So I ran every gate stage on its own:

  • The fixed-port stages (smoke, smoke:web:firefox, local:storybook) ran under the lease. All green.
  • The rest (local:validate, verify:skills:cli, coverage:{web,cli,tui,launcher}, verify:build-gate, verify:bundle-externals) ran without it. All green except one environmental failure, below.

Two kinds of failure came from the sandbox, not this diff:

  • Sandbox HTTP proxy. It returns 502 on loopback tunnels, which broke transport.test.ts, server-extra-coverage.test.ts and the CLI revocation tests. With the proxy env unset, all of them pass, and so does coverage:cli.
  • Bind mount. secret-store-selection.test.ts > isOnMountPoint > is false when the path itself can't be resolved fails because the worktree is a bind mount in this container, so . really is on a mount point. That is core code this PR does not touch, and CI's runner is not a container.

🤖 Generated with Claude Code

Orchestrators (Docker, Kubernetes, process managers) had no cheap way to
check the web backend is up without exercising a real proxy/connect flow.
Both the prod Hono server and the dev Vite backend now answer GET/HEAD
/healthz with 200, a fixed {"status":"ok"} body and Cache-Control:
no-store.

The route sits outside /api/*, so the bearer-token and origin checks do
not apply (a probe cannot learn a per-start token), and it discloses
nothing beyond "up" -- no version, uptime, servers or config.

Closes #2438

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: cliffhall <cliff@futurescale.com>
@cliffhall cliffhall added the v2 Issues and PRs for v2 label Oct 5, 2026
@cliffhall cliffhall linked an issue Oct 5, 2026 that may be closed by this pull request
2 tasks done
@cliffhall
cliffhall requested a balanced review from Copilot October 5, 2026 08:42

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The new integration test file lacks the repository-required purpose header.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
What changed in this PR

Adds a consistent unauthenticated /healthz endpoint to the web development and production backends.

Changes:

  • Adds shared health-check matching and response helpers.
  • Integrates the endpoint into both backends with integration coverage.
  • Documents orchestrator usage and security behavior.
File Description
clients/​web/​server/​health.ts Defines the health-check contract.
clients/​web/​server/​server.ts Registers the production route.
clients/​web/​server/​vite-hono-plugin.ts Handles health probes during development.
clients/​web/​src/​test/​integration/​server/​health.test.ts Tests helpers and production behavior.
clients/​web/​README.md Documents endpoint semantics and usage.
docs/​docker.md Documents external orchestrator probes.

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread clients/web/src/test/integration/server/health.test.ts
Every file in the repo carries a header explaining its purpose (AGENTS.md,
Project Structure); the new health.test.ts was missing one. Copilot review
on #2585.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: cliffhall <cliff@futurescale.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The distinct development-server response path lacks automated coverage required by the repository’s testing rules.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
Resolved since last review (1)

Comment thread clients/web/server/vite-hono-plugin.ts Outdated
The dev Vite middleware wrote the health response inline, in a file that
is excluded from coverage, so its headers and HEAD handling had no test.
Move it into handleNodeHealthRequest() in health.ts and drive that from a
real node:http server in health.test.ts. health.ts stays at 100% on all
four dimensions. Copilot review on #2585.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: cliffhall <cliff@futurescale.com>
@cliffhall

Copy link
Copy Markdown
Member Author

Copilot review round 2, summary:

  • No automated test for the dev-server response path (medium): fixed in c22de80. The response writing moved into handleNodeHealthRequest in health.ts, which the Vite middleware delegates to and which is tested against a real node:http server. Replied in the thread.
  • No suppressed comments.

Verification for c22de80: clients/web npm run check is green, the health test file passes 16/16, health.ts coverage is 100/100/100/100, and I re-probed a live dev server by hand.

I queued the fixed-port smokes and Storybook under the gate lease again, and the lease gave up after 66 minutes without starting, the third time on this machine. Those stages passed under the lease on the first commit, and since then nothing changed except this adapter. Requesting round 3.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The endpoint is narrowly scoped, consistently implemented, documented, and thoroughly tested.

Review effort: Balanced
Findings: None

Resolved since last review (1)

@cliffhall

Copy link
Copy Markdown
Member Author

Copilot review loop closed. Round 3 (on c22de80) was clean: "Approval recommended", no findings, no inline comments, no suppressed comments. Per the review loop's stop rule, one clean round ends it. Rounds 1 and 2 each had one in-scope finding, both fixed and answered in their threads.

@cliffhall
cliffhall merged commit 49a513c into v2/main Oct 5, 2026
6 checks passed
@cliffhall
cliffhall deleted the v2/feat/2438-web-healthz branch October 5, 2026 14:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

v2 Issues and PRs for v2

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Web client's backend server has no health/readiness endpoint

2 participants