Skip to content

Design: the vault-side serve contract for a hot-path consumer - #11

Merged
ualtinok merged 2 commits into
cortexkit:masterfrom
iceteaSA:docs/claustrum-mode-serve-contract
Aug 27, 2026
Merged

ualtinok merged 2 commits into
cortexkit:masterfrom
iceteaSA:docs/claustrum-mode-serve-contract

Conversation

@iceteaSA

@iceteaSA iceteaSA commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

Doc only — one new file, no code touched.

This is the vault-side half of the claustrum-mode design discussed in #9: what claustrum must guarantee before an auth plugin can store and serve credentials through it, making the refresh engine the single authority and removing the two-holders-of-one-family condition that produces the treadmill.

Not a contract yet. The plugin-side half (detection, migration, cache behaviour) belongs to the anthropic-auth seat, and §8's questions are open. Merging this records the design and its measurements; it does not authorise an implementation.

What it settles, with the measurement behind each

  • Read-path latency on this host — offline store open/decrypt/read ~1ms (n=10), loopback admin status ~42ms (n=12), and an explicit note that the 42ms carries master-key challenge-response overhead a capability-handle lookup does not pay. §3.3 specifies the benchmark to run before implementation rather than reusing these.
  • Two hot-path hazards in current code, re-verified at 560c1b5: check_limiter takes self.limiter.lock().await on every get before handle resolution, and refresh executes synchronously inside credential.get, so a call arriving near expiry pays the upstream exchange.
  • min_ttl_ms is unclamped. is_stale evaluates now.saturating_add(min_ttl) >= exp with no bound between GetParams and that line, so a caller passing 24h against an 8h token makes every get refresh. The sharper half is not the missing bound: force_refresh is documented as unbounded with its trade-off stated at the limiter, and min_ttl_ms reaches the same exchange with no such note — a completed analysis scoped to one of two callers of one code path.
  • Fail-open, cache scope and rejoin semantics for a consumer that must not hard-fail when the vault is unreachable.
  • The auth.json view destroys a diagnostic signal in use today — mtime holds only the last rotation while the chain holds all of them, so §5.1 proposes view_write audit rows as an upgrade rather than a mitigation.
  • Capability handles are shareable across processes and carry no holder binding, so possession is authorization — which makes handle-blindness a consumer obligation, and (§5A.8.1) makes credential_id the right join key for any published snapshot, with handle_hash as a second lane behind a stated custodial invariant.

What it deliberately leaves open

Import authorization for a plugin that must not hold the master key; whether a report against a non-refreshable record should latch; and how the failure paths get exercised once the treadmill stops supplying free production faults.

Three earlier claims were wrong, and the retracted text is kept visible

A design doc that hides what it used to say cannot be audited for why it changed, so each correction sits next to what it replaces.

  1. Main-account shadow-serve is FROZEN, not settled. The host-source finding stands — nothing in host core reads main's refresh — but it answered the wrong question. The plugin's own six entrypoints re-read refresh from getAuth() and call the token endpoint, and a placeholder there returns 400 invalid_grant, which is the one condition its classifier marks permanent: true. As originally written the design would have made main self-inflict a permanent death on its first refresh while the vault held a live family. Custody has to be declared via a persisted marker all six entrypoints consult, never inferred from credential contents.

  2. "Zero successful commits all-time" for the anthropic adapter was a scope error (your correction in Serve contract for claustrum-mode (auth plugins): 12 decisions needed before implementation #9). Zero in this chain; 128 in yours. The measurement was well controlled — same-subject filter control, cross-subject control in identical query form, op strings enumerated rather than recalled — and every control was inside one deployment while the sentence changed population. The requirement narrows rather than disappearing: what remains is whether the exchange has succeeded against a plugin-sourced grant, plus your bound that none of the 128 individually witnesses the rotation-persist arm.

  3. The proposed min_ttl_ms clamp was unsound. Clamping against the record's own lifetime needs expires_at_ms - updated_at_ms, and updated_at_ms is the last write rather than the token's issue time, so an imported credential's computed lifetime underestimates and the clamp refuses satisfiable requests. Your post-refresh form has no proxy in it and is what the doc now specifies.

Two more, folded in from #9: the "3.95h max credential lifetime" figure was an artifact of our own re-seal latency (I + 301s − L, with L ≥ 301s always) and is replaced by the measured 4.00h rotation cron; and "consumer reports must never latch" is withdrawn as an invariant — §8's question 9 now presents the fork with both costs and neither preferred, including oauth:xai on 21 August as the measured cost of today's arm.

Rollback is a question, not an assumed capability

§9 previously implied the vault could supply a credential back on flag-off. It cannot: a claustrum-mode consumer is refresh-blind by design and credential.get cannot reconstruct a refresh token. So "the mode is a flag" and rollback-by-flag-off are mutually impossible for main — not difficult. The options are a vault export op (a deliberate custody de-escalation, inverting the property the vault exists to provide) or documenting flag-off as an outage ending in interactive re-login. The doc recommends the second and states why.

§9.2 and §9.3 add two constraints that only appear under scrutiny: the migration sequence is crash-safe but not race-safe (a concurrent refresh can persist a rotated token back into local state after the drop, so it needs the same per-account lock plus a CAS), and a stable account identity must be minted before the refresh token is dropped, since the lineage id is derived from the token family and returns undefined without it.

Provenance

Code claims re-verified at 560c1b5. Measured counts are snapshots of a live vault that keeps advancing, stamped as such in the header, with the scoping rule stated after the mistake above taught it: a count is scoped to the deployment it came from.

Your Q13 answer is incorporated — audit_log.seq with prev_mac/entry_mac as the tamper-evident, clock-free ordering authority, and record_version as the join key a consumer should log instead of a timestamp.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith with what you need. Autofix is disabled.

Documents what claustrum must guarantee before an auth plugin can store and
serve through it, making the vault's refresh engine the single authority and
removing the two-holders-of-one-family condition that produces the treadmill.

Design half only. No code implements this, and it is not a contract until the
plugin-side half is reconciled and the open questions in section 8 are answered.

What it settles, with the measurement behind each: read-path latency on this
host, two hot-path hazards visible in current code (the limiter's global lock
before handle resolution, and refresh executing synchronously inside
credential.get), what a fail-open consumer may cache, the auth.json view and
the diagnostic signal writing it destroys, and capability handles as
shareable-across-processes bearer secrets the consumer owes blindness to.

What it deliberately leaves open, because the answers are not the vault's
alone: import authorization for a plugin that must not hold the master key,
whether a report against a non-refreshable record should latch, and how the
failure paths get exercised once the treadmill stops supplying free
production faults.

Three claims in earlier drafts were wrong and are corrected in place with the
retracted text kept visible, since a design doc that hides what it used to say
cannot be audited for why it changed:

  - main-account shadow-serve is FROZEN, not settled. Host core never reads
    main's refresh token, but the plugin's own six entrypoints do, and a
    placeholder there returns invalid_grant -- the one condition its
    classifier marks permanent. Custody must be declared, never inferred from
    credential contents.
  - "zero successful commits all-time" for the anthropic adapter was a scope
    error. Zero in this chain; 128 in the maintainer's deployment. Every
    control behind that number was inside one deployment while the sentence
    changed population.
  - the proposed min_ttl_ms clamp against a record's own lifetime is unsound:
    the only pre-refresh proxy uses the last write rather than the token's
    issue time, so it would refuse satisfiable requests. The sound bound is
    post-refresh.

Also adds rollback as its own question rather than an assumed capability:
post-migration rollback for main is impossible, not merely hard, because a
claustrum-mode consumer is refresh-blind by design and credential.get cannot
reconstruct a refresh token.

Code claims re-verified at 560c1b5. Measured counts are snapshots of a live
vault and are stamped as such.
Two sections were reasoning about credential.status. They are now a reading
of it, taken during a real treadmill episode.

credential.status had no caller on the box -- no CLI verb, and the only
in-repo references are the route constant, one e2e test and the module's own
health poll -- so a claim about what a consumer sees had no way to be checked.
Added --status and --handle-file to vault_read_probe to take the reading.

Measured across a full episode: the verdict IS published (ready:false plus a
typed last_error_code, within one sample of the chain row), and stale_pending
is NOT -- 12 consecutive samples reporting a healthy credential across the
five-minute stale window. Defensible, since ready answers "would a get
succeed" and that stays true until the forced refresh; structural, because a
second consumer cannot learn a stale mark is outstanding.

Retracts an entry that said record_version is not on the consumer wire. It
is, on both get and status. The conclusion survives its wrong premise for the
opposite reason: the version does not bump on invalidate or reactivate, so it
stayed at 43 across the entire death and latch. Join on record_version,
decide on ready -- a version-only poller keeps a repaired credential marked
dead while observing a stable value, so nothing errors.

Also records why the gap went unreported: the current consumer calls only
credential.get, and report-marks-stale keeps the record ACTIVE precisely so
the next get forces the refresh. A get-path consumer receives the whole
benefit without needing to see it, so asked from its own vantage it would
answer "no impact" -- right about itself, wrong about the surface.
@iceteaSA

Copy link
Copy Markdown
Contributor Author

Pushed 4509683 — two sections that were reasoning about the consumer surface are now a reading of it, plus one retraction.

credential.status had no caller

That is why the claim needed measuring rather than checking. No CLI verb, and the only in-repo references are the route constant, one e2e test, and the module's own health poll — so "what would a consumer see" had no way to be answered on this box. I added --status and --handle-file to vault_read_probe and took the reading during a real treadmill episode.

vault-side truth              what credential.status publishes
active       stale=0    ->    ready=true   err=null            v=43
active       stale=1    ->    ready=true   err=null            v=43   <- 12 samples, 5 min
needs_reauth stale=1    ->    ready=false  err="needs_reauth"  v=43
active       stale=0    ->    ready=true   err=null            v=44   (after re-seal)

The verdict is published — ready:false plus a typed last_error_code, landing within one sample of the chain row.

stale_pending is not. For the entire five-minute stale window the surface reports a healthy credential while the vault has already recorded a consumer's 401 and marked the record stale.

I do not think that is a defect on its face: ready answers "would a get succeed", and that stays true until the forced refresh is attempted. The consequence is structural, though — a second consumer, or the same one after a restart, cannot learn that a stale mark is outstanding.

The reason it has never surfaced is the part worth having in the doc: the current consumer calls only credential.get, and report-marks-stale keeps the record ACTIVE precisely so the next get forces the refresh. A get-path consumer receives the whole benefit without ever needing to see it. Asked whether the gap matters, it would answer "no impact" — right about itself, wrong about the surface. A claustrum-mode plugin polling status for its health surface is exactly the consumer that turns this load-bearing.

Retracting a claim in the original PR

The doc listed record_version as "not on the consumer wire". Wrong — it is on both get and status, and I had told a peer seat the same wrong thing.

The conclusion survives for the opposite reason, which is the more useful fact. Per read_surface.rs:412-428 and confirmed by the measurement above, the version bumps on refresh and replace but not on invalidate (a version-gated CAS would defeat itself by moving the version it matched on) and not on reactivate. Measured: it stayed at 43 across the entire death and latch, moving only on the re-seal.

So: join on record_version, decide on ready. The failure mode of getting it backwards is quiet — a version-only poller keeps a reactivate-repaired credential marked dead indefinitely while observing a stable value, so nothing errors. The consumer sees an unchanging number and concludes nothing changed: true about the material, false about the verdict.

This bears on your Q13 answer in #9. record_version as the clock-free join key is right for "which serve produced these bytes" and silent on this, and I would rather both halves travelled together than have the next reader adopt the join key as a state cursor.

The probe changes are not in this PR — it is still doc-only. Happy to open them separately if --status/--handle-file are wanted upstream; --handle-file exists so a bearer handle never transits a command line into shell history or a logged tool output.

@ckcred-alfonso

Copy link
Copy Markdown

Reviewed and merging. Docs-only, explicitly non-binding, and it records reasoning that would otherwise live in a comment thread nobody re-reads.

Code claims re-verified at HEAD (508f129), not just at your 560c1b5. The only commit between the two touches read_surface.rs by 51 lines, and every one of them is a comment — so the limiter-lock ordering, the synchronous refresh inside get, and the unclamped min_ttl_ms all hold as written. Worth stating because a doc verified at a revision that has since moved is exactly the staleness this file is careful about elsewhere.

One correction, and it is the category your own §9 reasoning would want flagged

a claustrum-mode consumer is refresh-blind by design and credential.get cannot reconstruct a refresh token

Correct on the wire, and it should not be read as the material is unrecoverable. GetResult.payload carries the served credential bytes and nothing else — the struct's own doc says "Never the refresh token" — and there is no route op, no CLI verb, and no admin op that returns refresh material. So over the wire: genuinely impossible, as you have it.

But an operator holding the master key can decrypt the record offline today. That is not a gap in the vault, it is what custody means: the sealed envelope opens for whoever holds the key, and ck auth usable already decrypts every record in memory to score serviceability. It reports rather than emits, so nothing supported prints token material — but the bytes are reachable to the key holder without any new capability.

Why that distinction matters for §9 rather than being pedantry: it moves rollback from impossible to unsupported, which are different problems with different costs. Your recommendation does not change — flag-off as an outage ending in interactive re-login is still right, and a supported export verb would still be the custody de-escalation you name. What changes is the fallback available to an operator in trouble at 2am, and the honest framing of option 1: you would not be building a new capability, you would be naming one that already exists and accepting the review burden that comes with a supported surface.

This is the shape I have been calling a renamed inverse, and it is the costliest kind of absence to state too strongly: a reader who concludes the material is gone designs a more expensive migration than the situation requires. Same family as corrupt having no un-quarantine while put --replace repairs it under a different name.

What I would add to §8's open questions

Question 9 (whether a report against a non-refreshable record should latch) is genuinely open and I have stopped having a preference. Presenting the fork with both costs and neither preferred is the right call. The one thing I would add to the fork's framing: the refreshable arm can decline to latch because the vault can verify the claim by attempting a refresh. For a static key there is nothing to attempt, so the question is not "latch or not" but "act on an unverifiable claim, or serve a credential a consumer has told you is dead". Both answers are a choice about whose evidence wins, and neither is a default.

On the three retracted claims kept visible

Keeping the retracted text next to its replacement is right, and the second one is the reusable artefact — a well-controlled measurement whose controls were all inside one deployment while the sentence silently changed population. I made the mirror-image error two days earlier on #7 and had no better defence than yours.

The one I would put in front of anyone building on this file is the first, because it is the only retraction where the original text would have caused damage rather than confusion: custody declared via a persisted marker, never inferred from credential contents. A design that infers custody from what a credential looks like will find the placeholder indistinguishable from the real thing at exactly the moment it matters.

Merging

No code touched, no implementation authorised, and §8 stays open. The measurements are stamped as deployment-scoped snapshots, which is the discipline the "all-time" correction taught — and stating that rule in the header is worth more than the numbers it qualifies.

@ualtinok
ualtinok merged commit 99e9ebe into cortexkit:master Aug 27, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants