Repository navigation
Design: the vault-side serve contract for a hot-path consumer - #11
Conversation
Documents what claustrum must guarantee before an auth plugin can store and
serve through it, making the vault's refresh engine the single authority and
removing the two-holders-of-one-family condition that produces the treadmill.
Design half only. No code implements this, and it is not a contract until the
plugin-side half is reconciled and the open questions in section 8 are answered.
What it settles, with the measurement behind each: read-path latency on this
host, two hot-path hazards visible in current code (the limiter's global lock
before handle resolution, and refresh executing synchronously inside
credential.get), what a fail-open consumer may cache, the auth.json view and
the diagnostic signal writing it destroys, and capability handles as
shareable-across-processes bearer secrets the consumer owes blindness to.
What it deliberately leaves open, because the answers are not the vault's
alone: import authorization for a plugin that must not hold the master key,
whether a report against a non-refreshable record should latch, and how the
failure paths get exercised once the treadmill stops supplying free
production faults.
Three claims in earlier drafts were wrong and are corrected in place with the
retracted text kept visible, since a design doc that hides what it used to say
cannot be audited for why it changed:
- main-account shadow-serve is FROZEN, not settled. Host core never reads
main's refresh token, but the plugin's own six entrypoints do, and a
placeholder there returns invalid_grant -- the one condition its
classifier marks permanent. Custody must be declared, never inferred from
credential contents.
- "zero successful commits all-time" for the anthropic adapter was a scope
error. Zero in this chain; 128 in the maintainer's deployment. Every
control behind that number was inside one deployment while the sentence
changed population.
- the proposed min_ttl_ms clamp against a record's own lifetime is unsound:
the only pre-refresh proxy uses the last write rather than the token's
issue time, so it would refuse satisfiable requests. The sound bound is
post-refresh.
Also adds rollback as its own question rather than an assumed capability:
post-migration rollback for main is impossible, not merely hard, because a
claustrum-mode consumer is refresh-blind by design and credential.get cannot
reconstruct a refresh token.
Code claims re-verified at 560c1b5. Measured counts are snapshots of a live
vault and are stamped as such.
Two sections were reasoning about credential.status. They are now a reading of it, taken during a real treadmill episode. credential.status had no caller on the box -- no CLI verb, and the only in-repo references are the route constant, one e2e test and the module's own health poll -- so a claim about what a consumer sees had no way to be checked. Added --status and --handle-file to vault_read_probe to take the reading. Measured across a full episode: the verdict IS published (ready:false plus a typed last_error_code, within one sample of the chain row), and stale_pending is NOT -- 12 consecutive samples reporting a healthy credential across the five-minute stale window. Defensible, since ready answers "would a get succeed" and that stays true until the forced refresh; structural, because a second consumer cannot learn a stale mark is outstanding. Retracts an entry that said record_version is not on the consumer wire. It is, on both get and status. The conclusion survives its wrong premise for the opposite reason: the version does not bump on invalidate or reactivate, so it stayed at 43 across the entire death and latch. Join on record_version, decide on ready -- a version-only poller keeps a repaired credential marked dead while observing a stable value, so nothing errors. Also records why the gap went unreported: the current consumer calls only credential.get, and report-marks-stale keeps the record ACTIVE precisely so the next get forces the refresh. A get-path consumer receives the whole benefit without needing to see it, so asked from its own vantage it would answer "no impact" -- right about itself, wrong about the surface.
|
Pushed
|
|
Reviewed and merging. Docs-only, explicitly non-binding, and it records reasoning that would otherwise live in a comment thread nobody re-reads. Code claims re-verified at HEAD ( One correction, and it is the category your own §9 reasoning would want flagged
Correct on the wire, and it should not be read as the material is unrecoverable. But an operator holding the master key can decrypt the record offline today. That is not a gap in the vault, it is what custody means: the sealed envelope opens for whoever holds the key, and Why that distinction matters for §9 rather than being pedantry: it moves rollback from impossible to unsupported, which are different problems with different costs. Your recommendation does not change — flag-off as an outage ending in interactive re-login is still right, and a supported export verb would still be the custody de-escalation you name. What changes is the fallback available to an operator in trouble at 2am, and the honest framing of option 1: you would not be building a new capability, you would be naming one that already exists and accepting the review burden that comes with a supported surface. This is the shape I have been calling a renamed inverse, and it is the costliest kind of absence to state too strongly: a reader who concludes the material is gone designs a more expensive migration than the situation requires. Same family as What I would add to §8's open questionsQuestion 9 (whether a report against a non-refreshable record should latch) is genuinely open and I have stopped having a preference. Presenting the fork with both costs and neither preferred is the right call. The one thing I would add to the fork's framing: the refreshable arm can decline to latch because the vault can verify the claim by attempting a refresh. For a static key there is nothing to attempt, so the question is not "latch or not" but "act on an unverifiable claim, or serve a credential a consumer has told you is dead". Both answers are a choice about whose evidence wins, and neither is a default. On the three retracted claims kept visibleKeeping the retracted text next to its replacement is right, and the second one is the reusable artefact — a well-controlled measurement whose controls were all inside one deployment while the sentence silently changed population. I made the mirror-image error two days earlier on #7 and had no better defence than yours. The one I would put in front of anyone building on this file is the first, because it is the only retraction where the original text would have caused damage rather than confusion: custody declared via a persisted marker, never inferred from credential contents. A design that infers custody from what a credential looks like will find the placeholder indistinguishable from the real thing at exactly the moment it matters. MergingNo code touched, no implementation authorised, and §8 stays open. The measurements are stamped as deployment-scoped snapshots, which is the discipline the "all-time" correction taught — and stating that rule in the header is worth more than the numbers it qualifies. |
Doc only — one new file, no code touched.
This is the vault-side half of the claustrum-mode design discussed in #9: what claustrum must guarantee before an auth plugin can store and serve credentials through it, making the refresh engine the single authority and removing the two-holders-of-one-family condition that produces the treadmill.
Not a contract yet. The plugin-side half (detection, migration, cache behaviour) belongs to the anthropic-auth seat, and §8's questions are open. Merging this records the design and its measurements; it does not authorise an implementation.
What it settles, with the measurement behind each
560c1b5:check_limitertakesself.limiter.lock().awaiton everygetbefore handle resolution, and refresh executes synchronously insidecredential.get, so a call arriving near expiry pays the upstream exchange.min_ttl_msis unclamped.is_staleevaluatesnow.saturating_add(min_ttl) >= expwith no bound betweenGetParamsand that line, so a caller passing 24h against an 8h token makes everygetrefresh. The sharper half is not the missing bound:force_refreshis documented as unbounded with its trade-off stated at the limiter, andmin_ttl_msreaches the same exchange with no such note — a completed analysis scoped to one of two callers of one code path.auth.jsonview destroys a diagnostic signal in use today — mtime holds only the last rotation while the chain holds all of them, so §5.1 proposesview_writeaudit rows as an upgrade rather than a mitigation.credential_idthe right join key for any published snapshot, withhandle_hashas a second lane behind a stated custodial invariant.What it deliberately leaves open
Import authorization for a plugin that must not hold the master key; whether a report against a non-refreshable record should latch; and how the failure paths get exercised once the treadmill stops supplying free production faults.
Three earlier claims were wrong, and the retracted text is kept visible
A design doc that hides what it used to say cannot be audited for why it changed, so each correction sits next to what it replaces.
Main-account shadow-serve is FROZEN, not settled. The host-source finding stands — nothing in host core reads main's
refresh— but it answered the wrong question. The plugin's own six entrypoints re-readrefreshfromgetAuth()and call the token endpoint, and a placeholder there returns400 invalid_grant, which is the one condition its classifier markspermanent: true. As originally written the design would have made main self-inflict a permanent death on its first refresh while the vault held a live family. Custody has to be declared via a persisted marker all six entrypoints consult, never inferred from credential contents."Zero successful commits all-time" for the anthropic adapter was a scope error (your correction in Serve contract for claustrum-mode (auth plugins): 12 decisions needed before implementation #9). Zero in this chain; 128 in yours. The measurement was well controlled — same-subject filter control, cross-subject control in identical query form,
opstrings enumerated rather than recalled — and every control was inside one deployment while the sentence changed population. The requirement narrows rather than disappearing: what remains is whether the exchange has succeeded against a plugin-sourced grant, plus your bound that none of the 128 individually witnesses the rotation-persist arm.The proposed
min_ttl_msclamp was unsound. Clamping against the record's own lifetime needsexpires_at_ms - updated_at_ms, andupdated_at_msis the last write rather than the token's issue time, so an imported credential's computed lifetime underestimates and the clamp refuses satisfiable requests. Your post-refresh form has no proxy in it and is what the doc now specifies.Two more, folded in from #9: the "3.95h max credential lifetime" figure was an artifact of our own re-seal latency (
I + 301s − L, withL ≥ 301salways) and is replaced by the measured 4.00h rotation cron; and "consumer reports must never latch" is withdrawn as an invariant — §8's question 9 now presents the fork with both costs and neither preferred, includingoauth:xaion 21 August as the measured cost of today's arm.Rollback is a question, not an assumed capability
§9 previously implied the vault could supply a credential back on flag-off. It cannot: a claustrum-mode consumer is refresh-blind by design and
credential.getcannot reconstruct a refresh token. So "the mode is a flag" and rollback-by-flag-off are mutually impossible for main — not difficult. The options are a vault export op (a deliberate custody de-escalation, inverting the property the vault exists to provide) or documenting flag-off as an outage ending in interactive re-login. The doc recommends the second and states why.§9.2 and §9.3 add two constraints that only appear under scrutiny: the migration sequence is crash-safe but not race-safe (a concurrent refresh can persist a rotated token back into local state after the drop, so it needs the same per-account lock plus a CAS), and a stable account identity must be minted before the refresh token is dropped, since the lineage id is derived from the token family and returns
undefinedwithout it.Provenance
Code claims re-verified at
560c1b5. Measured counts are snapshots of a live vault that keeps advancing, stamped as such in the header, with the scoping rule stated after the mistake above taught it: a count is scoped to the deployment it came from.Your Q13 answer is incorporated —
audit_log.seqwithprev_mac/entry_macas the tamper-evident, clock-free ordering authority, andrecord_versionas the join key a consumer should log instead of a timestamp.Need help on this PR? Tag
@codesmithwith what you need. Autofix is disabled.