Skip to content

Merge upstream/main (4 commits, through b44550e3) into fork/consolidated - #292

Merged
bompus merged 5 commits into
fork/consolidatedfrom
merge/upstream-2085
Sep 29, 2026
Merged

bompus merged 5 commits into
fork/consolidatedfrom
merge/upstream-2085

Conversation

@bompus

@bompus bompus commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Merges upstream/main through b44550e (4 commits) into fork/consolidated:

Conflict resolution:

README: the upstream merge point moves to b44550e. No fork-vs-upstream row or feature list changes; the "About this fork" section was checked.

Checks: npm run build and the full suite (386 files, 6,350 passed, 34 skipped; kernel golden dumps unchanged).

danusha2345 and others added 5 commits September 29, 2026 03:01
…aced file (colbymchenry#1902) (colbymchenry#1917)

* fix(watcher): follow a rebuilt index instead of syncing into the replaced file (colbymchenry#1902)

`codegraph index` rebuilds through `CodeGraph.recreate()`, which unlinks
`.codegraph/codegraph.db` and creates a new file at the same path. A running
MCP daemon kept its handle on the unlinked inode and its watcher kept
"auto-syncing" into it, so no edit reached the on-disk index. The colbymchenry#925
self-heal only ran on the tool-call path and ran no catch-up after reopening.

- `sync()` now checks, once it holds the index mutex and the write lock,
  whether the database was replaced on disk (one stat, the same inode check
  `reopenIfReplaced` uses). If so it reopens the live file in place and widens
  a scoped sync to a full reconcile. A failed reopen (rebuild mid-way) returns
  the lock-busy shape so the watcher keeps its pending files and retries.
- `reopenIfReplaced()` refuses while an index/sync holds the mutex, so the
  tool-call path never closes the handle an in-flight sync writes through.
- When a tool call's `freshen()` does reopen, the engine runs its existing
  catch-up sync (serialized on the same mutex), but only for its own watched
  instance, so an engine that is not the project's writer never writes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sync): step aside for a rebuild that has recreated the file but not locked it yet

`codegraph index` recreates the database, then takes the write lock inside
indexAll. With the watcher now following a replaced file, a sync landing in
that gap reopened the empty file and started a full reconcile, holding the
lock the rebuild was about to request — `codegraph index` could then fail
with "another process may be indexing". Review flagged the gap; this closes
it. After following a replaced file, a sync that finds no index_state yet
(the rebuild has not started indexing) and a file written within the last
two minutes returns the lock-busy shape so the watcher retries; the first
sync after that reconciles the whole tree once, then syncs go back to scoped.
The time bound keeps a rebuild that died before indexing from parking the
watcher forever.

New test: fails without the guard, passes with it; 136 related tests pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sync): keep the full catch-up until a full sync completes (colbymchenry#1902)

Two gaps in following a replaced database:
- Only a reopen done by sync() itself set pendingFullReconcile. The tool-call
  self-heal (reopenIfReplaced) swapped the handle without it, so the next
  scoped watcher sync skipped whatever the old handle had absorbed. The flag
  is now set wherever the database is reopened.
- The flag was cleared before the sync ran, so a sync that threw lost the
  catch-up. It is now cleared only after a full sync completes.
Ported from the maintainer's hardening of this fix in colbymchenry#1919.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sync): report a rebuild in progress as lock contention, not an empty sync (colbymchenry#1902)

Since colbymchenry#2014 the watcher tells lock contention apart from success only by
LockUnavailableError; it no longer recognises the all-zero result. Following
a rebuilt index still returned that shape when the reopen failed or when the
recreated file had not been indexed yet, so the watcher counted the step-aside
as a clean sync and dropped the pending edits. Both paths now throw
LockUnavailableError. A watcher test proves the edit stays pending through the
rebuild gap and is reconciled once the rebuild finishes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: danusha2345 <danusha2345@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: danusha2345 <ewidusoc498@gmail.com>
….ts is not TypeScript (colbymchenry#1910) (colbymchenry#2082)

* fix(extraction): an MPEG-TS video named .ts is not TypeScript (colbymchenry#1910)

`.ts` mapped to TypeScript by extension alone, so MPEG transport-stream
fixtures (`testdata/*.ts`) were parsed by tree-sitter: ~28 s of CPU for a
900 KB clip, minutes for a directory of them, for zero symbols.

Recognise the stream from the head of the file — the 0x47 sync byte at
offsets 0, 188, 376 and 564 plus a NUL byte, which every stream carries in
its first packets and UTF-8 source never does — and drop it at discovery:
not indexed, not parsed, not counted, not tallied as an unsupported
language. The check reads under 1 KB and only for `.ts` files; the batch
reader sniffs the bytes it already read, and the single-file path (sync,
watcher) does the same head read, so a clip handed in by name is skipped
too. The skip happens before parse dispatch, so no kernel mirror is needed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(extraction): a video .ts is never pending on the git path or a scoped sync

Review found two places that decide "is this a source file" without the
MPEG-TS check the scan applies: the git fast path's candidate loop in
getChangedFiles and the scoped-sync path filter (watcher events). An
untracked clip in a git repo was reported as `added`, skipped by sync
without a record, and reported as `added` again on every status; a tracked
.ts that became a clip stayed `modified` with its stale nodes. Both now
treat the clip as non-source: removed when tracked, ignored otherwise.

Regression tests fail without this change and pass with it, on both arms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(extraction): require a binary head before calling a .ts file video (colbymchenry#1910)

Review on colbymchenry#1915: the sniff took four aligned 0x47 sync bytes plus any NUL
as proof of a transport stream, so TypeScript with a `G` at four 188-byte
strides and one raw NUL in a comment was dropped unindexed. The sniff now
needs the sync byte on sixteen consecutive packets (3 KB of head) and at
least 1/64 of the head to be control bytes. Compressed payload carries
8-20% of them (checked on ffmpeg-made H.264/AAC, MPEG-2 and MP2 streams);
source text carries none. A shorter clip is cheap to parse anyway.

Also drops the benchmark figures from the changelog entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(extraction): a file over the size limit is never read (colbymchenry#1910)

The size gate ran after `readFile`, so a committed video or blob fixture was
decoded in full — 3.4 GB of RSS for a 400 MB file, `Invalid string length`
past ~512 MB — only to be stored as skipped. It was then read again by every
pass that scans files by name: the framework detectors' `readFile` and the
resolver's cached reader, three full passes on the reporter's fixtures.

Stat first, everywhere a source file is read by path: the batch reader and
`extractFile` store an oversize file with a size stamp in place of its
content, change detection hashes the same stamp (so a same-size rewrite of a
file nothing is indexed from is not a change, while crossing the limit in
either direction is), and both by-name readers return null above the limit.
The limit moves to `src/file-limits.ts` so the three sites share one number.

strace on a sparse 400 MB `.ts`: 153 603 reads before, 0 after; indexing it
next to a real file costs no measurable RSS. Stacked on colbymchenry#1914, which shares
the batch-reader hunk.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(extraction): one-purpose size-stamp reader, comments in operation order

`readSourceOrStamp` took a `stats` argument no caller passed and returned
stats no caller read; it now returns the text or stamp change detection
hashes. The batch reader's two comments are back in the order its code runs:
the size gate, then the MPEG-TS head check. No behaviour change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(extraction): bound the read itself, not only the stat before it (colbymchenry#1910)

A file can grow between the stat that gates the size limit and the read
that follows (a log, a download, a build output being written). Every
source read now goes through readBoundedSource/readBoundedSourceSync: the
open descriptor is re-checked, and the read stops one byte past the limit,
so an oversize file is still stamped rather than decoded. The extraction
batch reader, single-file extraction, framework detectors and the resolver's
file cache all use it. Ported from the maintainer's hardening of this fix
in colbymchenry#1919.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(extraction): bound the parse-retry reads too (colbymchenry#1910)

The retry passes for files whose worker crashed or timed out re-read the
source with an unbounded fsp.readFile. The file can have grown since the
first, bounded read; an oversize file is now skipped by the retry instead
of being decoded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(extraction): check an oversize file against its size stamp everywhere (colbymchenry#1910)

Review on colbymchenry#1915:

- The viewer (readFileShape, hasDriftedOnDisk, the source endpoint) and
  MCP's drift gate hashed a file's real bytes, but the index stores the
  size stamp for a file over the limit, so every unchanged 1-8 MB file read
  as "changed on disk". oversizeStamp moves to file-limits.ts with an
  indexedHashInput helper, and all four checks hash what the index stored.
  An oversize file is now compared without being read.
- git-index-currency's "keeps a committed path pending when sync cannot
  read it" injected its failure only into readFileSync, which the bounded
  reader no longer calls, so it failed. The failure now reaches openSync
  too.
- The changelog entry drops its benchmark figures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mcp): check an oversize answer file against its size stamp (colbymchenry#1910)

colbymchenry#2043's answer-freshness check hashes each contributing file whole and
compares it with the stored contentHash. For a file over the index's size
limit the stored hash is the size stamp, so an unchanged oversize file
would read as stale. Compare it by its stamp, without reading it, like
the viewer and the MCP drift gate do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(extraction): judge a .ts clip from the bytes it is read for, not at discovery (colbymchenry#1910)

Discovery sniffed the head of every `.ts` file on every full scan — an
open + 3 KB read + close per file, before anything asked for its content.
On a 10k-file TypeScript tree that took a warm-cache scan from ~105 ms to
~243 ms; on a slow disk or a network share it is a random read per file
per scan, the access pattern colbymchenry#1231 and the SMB reports (colbymchenry#1014, colbymchenry#448) are
about.

The bytes are now judged where they are read anyway, so the check costs no
I/O and runs only for a file that is new or changed: the batch reader (as
before), `indexFile`, the sync reconcile (an untracked clip is skipped, a
tracked file that became one is removed through the same removal path as
a deletion), and both `getChangedFiles` paths. The scan lists `.ts` files
by name again; a test pins that discovery opens none of them on the walk
or the git path. Scan time is back to main's (~98 ms), and a full index is
unchanged within noise.

Also collapses the four colbymchenry#1910 changelog bullets (colbymchenry#1914 and colbymchenry#1915 each added
both) to two, without the benchmark figures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: danusha2345 <danusha2345@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: danusha2345 <ewidusoc498@gmail.com>
…itialized (colbymchenry#1895) (colbymchenry#2083)

* fix(directory): a schema-less codegraph.db does not make a project initialized (colbymchenry#1895)

isInitialized() accepted any existing .codegraph/codegraph.db, so an empty or
table-less file in an ancestor directory (an interrupted init, a stray touch,
a never-populated ~/.codegraph/) captured the upward walk of
findNearestCodeGraphRoot / resolveProjectPath / resolveServerRoot for every
project beneath it: the real index of a sub-project became unreachable, the
sub-project down-scan never ran, and commands either operated on the broken
file or failed opening it instead of showing the not-initialized guidance.

An initialized project is now one whose db carries the schema: the probe is
gated by file size (>= the 100-byte SQLite header), then the header magic,
then a read-only node:sqlite open with a single sqlite_master lookup for the
`nodes` table, memoized per path + mtime + size so hot callers (the prompt
hook, MCP root resolution on every call) never reopen an unchanged db. The
handle is closed immediately so the probe never holds the db on Windows; a
locked/busy db counts as live.

createDirectory() uses the same predicate, so `codegraph init` in the
directory with the broken file repairs it in place (the schema is applied
with CREATE ... IF NOT EXISTS) instead of refusing with "Already
initialized"; the CLI says what it found. Nothing is deleted automatically.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(directory): a database we cannot open read-only still counts as initialized

The schema probe failed CLOSED: a real WAL codegraph.db in a directory the
process cannot write (a read-only checkout, a mount, another user's tree)
made the read-only node:sqlite open throw — SQLite must create `-shm` to
open a WAL database — and isInitialized() reported the project as not
initialized, where main's file-exists check had been right.

The probe now fails OPEN. Once the size gate and the `SQLite format 3\0`
header have passed, the file is a SQLite database; only a SUCCESSFUL
sqlite_master query proving the `nodes` table absent may return false. Any
open/prepare error (locked, busy, cantopen, readonly, disk I/O) returns
true, the pre-existing behaviour for a database we cannot inspect. No error
strings are matched. The per path + mtime + size memo is unchanged, and
schema-less/empty files still return false, so `codegraph init` still
repairs them and never touches a real db.

Test: a WAL db copied into a `.codegraph/` made read-only with chmod stays
initialized — skipped on win32 and when running as root (root ignores mode
bits).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(init): refuse a codegraph.db that is not SQLite instead of promising a rebuild

Review found that a `.codegraph/codegraph.db` holding bytes that are not
SQLite at all counted as "schema-less", so `init` printed "rebuilding it"
and then failed with SQLite's "file is not a database". Only an empty file
or a SQLite database without the codegraph tables can be rebuilt in place;
`hasSchemalessDb` now means exactly that, and the new `hasForeignDbFile`
names the other case. `init` stops on it with exit 1, says which file it
is and to move or delete it, and never touches it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(directory): ask SQLite what codegraph.db is instead of reading its header ourselves (colbymchenry#1895)

The header check opened and closed a descriptor of our own on
`codegraph.db`. Closing any descriptor on a database file drops every
POSIX lock the process holds on it, including those of a connection it
already has open (sqlite.org/howtocorrupt.html §2.2.1) — and the MCP
server resolves projects through isInitialized on every call while it
holds the index as its writer. mcp-projectpath-lifecycle's takeover tests
caught it: after a real owner exits, the taking-over engine's next read
failed with "disk I/O error" (2 of 21 fail with the header read; 21/21
on main and with this change).

One read-only SQLite probe now answers all three questions: the schema
is there, SQLite opened it and the schema is not (repairable), or SQLite
refused it as not a database (SQLITE_NOTADB). Anything else — locked,
busy, no `-shm` possible — still fails open as initialized. SQLite's own
connections share one lock table per file, so opening and closing a
second one leaves the first one's locks alone. A test pins that
isInitialized and both init helpers never open the file themselves while
this process holds it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): keep the colbymchenry#1895 entry clear of colbymchenry#1917's, so the carries merge in any order

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: danusha2345 <danusha2345@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…olbymchenry#2085)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@bompus
bompus merged commit 5239cfe into fork/consolidated Sep 29, 2026
@bompus
bompus deleted the merge/upstream-2085 branch September 29, 2026 03:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants