Gate the undocumented components reason on a real semantic layer - #3828
ayushcodes10 wants to merge 3 commits into
Conversation
GRAPH_REPORT.md's Knowledge Gaps section always offered undocumented components as a possible explanation for an isolated node, regardless of whether the graph could ever actually have one. That phrase only means anything when a semantic layer exists, document, paper or image nodes an LLM extracted meaning from, where an isolated node genuinely could be mentioned but never explained further. A code only graph has no such nodes at all, so offering it as a live possibility on every report is misleading. The reason now checks whether any node in the graph carries a document, paper or image file_type before including that clause, using the same VALID_FILE_TYPES distinction validate.py already draws between the semantic layer and code/rationale/concept nodes, so no new build time flag is needed. Fixes issue 3801.
Covers issue 3801 directly: a code plus concept only graph (no document, paper or image nodes anywhere) must not offer undocumented components as a reason for an isolated node, only possible missing edges. Adding a single document node elsewhere in the same graph shape brings the fuller reason back, confirming the gate tracks the graph's actual content rather than being tied to a particular test fixture shape. Confirmed against pre fix code via git stash: the first test fails there and passes only with the fix applied.
There was a problem hiding this comment.
Graphify reviewed this change.
Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).
Formal verification. No changes could be formally verified in this run.
Graphify review — findings
Makes the Knowledge Gaps section of GRAPH_REPORT.md drop "undocumented components" from the isolated-node explanation unless the graph actually contains document/paper/image nodes, so code-only runs now read "possible missing edges" only. Adds tests covering both the omitted case (code+concept fixture) and the restored case once a semantic node exists.
No blocking issues surfaced.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 594 functions depend on the 237 functions this change touches.
Health — this change adds coupling hotspots:
- new:
_rebuild_code()— 144 callers, 55 callees - new:
main()— 98 callers, 3 callees - new:
dispatch_command()— 2 callers, 125 callees - new:
generate()— 35 callers, 7 callees - new:
run_pipeline()— 8 callers, 13 callees - new:
watch()— 5 callers, 7 callees - new:
load_learning_for_report()— 4 callers, 3 callees - new:
test_report_shows_avg_confidence_for_inferred()— 0 callers, 7 callees - …and 2 more — each is listed as a finding
Verification — 594 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 417 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
300 of 300 test file(s) selected (100%) via static blast radius.
Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.
tests/test_affected_cli.py— full-run-safetytests/test_affected_member_seed.py— full-run-safetytests/test_agents_platform.py— full-run-safetytests/test_analyze.py— full-run-safetytests/test_anthropic_custom_endpoint.py— full-run-safetytests/test_antigravity_install.py— full-run-safetytests/test_apm_fallback_version.py— full-run-safetytests/test_architecture_doc.py— full-run-safetytests/test_astro_extraction.py— full-run-safetytests/test_astro_import_ids.py— full-run-safetytests/test_atomic_canvas_export.py— full-run-safetytests/test_atomic_version_stamp.py— full-run-safetytests/test_atomic_writes.py— full-run-safetytests/test_backend_env_isolation.py— full-run-safetytests/test_backend_extras.py— full-run-safetytests/test_benchmark.py— full-run-safetytests/test_benchmark_raw_graph.py— full-run-safetytests/test_build.py— full-run-safetytests/test_build_merge_dedup_scope.py— full-run-safetytests/test_build_merge_hyperedges_and_prune.py— full-run-safetytests/test_build_merge_shrink_guard.py— full-run-safetytests/test_builtin_global_type_refs.py— full-run-safetytests/test_cache.py— full-run-safetytests/test_callflow_html.py— full-run-safetytests/test_cargo_introspect.py— full-run-safetytests/test_cargo_missing_manifest.py— full-run-safetytests/test_carried_hyperedge_remap.py— full-run-safetytests/test_case_sensitive_resolution.py— full-run-safetytests/test_charmap_encoding.py— full-run-safetytests/test_chunking.py— full-run-safetytests/test_cjs_module_extension.py— full-run-safetytests/test_claude_cli_backend.py— full-run-safetytests/test_claude_md.py— full-run-safetytests/test_cli_broken_pipe.py— full-run-safetytests/test_cli_export.py— full-run-safetytests/test_cli_help.py— full-run-safetytests/test_cluster.py— full-run-safetytests/test_cobol_extractor.py— full-run-safetytests/test_codebuddy.py— full-run-safetytests/test_community_hub_labels.py— full-run-safetytests/test_community_labels_skill.py— full-run-safetytests/test_confidence.py— impact, full-run-safetytests/test_corrupt_graph_json.py— full-run-safetytests/test_cpp_nested_and_cli.py— full-run-safetytests/test_cpp_objc_cross_file_calls.py— full-run-safetytests/test_cpp_preprocess.py— full-run-safetytests/test_cross_extension_reexport_self_cycle.py— full-run-safetytests/test_cross_language_call_resolution.py— full-run-safetytests/test_cross_repo_external_call_guards.py— full-run-safetytests/test_cross_repo_member_calls.py— full-run-safety- … and 250 more
non-code file(s) changed (
CHANGELOG.md) → running the full suite for safety (a code graph can't see config/fixture/data deps)
changed code file(s) with no mapped test (
CHANGELOG.md) — a coverage gap or a missing link — running the full suite rather than only the selected tests
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Formal verification
Could not verify: Could not verify generate.
The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 200 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)
· 10 more finding(s) on lines outside this diff (see the check run).
Windows hash-seed re-exec wait (#3816/#3799), Terraform secret-named variable-default/output redaction (#3817), COBOL fixed-format sequence-number detection (#3813) and PERFORM...THRU both-endpoint linking (#3814), and gate the GRAPH_REPORT undocumented-components reason on a real semantic layer (#3828/#3801). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
Shipped in v0.9.68 (now on PyPI) via an authorship-preserving cherry-pick, so your commit keeps contributor-graph credit. Thanks @ayushcodes10! Gates the undocumented-components report reason on a real semantic layer. Release: https://github.com/Graphify-Labs/graphify/releases/tag/v0.9.68 |
Fixes #3801.
Summary
GRAPH_REPORT.md's Knowledge Gaps section always offered "undocumented components" as a possible reason for an isolated node, regardless of whether the graph could ever actually have one. That phrase only means anything when a semantic layer exists —document/paper/imagenodes an LLM extracted meaning from, where an isolated node genuinely could be "mentioned but never explained further." A code-only graph (--code-only, or any run where semantic extraction never happened) has no such nodes at all, so offering it as a live possibility on every report is misleading.Fix
The reason now checks whether any node in the graph carries a
document/paper/imagefile_typebefore including that clause, reusing the sameVALID_FILE_TYPESdistinctionvalidate.pyalready draws between the semantic layer andcode/rationale/conceptnodes — no new build-time flag needed, matching the issue's own suggested fix.Testing
tests/test_report_gap_thresholds.py: one confirming a code+concept-only graph's isolated-node message omits "undocumented components" and keeps "possible missing edges", and one confirming adding a singledocumentnode elsewhere in the same graph brings the fuller reason back.git stash.test_report.py/test_report_gap_thresholds.pysuite passes.tree-sitter-language-packgrammars not installed in this environment).python3 -m tools.skillgen --checkpasses.