Reviewers 1.2 / 1.3: transport scoring limits, and what template scope actually means - #297
Open
freiburgermsu wants to merge 3 commits into
Open
freiburgermsu wants to merge 3 commits into
freiburgermsu wants to merge 3 commits into
Conversation
Reviewer 1 raised two points that the manuscript answered in a single clause
each. Both are expanded here, with measurements rather than assurances.
TRANSPORT (reviewer 1, comment 2). The caveat was one sentence at the end of
the conditions paragraph in M05 and was never picked up again. It now has its
own paragraph, and states three things the reviewer asked for:
- the 0/1 convention. Compartment indices are RELATIVE, not named, which is
what lets one record serve a bacterial membrane, a mitochondrion and a
plastid envelope. Rhea's in/out maps onto them directly. REACTIONS.md
described "m" only as "a compartment index number" and now explains this.
- what scoring from stoichiometry alone costs. 95% of the gold-graded
transport reactions carry ATP and ADP, and only 2.5% translocate a proton:
the grade is earned by the hydrolysis chemistry, which every predictor
estimates tightly, while the translocation -- the part needing a membrane
potential and a pH gradient -- contributes nothing to the score.
- that the scheme does not over-claim on the trivial cases. Reactions whose
only change is a change of compartment have no net chemistry, and none is
graded gold.
Also recorded: dGPredictor commits on 1.2% of transport reactions against 17.7%
elsewhere, so transport grades rest on fewer independent sources than anything
else in the database.
TEMPLATE SCOPE (reviewer 1, comment 3). The reviewer read "roughly 9,000
reactions used in the ModelSEED and PlantSEED reconstruction templates" as a
claim that the other 47,000 are unusable. That sentence was scoping the
curation effort, not the database, but it invited the reading, so it is
rewritten and the question is answered properly in M13.
Neither of the reviewer's two hypotheses is the reason:
- not GPR. The Gram-negative template types its 8,584 reactions explicitly,
and 2,258 of them are `gapfilling` -- present with no gene requirement at
all. A template is not a set of reactions with genes attached.
- not connectivity. Nothing prunes on graph connectivity.
The real limits are mass and charge balance (58% of reactions) and association
to an annotated functional role (41%). 14,041 reactions now satisfy both, 8,089
of them outside every current template -- which is the headroom the structure
curation in this paper creates.
Page budget. The transport and scope paragraphs are offset by trimming two
passages that said the same thing twice: M05's Uncertainty paragraph repeated
the three median values verbatim from M11 Results (they now appear once, in
Results), and the Experimental anchors paragraph is tightened without losing a
claim. A stray "))" in the Figure 2C reference is fixed in passing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Written before the population decision on the figures PR; the paper is now on the live-only basis and these were the all-records values. 95% -> 96% ATP-coupled, 2.5% -> 2.1% translocating a proton, dGPredictor 1.2% -> 1.4% on transport against 17.7% -> 17.3% elsewhere. The finding is unchanged on either basis, and uniport-like reactions are graded gold on neither (0 of 4,056 and 0 of 2,688). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…n-invariant tree The paper now describes the database as #7 (1ad34ad, on top of the merged ModelSEED#296) leaves it, so the numbers this branch introduced were re-measured there: M05 gold-graded transport that translocates a proton 2.1% -> 2.7% M13 reactions satisfying mass and charge balance 58% -> 69% balanced and carrying an annotated role 14,041 -> 15,128 of which outside every current template 8,089 -> 8,826 The balance share moves because #7 repairs the reaction refresh: Rebuild_Stoichiometry.py had never refreshed the charge embedded in each reagent, and Rebalance_Reactions.py rejected its own `save` flag, so the proton-adjustment step had not been persisting. Between dev (41b20c2) and 1ad34ad, 5,428 live reactions go from imbalanced to balanced: 3,130 by a proton or water adjustment the fixed step now writes, 969 through a participant whose record the re-pick changed, and 1,329 whose flag had been stale against unchanged records and stoichiometry; 25 go the other way, every one through a changed participant. That is the 58% -> 69%. The template union (8,597), the ATP-coupled share of gold transport (96%), dGPredictor's commit rates on and off transport (1.4% / 17.3%) and the role-annotation share (41%) did not move. Measured with Papers/NAR_Update_2026/analysis/review_transport_and_llm.py --live and the reaction records of that tree. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Answers reviewer 1's comments 2 (transport) and 3 (why only ~9,000 reactions reach a reconstruction). Text and docs only — no data or schema changes.
Every number below is measured on the live population (non-obsolete records) by
Papers/NAR_Update_2026/analysis/review_transport_and_llm.py --live, added in the figures PR of this series, against the protonation-invariant tree (freiburgermsu#7,1ad34adc). The dated sections at the end record how the numbers moved as the population and the tree changed.Reviewer 1, comment 2 — transport
Where it came from. One sentence at the end of the conditions paragraph in
M05_methods_thermodynamics.tex:It sat at the end of a paragraph about reporting conditions and was never picked up in Results or Discussion. The reviewer is right that this is under-treated.
What changed. The caveat is promoted to its own
\paragraph{Transport.}and now makes three claims, each measured:Biochemistry/REACTIONS.md, which previously calledmonly "a compartment index number"The first is the reviewer's factual point and it is a documentation gap, not a data gap: indices are relative so that one record can be instantiated against a bacterial plasma membrane, a mitochondrial inner membrane or a plastid envelope when a template builds a model. Rhea's
in/outmaps onto them directly.The second is the substantive limitation, and it is sharper than "transport is approximate". A primary active transporter carries ATP hydrolysis in its stoichiometry; every predictor estimates that hydrolysis tightly; the reaction therefore earns a confident grade from chemistry that was never in question, while the translocation contributes nothing to the score. Transport is graded gold 34.4% of the time against 5.4% elsewhere, and that gap is almost entirely ATP coupling.
Also now recorded: dGPredictor commits to a direction on 1.4% of transport reactions against 17.3% elsewhere, so transport grades rest on fewer independent sources than anything else in the database — directly relevant to the reviewer's point that energy metabolism depends on these reactions.
Deliberately not done here. Flagging transport grades in the released records (an
evidence_caveatkey, or capping transport at silver) changes the shipped schema and is a call for the corresponding authors, not a reviewer-response edit. The paragraph is worded to describe only what ships today —is_transport, which is already in every record. Raised as a follow-up rather than assumed.Reviewer 1, comment 3 — the ~9,000 reactions
Where it came from. The opening sentence of
M10_results_structure_curation.tex, which was scoping the curation pipeline — we prioritised curation on the reactions the templates touch — not asserting anything about usability. The reviewer read it as a coverage claim. Since a careful reader drew that conclusion, the sentence is rewritten rather than the reviewer corrected.What changed.
M10's opening now says plainly that the scope is a curation priority, and points atM13, where the question is answered.Both of the reviewer's hypotheses are wrong, and the data says so:
conditional, 2,258gapfilling, 31spontaneous, 11universal. A quarter of the template is present with no gene requirement at all. A template is not a set of reactions with genes attached, so missing GPRs cannot be the gate.The real limits, now stated in
M13:That last row is the useful answer: roughly a doubling of template scope is already available, and the structure curation this paper reports is what makes it reachable.
M14gains one sentence naming compartment-aware scoring and the ~50% of reactions with no direction as the two standing limits.Page budget
The NAR checklist in
main.texnotes the draft is already at the limit. The two new paragraphs are offset by removing genuine duplication rather than by cutting content:M05's Uncertainty paragraph repeated the three median uncertainties (0.63 / 10.41 / 17.01 kcal mol⁻¹) verbatim fromM11Results. They now appear once, in Results.))in the Figure 2C cross-reference is fixed in passing.Net effect: body text still ends on page 6; six reference entries now spill to page 7 where the baseline fit in six pages. Three of the citations on that page (
mavrovouniotis1993,xu2008,noor2014) are the orphaned ones addressed in the LLM PR of this series, so the spill is expected to resolve once that lands.\lastpage{6}should be re-checked after all four PRs merge.Verification
latexmk -pdf main.texclean; no undefined references or citations.🤖 Generated with Claude Code
Merge note (added after the live-basis conversion)
This branch and #297 both touch
M13_results_mapping.tex, which has a single-line body, so git cannot auto-merge them:\paragraph{Reconstruction scope.}after the atom-mapping sentenceThe resolution is to keep both, in this order:
No other pair of branches in this series conflicts — verified with
git merge-treeacross all four.Review pass (2026-09-23)
Self-review against the shipped data after the population decision in #299. One defect: the four percentages in the Transport paragraph were all-records values; they are now live — 96% ATP-coupled, 2.1% translocating a proton, dGPredictor 1.4% on transport vs 17.3% elsewhere. Re-verified unchanged: the template-scope numbers (8,597 / 2,258
gapfilling/ 58% / 41% / 14,041 / 8,089) were computed on the live basis already; uniport-like reactions are graded gold on neither population (0 of 4,056 all-records, 0 of 2,688 live).Re-measured against the protonation-invariant tree (2026-09-25)
#296 merged into
devon 2026-09-23. Its follow-up, freiburgermsu#7 (protonation-invariant-gate,1ad34adc), fixes four defects in the structure pipeline and re-runs the reaction refresh, and it is the tree the paper now describes. Every number this branch introduced was re-measured there (commit e1c4f6b), and the tables above now show the right-hand column:dev(41b20c2)Unchanged: the template union (8,597), 2,258
gapfilling, 96% ATP-coupled, dGPredictor 1.4% / 17.3%, EC share 41%.Why the balance share jumps eleven points. #7 repairs the reaction refresh itself:
Rebuild_Stoichiometry.pyhad never refreshed the charge embedded in each reagent, andRebalance_Reactions.pyrejected its ownsaveflag, so the proton-adjustment step had not been persisting. Between the two trees 5,428 live reactions go from imbalanced to balanced — 3,130 by a proton or water adjustment the fixed step now writes, 969 through a participant whose record the re-pick changed, and 1,329 whose flag had been stale against unchanged records and stoichiometry — and 25 go the other way, every one through a changed participant. The paragraph's claim, that balance is the first gate on template promotion, is unchanged; the gate is now measured on a refresh that actually ran.Merge order. This branch is text-only against
devand merges cleanly onto #7 (git merge-tree, verified), but its numbers do not reproduce ondevalone:review_transport_and_llm.py --liverun on this branch by itself returns the left column. Merge #7 first, then this. TheM13overlap with #299 is unchanged and resolves as the merge note above says, with the atom-mapping sentence at its re-measured values (26,242 / 19,809 / 6,433).