Give a bitten fork a DHT of its own - #26
Merged
Merged
Conversation
Stripping bootNodes was meant to keep a fork off the source network and does not. Protocol names are built from the genesis hash plus the fork id, and a fork keeps the source's genesis hash, so without a fork id its names are byte-identical to the source's: it sits on the source's DHT and finds nodes there. An asset-hub collator was found holding established connections to public mainnet nodes while its spec carried zero bootNodes, which is a forked node rejoining the source network and following its longer chain: success on every metric while not being a fork at all. The chain spec's fork_id field is what takes it off that DHT. Nodes only synchronize with other nodes holding the same value, and it names protocols and nothing else, so it is not part of the genesis hash or the state and snapshots taken before it still restore. Leaving the source's DHT is also what makes the missing bootNodes fatal, which is why the fork id and the bootnode land together. A fork with a fork id has a DHT of its own, and Kademlia inserts peers manually, so connecting to a peer never adds it to the routing table: the table fills from bootnodes or from identify, and a spec shipping neither leaves every node's table empty and authority discovery resolving no addresses. Measured on a fork with a fork id and no bootnode, every one of a validator's random Kademlia walks returned peers-not-found. With one bootnode injected, the same node reported nine walks and nine peers-found, six records stored, and published its own addresses. Only the first relay node can be named ahead of the spawn, because it is the only one whose p2p port the generator pins; the rest take ephemeral ports that change every time. One entry is enough, since Kademlia walks out from it to the rest. Every node then bootstraps through that one, which the fork already depends on anyway: it is the node every collator reaches the relay through. Parachain specs ship none, because a fork runs a single collator per parachain and has no parachain peer to find. Only the shipped specs get a fork id or a bootnode. The copies the bite itself runs against keep the source's bootNodes and no fork id, because it has to reach the real network to warp-sync from it. These changes need a patched polkadot-parachain, and the fork runs one. Cumulus hardcodes fork_id: None when it builds the relay node a collator embeds (relay-chain-minimal-node/src/network.rs), so a stock collator keeps its relay node on /<genesis>/kad while the validators move to /<genesis>/<fork id>/kad. Its collation protocol still carries the fork id, so it connects to them and then resolves none of them, and the relay keeps producing blocks while the parachains stop. I will fix that upstream. Signed-off-by: Cosmin Paraschiv <cosmin@parity.io>
mordamax
approved these changes
Sep 15, 2026
csmnprschv
added a commit
that referenced
this pull request
Sep 17, 2026
This reverts commit a6118da. The fork id needs a patched polkadot-parachain, which CI does not have, so every parachain stalls in the fork-e2e jobs. I will restore it once the cumulus fork id fix lands in a weekly release. Signed-off-by: Cosmin Paraschiv <cosmin@parity.io>
csmnprschv
added a commit
that referenced
this pull request
Sep 17, 2026
Stripping bootNodes was meant to keep a fork off the source network and does not. Protocol names are built from the genesis hash plus the fork id, and a fork keeps the source's genesis hash, so without a fork id its names are byte-identical to the source's: it sits on the source's DHT and finds nodes there. An asset-hub collator was found holding established connections to public mainnet nodes while its spec carried zero bootNodes, which is a forked node rejoining the source network and following its longer chain: success on every metric while not being a fork at all. The chain spec's fork_id field is what takes it off that DHT. Nodes only synchronize with other nodes holding the same value, and it names protocols and nothing else, so it is not part of the genesis hash or the state and snapshots taken before it still restore. Leaving the source's DHT is also what makes the missing bootNodes fatal, which is why the fork id and the bootnode land together. A fork with a fork id has a DHT of its own, and Kademlia inserts peers manually, so connecting to a peer never adds it to the routing table: the table fills from bootnodes or from identify, and a spec shipping neither leaves every node's table empty and authority discovery resolving no addresses. Measured on a fork with a fork id and no bootnode, every one of a validator's random Kademlia walks returned peers-not-found. With one bootnode injected, the same node reported nine walks and nine peers-found, six records stored, and published its own addresses. Only the first relay node can be named ahead of the spawn, because it is the only one whose p2p port the generator pins; the rest take ephemeral ports that change every time. One entry is enough, since Kademlia walks out from it to the rest. Every node then bootstraps through that one, which the fork already depends on anyway: it is the node every collator reaches the relay through. Parachain specs ship none, because a fork runs a single collator per parachain and has no parachain peer to find. Only the shipped specs get a fork id or a bootnode. The copies the bite itself runs against keep the source's bootNodes and no fork id, because it has to reach the real network to warp-sync from it. These changes need a patched polkadot-parachain, and the fork runs one. Cumulus hardcodes fork_id: None when it builds the relay node a collator embeds (relay-chain-minimal-node/src/network.rs), so a stock collator keeps its relay node on /<genesis>/kad while the validators move to /<genesis>/<fork id>/kad. Its collation protocol still carries the fork id, so it connects to them and then resolves none of them, and the relay keeps producing blocks while the parachains stop. I will fix that upstream. Signed-off-by: Cosmin Paraschiv <cosmin@parity.io>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stripping bootNodes was meant to keep a fork off the source network and
does not. Protocol names are built from the genesis hash plus the fork
id, and a fork keeps the source's genesis hash, so without a fork id its
names are byte-identical to the source's: it sits on the source's DHT
and finds nodes there. An asset-hub collator was found holding
established connections to public mainnet nodes while its spec carried
zero bootNodes, which is a forked node rejoining the source network and
following its longer chain: success on every metric while not being a
fork at all.
The chain spec's fork_id field is what takes it off that DHT. Nodes only
synchronize with other nodes holding the same value, and it names
protocols and nothing else, so it is not part of the genesis hash or the
state and snapshots taken before it still restore.
Leaving the source's DHT is also what makes the missing bootNodes fatal,
which is why the fork id and the bootnode land together. A fork with a
fork id has a DHT of its own, and Kademlia inserts peers manually, so
connecting to a peer never adds it to the routing table: the table fills
from bootnodes or from identify, and a spec shipping neither leaves
every node's table empty and authority discovery resolving no addresses.
Measured on a fork with a fork id and no bootnode, every one of a
validator's random Kademlia walks returned peers-not-found. With one
bootnode injected, the same node reported nine walks and nine
peers-found, six records stored, and published its own addresses.
Only the first relay node can be named ahead of the spawn, because it is
the only one whose p2p port the generator pins; the rest take ephemeral
ports that change every time. One entry is enough, since Kademlia walks
out from it to the rest. Every node then bootstraps through that one,
which the fork already depends on anyway: it is the node every collator
reaches the relay through. Parachain specs ship none, because a fork
runs a single collator per parachain and has no parachain peer to find.
Only the shipped specs get a fork id or a bootnode. The copies the bite
itself runs against keep the source's bootNodes and no fork id, because
it has to reach the real network to warp-sync from it.
These changes need a patched polkadot-parachain, and the fork runs one.
Cumulus hardcodes fork_id: None when it builds the relay node a collator
embeds (relay-chain-minimal-node/src/network.rs), so a stock collator
keeps its relay node on
/<genesis>/kadwhile the validators move to/<genesis>/<fork id>/kad. Its collation protocol still carries the forkid, so it connects to them and then resolves none of them, and the relay
keeps producing blocks while the parachains stop. I will fix that
upstream.