Repository navigation
client must silently blackhole messages on stale destinations #112
Description
Activity
Hi Tiara,
I roughly looked into this, and I think there are two separate problems here, plus a few things in the test setup that make them hard to see.
Two problems
Problem 1: a stale dst is reused on a route-cache hit
homa_route_get()returns a cachedhoma_routeon an rhashtable hit without callingdst_check()(homa_peer.c:441-451). This came in with f30a0ef, which introducedhoma_route. Its design is that an invalid dst means discarding the route and re-establishing it. Today onlyhoma_route_validate()does that, and it runs on the transmit path (homa_outgoing.c:404,:464), not on lookup.When the egress device is removed, the kernel marks its dsts DEAD and repoints them at
blackhole_netdev(MTU 68, output =dst_discard_out). The nextsendmsgreuses that dst:homa_message_out_init()computesmax_seg_data = 68 - 20 - 56 = -8.homa_tx_copy_from_user()keeps writing segment headers until the buffer is full, then returnsEINVAL.- Nothing is transmitted, so
homa_route_validate()never runs and the stale route stays cached.
Reproduced on 6.17.8 (IPv4, veth/netns as in your steps, default route present) with a request/response client: after the veth is deleted, every
sendmsgreturnsEINVAL, indefinitely.To make the expected behaviour concrete, I briefly summarized how other kernel transports behave in the same situations, which may serve as a reference. The last two columns are what I think Homa should do, and what it does today. For each case, the table shows what
sendmsgreturns and what a laterrecvmsgreturns.Case TCP UDP SCTP Expected for Homa (IMO) Homa today (53f7963) Cached dst invalid, peer alive (e.g. route moved to another interface) sendreturns bytes; route re-resolved at transmit;recvgets datasendmsgre-resolves inside the call, returns bytes;recvmsggets datasendmsgreturns bytes; route re-resolved at transmit;recvmsggets datasendmsgre-resolves inside the call, returns 0;recvmsggets the responseNo RPC outstanding: every sendmsgreturnsEINVAL, nothing to receive (reproduced). RPC outstanding: a RESEND re-routes first, thensendmsgreturns 0 and the outcome depends on where the new route leads (reproduced)Cached dst invalid, re-lookup finds no route sendreturns bytes (queued); error surfaces latersendmsgreturns-ENETUNREACHsendmsgreturns bytes; packet dropped; later association-loss notificationsendmsgreturns-ENETUNREACHCached peer: no re-lookup happens, so same as the row above. New peer: sendmsgfails with "couldn't find route for peer" (from code, not tested)Route valid, peer dead sendreturns bytes; after retransmits (~15 min by default)recvreturnsETIMEDOUTsendmsgreturns bytes;recvmsgblocks, no errorsendmsgreturns bytes; after path retransmitsrecvmsggetsSCTP_COMM_LOSTsendmsgreturns 0;recvmsgreturnsETIMEDOUTafter ~100 msSame as expected: recvmsgreturnsETIMEDOUTafter 100–190 ms (reproduced)None of these transports fail
sendmsgbecause a cached dst went stale, or because the peer is dead.sendmsgfails synchronously only when a fresh route lookup fails.Device removal is only one of the ways a cached dst can become invalid, just the most visible one. Below are the ones I found, with how the kernel invalidates the dst and what would happen if Homa keeps using the old one:
Event How the kernel invalidates the dst Effect if Homa keeps using it Device unregistered dst marked DEAD, moved to blackhole_netdev(dst.c:148-153)negative max_seg_data→EINVAL, or silent drop (reproduced)Route added / changed / deleted rt_genidbump (fib_trie.c:1316,1379,1749); deleted route's dsts marked DEADold egress / next hop kept Address added / removed, link up / down rt_genidbump (fib_frontend.c:1457,1474,1485,1521,1532)old source address or a down interface kept Local MTU change rt_genidbump (fib_frontend.c:1536)old MTU kept for segmentation First PMTU exception for a destination (learned by any protocol) shared nexthop dsts marked KILL ( route.c:720-728)new PMTU not picked up ICMP redirect marked KILL ( route.c:800)new next hop not used fib rules, nexthop objects, ARP, sysctls, IPsec policy rt_genidbumpold lookup result kept Problem 2: Homa does not handle PMTU ICMPs
The IP layer stores PMTU, per destination in the fib nexthop exception (
fnhe_pmtu), and shares it across protocols. But it does not write that value itself when an ICMP arrives.icmp_unreach()only extracts the MTU and hands the error to the protocol'serr_handler(icmp.c:846-850,903-921). The transport is expected to callipv4_sk_update_pmtu()/ip6_sk_update_pmtu().Homa sends with DF set by default (
IP_PMTUDISC_WANT,ip_output.c:510), so a smaller-MTU hop drops the packet and replies with an ICMP. For comparison, here is roughly how TCP and UDP handle that ICMP, next to what Homa does today:Protocol On "Fragmentation Needed" / "Packet Too Big" PMTU updated Result TCP tcp_v4_err→tcp_v4_mtu_reduced→inet_csk_update_pmtu(tcp_ipv4.c:583-593)yes resegments and retransmits; connection continues UDP udp_err→ipv4_sk_update_pmtu(udp.c:982-983)yes later sends use the new MTU; EMSGSIZEreported withIP_RECVERRHoma, IPv4 homa_err_handler_v4treats it as a generic DEST_UNREACH (homa_plumbing.c:1757-1788)no all RPCs to that host aborted; recvmsgreturnsEHOSTUNREACHHoma, IPv6 homa_err_handler_v6does not handle PKT_TOOBIG (homa_plumbing.c:1803-1827)no ICMP ignored; large packets keep being dropped; recvmsgeventually returnsETIMEDOUTThe two problems interact. Even when another protocol on the host has already recorded the new PMTU and the kernel has marked the shared dst KILL, Homa's cache hit (Problem 1) keeps the old dst. The Homa rows in this table are from reading the code; I haven't tested them.
Other chaos your test setup brought, which may confuse you on this problem
- The response length is 0 and the client never calls
recvmsg. Every RPC therefore stays outstanding. Homa's timer sends RESENDs for outstanding RPCs every few ms, and that path callshoma_route_validate(), which replaces the stale route before the nextsendmsg. As a result theEINVALfrom Problem 1 never shows up with this client. With a client that waits for each response, it shows up on every send. - The host has a default route. Deleting the veth also removes the connected route for 10.80.0.0/24. When Homa looks up 10.80.0.2 again, the lookup matches the default route and succeeds, so Homa gets a perfectly valid dst that points at the default gateway.
sendmsgreturns success and the packets go to the gateway, which drops them because 10.80.0.2 is not reachable there. That is the "blackhole": Homa isn't hiding an error, the kernel really did give it a route. Without a default route, the lookup would fail andsendmsgwould return-ENETUNREACH. - The link and the server are removed in the same step. That mixes two different events, and Homa reports them in different places. A broken route is handled inside the next
sendmsg(re-resolve, or-ENETUNREACHif there is no route). A dead peer can't be known atsendmsgtime at all, so it is reported later byrecvmsgasETIMEDOUT, about 100 ms after the response fails to arrive. Since the client never callsrecvmsg, the second error never shows up. Testing them separately makes this clearer: remove only the link while the server stays reachable another way, or kill only the server and leave the network alone. - On the note in client must silently blackhole messages on stale destinations #112: the
EINVALdescribed in What evidence is useful for reporting NIC support status? #107 is stock behaviour (Problem 1). It appears whenever no RPC is outstanding when the device goes away, so I don't think it came from your patch.
Thanks to both of you for all this analysis. I have not considered all of these conditions before (such as an MTU so small that max_seg_data becomes negative) so this is super-helpful. It may take me a bit before I can work through all of these. In the meantime I have a question for you, Tianyi.
You said:
it runs on the transmit path (homa_outgoing.c:404, :464), not on lookup.
What do you mean by "lookup"?
- Hi John, By "lookup" I meant the route-cache lookup in homa_route_get(). - Tianyi…________________________________ From: John Ousterhout ***@***.***> Sent: Monday, 05 October 2026 17:31:26 To: PlatformLab/HomaModule ***@***.***> Cc: breakertt ***@***.***>; Comment ***@***.***> Subject: Re: [PlatformLab/HomaModule] client must silently blackhole messages on stale destinations (Issue #112) [https://avatars.githubusercontent.com/u/10283530?s=20&v=4]johnousterhout left a comment (PlatformLab/HomaModule#112)<#112 (comment)> Thanks to both of you for all this analysis. I have not considered all of these conditions before (such as an MTU so small that max_seg_data becomes negative) so this is super-helpful. It may take me a bit before I can work through all of these. In the meantime I have a question for you, Tianyi. You said: it runs on the transmit path (homa_outgoing.c:404, :464), not on lookup. What do you mean by "lookup"? — Reply to this email directly, view it on GitHub<#112?email_source=notifications&email_token=AEHMMGBAF2MIYLCHANNKGEL5SPEF5A5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKOJZHA3DSNJVGA32M4TFMFZW63VHMNXW23LFNZ2KKZLWMVXHJLDGN5XXIZLSL5RWY2LDNM#issuecomment-5998695507>, or unsubscribe<https://github.com/notifications/unsubscribe-auth/AEHMMGH6RWN3ARPYSZ5STTT5SPEF5AVCNFSNUABFKJSXA33TNF2G64TZHMYTGOJWGM3DANZXHNEXG43VMU5TKNRZHA4DQNRRGI32C5QC>. You are receiving this because you commented.Message ID: ***@***.***>
- added 2 commits that reference this issue
on Oct 6, 2026 I believe that my latest push fixes 2 problems from this issue:
- Routes get revalidated in homa_message_out_init before computing packet geometry for an RPC.
- homa_message_out_init checks for unusuably-small MTUs and fails the RPC with EHOSTUNREACH.
I have not yet addressed the issue of changes in MTU after an RPC has been initialized, or handling PMTU. This is a bigger issue that I'll try to schedule at some point in the future.
In the meantime, Tianyi, if you have time and interest, I'd love for you to explore what happens when the MTU does change in the middle of an RPC (I assume that the only problem is if it gets smaller). I have no idea whether Homa handles this in a reasonable way (vs. just ending up with a hung RPC that never completes).
For now I'm going to leave this issue open as a reminder.
Reacted by breakerttIn case you already started looking at the push above, I had a change of heart and decided to add route validation to homa_route_get ("lookup", as you put it). With this addition, I don't think that homa_message_out_init needs to perform an additional validation, so I removed that code. This will result in route validations in other situations that weren't covered by my first push, such as in homa_xmit_unknown and homa_need_ack_pkt.
In case you already started looking at the push above, I had a change of heart and decided to add route validation to homa_route_get ("lookup", as you put it). With this addition, I don't think that homa_message_out_init needs to perform an additional validation, so I removed that code. This will result in route validations in other situations that weren't covered by my first push, such as in homa_xmit_unknown and homa_need_ack_pkt.
Hi John,
After checking I think it may be incorrect to remove validation in
homa_message_out_init. When server replies the response, it may been already a while since the server rpc allocation (hencehoma_route_get) and route status may changed.Cheers,
Tianyi- I considered this issue but convinced myself that it isn't a game changer. Changes in route won't be a problem because the route is revalidated before sending packets, which is even later than homa_message_out_init. So, I think the only potential issue is changes in MTU (which affect packet geometry); those can still happen after homa_message_out_init finishes, so the only question is the size of the window of vulnerability. The worst that happens is more RPCs fail because of MTU changes; this seems like an uncommon event in either case, and we can't completely eliminate this issue until support for PMTU is added. Once that happens, it won't matter whether the MTU problem is detected in homa_message_out_init or at packet transmission time. And, in any event, new RPCs will get the correct MTU so any RPC lossage will be temporary. Given this, I decided to drop the check in homa_message_out_init, hoping to limit code clutter and inefficiency from unneeded checks. Let me know if there's something that I'm missing?…-John-On Tue, Oct 6, 2026 at 5:04 PM breakertt ***@***.***> wrote: *breakertt* left a comment (PlatformLab/HomaModule#112) <#112 (comment)> In case you already started looking at the push above, I had a change of heart and decided to add route validation to homa_route_get ("lookup", as you put it). With this addition, I don't think that homa_message_out_init needs to perform an additional validation, so I removed that code. This will result in route validations in other situations that weren't covered by my first push, such as in homa_xmit_unknown and homa_need_ack_pkt. Hi John, After checking I think it may be incorrect to remove validation in homa_message_out_init. When server replies the response, it may been already a while since the server rpc allocation (hence homa_route_get) and route status may changed. Cheers, Tianyi — Reply to this email directly, view it on GitHub <#112?email_source=notifications&email_token=ACOOUCXVO7D4RFFVK3L6GED5SWCARA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTMMBSG44DCMRXG4YKM4TFMFZW63VHMNXW23LFNZ2KKZLWMVXHJLDGN5XXIZLSL5RWY2LDNM#issuecomment-6027812770>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/ACOOUCSIHDGU6MB4JPVSE735SWCARAVCNFSNUABFKJSXA33TNF2G64TZHMYTGOJWGM3DANZXHNEXG43VMU5TKNRZHA4DQNRRGI32C5QC> . Triage notifications, keep track of coding agent tasks and review pull requests on the go with GitHub Mobile for iOS <https://github.com/notifications/mobile/ios/ACOOUCWHL3NMR33FCFZHMND5SWCARA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTMMBSG44DCMRXG4YKM4TFMFZW63VHMNXW23LFNZ2KKZLWMVXHJKTGN5XXIZLSL5UW64Y> and Android <https://github.com/notifications/mobile/android/ACOOUCTTSAL636NTLRFE7WD5SWCARA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTMMBSG44DCMRXG4YKM4TFMFZW63VHMNXW23LFNZ2KKZLWMVXHJLTGN5XXIZLSL5QW4ZDSN5UWI>. Download it today! You are receiving this because you commented.Message ID: <PlatformLab/HomaModule/issues/112/6027812770 ***@***.*** com>
- Hi John, That sounds fair to me, thank you! Cheers, Tianyi
When a client-server route goes stale, the client must silently blackhole messages.
Expected behavior
Client can fail loudly while sending messages when destination route is stale.
Actual behavior
Client must silently blackhole messages.
Reproduction
Git checkout against 53f7963.
A simple client-server setup where the client continuously sends messages to the server is enough.
This assumes the reproduction is done locally via veths/namespace networking, so step 3,5,7 differ slightly for testing with a physical link.
1. load kernel module
2. make server + client binaries
3. set up networking
4. start server + client
5. disrupt networking
with a physical link, just yoink the cable...
6. inspect logs
After the server died (
[SERVER DIED]in client.log), the client still reports successful sends.7. teardown + cleanup
Notes
Initial findings were already mentioned in #107. I falsely described there being two failure modes (
EINVALin addition to blackhole), however I mixed that up with the fix I already applied... Sorry...