Skip to content

client must silently blackhole messages on stale destinations #112

Description

@tiararodney

When a client-server route goes stale, the client must silently blackhole messages.

Expected behavior

Client can fail loudly while sending messages when destination route is stale.

Actual behavior

Client must silently blackhole messages.

Reproduction

Git checkout against 53f7963.

A simple client-server setup where the client continuously sends messages to the server is enough.

This assumes the reproduction is done locally via veths/namespace networking, so step 3,5,7 differ slightly for testing with a physical link.

1. load kernel module

2. make server + client binaries

make -C util server
sh <<- EOF
	cd util
	cc -x c -Wall -g -o 'client' - -I'..' << 'CC'
	#include <errno.h>
	#include <stdio.h>
	#include <stdlib.h>
	#include <string.h>
	#include <unistd.h>
	#include <arpa/inet.h>
	#include <sys/socket.h>
	#include <netinet/in.h>
	#include "homa.h"
	
	int main(int argc, char **argv)
	{
		if (argc != 3) { fprintf(stderr, "usage: %s <host> <port>\n", argv[0]); return 1; }
	
		int fd = socket(AF_INET, SOCK_DGRAM, IPPROTO_HOMA);
		if (fd < 0) { perror("socket"); return 1; }
	
		struct sockaddr_in dest = { .sin_family = AF_INET, .sin_port = htons(atoi(argv[2])) };
		if (inet_pton(AF_INET, argv[1], &dest.sin_addr) != 1) { fprintf(stderr, "bad addr\n"); return 1; }
	
		char buf[10000] = {};
	((int *)buf)[0] = sizeof(buf);
	((int *)buf)[1] = 0;
		for (unsigned long n = 1; ; n++) {
			struct homa_sendmsg_args sa = {};
			struct iovec iov = { .iov_base = buf, .iov_len = sizeof(buf) };
			struct msghdr msg = { .msg_name = &dest, .msg_namelen = sizeof(dest),
				.msg_iov = &iov, .msg_iovlen = 1,
				.msg_control = &sa, .msg_controllen = 0 };
	
			if (sendmsg(fd, &msg, 0) < 0)
				fprintf(stderr, "#%lu FAIL %s (errno %d)\n", n, strerror(errno), errno);
			else
				fprintf(stderr, "#%lu ok\n", n);
			usleep(500000);
		}
	}

	CC
EOF

3. set up networking

NS=homatestitest
VCLIENT=veth-homah
VSERVER=veth-homap
IPCLIENT=10.80.0.1
IPSERVER=10.80.0.2
ip netns add "$NS"
ip link add "$VCLIENT" type veth peer name "$VSERVER" netns "$NS"
ip addr add "$IPCLIENT/24" dev "$VCLIENT"
ip link set "$VCLIENT" up
ip -n "$NS" addr add "$IPSERVER/24" dev "$VSERVER"
ip -n "$NS" link set "$VSERVER" up
ip -n "$NS" link set lo up

4. start server + client

LOGDIR=$(mktemp -d)
ip netns exec "$NS" util/server --verbose --port 4000 >$LOGDIR/server.log 2>&1 & JID_SERVER=$!
util/client "$IPSERVER" 4000 >$LOGDIR/client.log 2>&1 & JID_CLIENT=$!
sleep 4

5. disrupt networking

ip netns del "$NS"
ip link del "$VCLIENT"
kill $JID_SERVER
wait $JID_SERVER
echo "[SERVER DIED]" >> $LOGDIR/client.log

with a physical link, just yoink the cable...

6. inspect logs

dmesg 2>/dev/null | grep -i homa
cat $LOGDIR/client.log

After the server died ([SERVER DIED] in client.log), the client still reports successful sends.

7. teardown + cleanup

kill $JID_SERVER $JID_CLIENT
wait $JID_SERVER $JID_CLIENT
ip netns del "$NS" || true
ip link del "$VCLIENT" || true
rm -r $LOGDIR

Notes

Initial findings were already mentioned in #107. I falsely described there being two failure modes (EINVAL in addition to blackhole), however I mixed that up with the fix I already applied... Sorry...

Activity

  1. breakertt commented on Oct 4, 2026

    @breakertt
    Contributor

    Hi Tiara,

    I roughly looked into this, and I think there are two separate problems here, plus a few things in the test setup that make them hard to see.

    Two problems

    Problem 1: a stale dst is reused on a route-cache hit

    homa_route_get() returns a cached homa_route on an rhashtable hit without calling dst_check() (homa_peer.c:441-451). This came in with f30a0ef, which introduced homa_route. Its design is that an invalid dst means discarding the route and re-establishing it. Today only homa_route_validate() does that, and it runs on the transmit path (homa_outgoing.c:404, :464), not on lookup.

    When the egress device is removed, the kernel marks its dsts DEAD and repoints them at blackhole_netdev (MTU 68, output = dst_discard_out). The next sendmsg reuses that dst:

    1. homa_message_out_init() computes max_seg_data = 68 - 20 - 56 = -8.
    2. homa_tx_copy_from_user() keeps writing segment headers until the buffer is full, then returns EINVAL.
    3. Nothing is transmitted, so homa_route_validate() never runs and the stale route stays cached.

    Reproduced on 6.17.8 (IPv4, veth/netns as in your steps, default route present) with a request/response client: after the veth is deleted, every sendmsg returns EINVAL, indefinitely.

    To make the expected behaviour concrete, I briefly summarized how other kernel transports behave in the same situations, which may serve as a reference. The last two columns are what I think Homa should do, and what it does today. For each case, the table shows what sendmsg returns and what a later recvmsg returns.

    Case TCP UDP SCTP Expected for Homa (IMO) Homa today (53f7963)
    Cached dst invalid, peer alive (e.g. route moved to another interface) send returns bytes; route re-resolved at transmit; recv gets data sendmsg re-resolves inside the call, returns bytes; recvmsg gets data sendmsg returns bytes; route re-resolved at transmit; recvmsg gets data sendmsg re-resolves inside the call, returns 0; recvmsg gets the response No RPC outstanding: every sendmsg returns EINVAL, nothing to receive (reproduced). RPC outstanding: a RESEND re-routes first, then sendmsg returns 0 and the outcome depends on where the new route leads (reproduced)
    Cached dst invalid, re-lookup finds no route send returns bytes (queued); error surfaces later sendmsg returns -ENETUNREACH sendmsg returns bytes; packet dropped; later association-loss notification sendmsg returns -ENETUNREACH Cached peer: no re-lookup happens, so same as the row above. New peer: sendmsg fails with "couldn't find route for peer" (from code, not tested)
    Route valid, peer dead send returns bytes; after retransmits (~15 min by default) recv returns ETIMEDOUT sendmsg returns bytes; recvmsg blocks, no error sendmsg returns bytes; after path retransmits recvmsg gets SCTP_COMM_LOST sendmsg returns 0; recvmsg returns ETIMEDOUT after ~100 ms Same as expected: recvmsg returns ETIMEDOUT after 100–190 ms (reproduced)

    None of these transports fail sendmsg because a cached dst went stale, or because the peer is dead. sendmsg fails synchronously only when a fresh route lookup fails.

    Device removal is only one of the ways a cached dst can become invalid, just the most visible one. Below are the ones I found, with how the kernel invalidates the dst and what would happen if Homa keeps using the old one:

    Event How the kernel invalidates the dst Effect if Homa keeps using it
    Device unregistered dst marked DEAD, moved to blackhole_netdev (dst.c:148-153) negative max_seg_data → EINVAL, or silent drop (reproduced)
    Route added / changed / deleted rt_genid bump (fib_trie.c:1316, 1379, 1749); deleted route's dsts marked DEAD old egress / next hop kept
    Address added / removed, link up / down rt_genid bump (fib_frontend.c:1457, 1474, 1485, 1521, 1532) old source address or a down interface kept
    Local MTU change rt_genid bump (fib_frontend.c:1536) old MTU kept for segmentation
    First PMTU exception for a destination (learned by any protocol) shared nexthop dsts marked KILL (route.c:720-728) new PMTU not picked up
    ICMP redirect marked KILL (route.c:800) new next hop not used
    fib rules, nexthop objects, ARP, sysctls, IPsec policy rt_genid bump old lookup result kept

    Problem 2: Homa does not handle PMTU ICMPs

    The IP layer stores PMTU, per destination in the fib nexthop exception (fnhe_pmtu), and shares it across protocols. But it does not write that value itself when an ICMP arrives. icmp_unreach() only extracts the MTU and hands the error to the protocol's err_handler (icmp.c:846-850, 903-921). The transport is expected to call ipv4_sk_update_pmtu() / ip6_sk_update_pmtu().

    Homa sends with DF set by default (IP_PMTUDISC_WANT, ip_output.c:510), so a smaller-MTU hop drops the packet and replies with an ICMP. For comparison, here is roughly how TCP and UDP handle that ICMP, next to what Homa does today:

    Protocol On "Fragmentation Needed" / "Packet Too Big" PMTU updated Result
    TCP tcp_v4_err → tcp_v4_mtu_reduced → inet_csk_update_pmtu (tcp_ipv4.c:583-593) yes resegments and retransmits; connection continues
    UDP udp_err → ipv4_sk_update_pmtu (udp.c:982-983) yes later sends use the new MTU; EMSGSIZE reported with IP_RECVERR
    Homa, IPv4 homa_err_handler_v4 treats it as a generic DEST_UNREACH (homa_plumbing.c:1757-1788) no all RPCs to that host aborted; recvmsg returns EHOSTUNREACH
    Homa, IPv6 homa_err_handler_v6 does not handle PKT_TOOBIG (homa_plumbing.c:1803-1827) no ICMP ignored; large packets keep being dropped; recvmsg eventually returns ETIMEDOUT

    The two problems interact. Even when another protocol on the host has already recorded the new PMTU and the kernel has marked the shared dst KILL, Homa's cache hit (Problem 1) keeps the old dst. The Homa rows in this table are from reading the code; I haven't tested them.

    Other chaos your test setup brought, which may confuse you on this problem

    • The response length is 0 and the client never calls recvmsg. Every RPC therefore stays outstanding. Homa's timer sends RESENDs for outstanding RPCs every few ms, and that path calls homa_route_validate(), which replaces the stale route before the next sendmsg. As a result the EINVAL from Problem 1 never shows up with this client. With a client that waits for each response, it shows up on every send.
    • The host has a default route. Deleting the veth also removes the connected route for 10.80.0.0/24. When Homa looks up 10.80.0.2 again, the lookup matches the default route and succeeds, so Homa gets a perfectly valid dst that points at the default gateway. sendmsg returns success and the packets go to the gateway, which drops them because 10.80.0.2 is not reachable there. That is the "blackhole": Homa isn't hiding an error, the kernel really did give it a route. Without a default route, the lookup would fail and sendmsg would return -ENETUNREACH.
    • The link and the server are removed in the same step. That mixes two different events, and Homa reports them in different places. A broken route is handled inside the next sendmsg (re-resolve, or -ENETUNREACH if there is no route). A dead peer can't be known at sendmsg time at all, so it is reported later by recvmsg as ETIMEDOUT, about 100 ms after the response fails to arrive. Since the client never calls recvmsg, the second error never shows up. Testing them separately makes this clearer: remove only the link while the server stays reachable another way, or kill only the server and leave the network alone.
    • On the note in client must silently blackhole messages on stale destinations #112: the EINVAL described in What evidence is useful for reporting NIC support status? #107 is stock behaviour (Problem 1). It appears whenever no RPC is outstanding when the device goes away, so I don't think it came from your patch.
  2. johnousterhout commented on Oct 5, 2026

    @johnousterhout
    Member

    Thanks to both of you for all this analysis. I have not considered all of these conditions before (such as an MTU so small that max_seg_data becomes negative) so this is super-helpful. It may take me a bit before I can work through all of these. In the meantime I have a question for you, Tianyi.

    You said:

    it runs on the transmit path (homa_outgoing.c:404, :464), not on lookup.

    What do you mean by "lookup"?

  3. breakertt commented on Oct 5, 2026

    @breakertt
    Contributor
  4. johnousterhout commented on Oct 6, 2026

    @johnousterhout
    Member

    I believe that my latest push fixes 2 problems from this issue:

    • Routes get revalidated in homa_message_out_init before computing packet geometry for an RPC.
    • homa_message_out_init checks for unusuably-small MTUs and fails the RPC with EHOSTUNREACH.

    I have not yet addressed the issue of changes in MTU after an RPC has been initialized, or handling PMTU. This is a bigger issue that I'll try to schedule at some point in the future.

    In the meantime, Tianyi, if you have time and interest, I'd love for you to explore what happens when the MTU does change in the middle of an RPC (I assume that the only problem is if it gets smaller). I have no idea whether Homa handles this in a reasonable way (vs. just ending up with a hung RPC that never completes).

    For now I'm going to leave this issue open as a reminder.

  5. johnousterhout commented on Oct 6, 2026

    @johnousterhout
    Member

    In case you already started looking at the push above, I had a change of heart and decided to add route validation to homa_route_get ("lookup", as you put it). With this addition, I don't think that homa_message_out_init needs to perform an additional validation, so I removed that code. This will result in route validations in other situations that weren't covered by my first push, such as in homa_xmit_unknown and homa_need_ack_pkt.

  6. breakertt commented on Oct 7, 2026

    @breakertt
    Contributor

    In case you already started looking at the push above, I had a change of heart and decided to add route validation to homa_route_get ("lookup", as you put it). With this addition, I don't think that homa_message_out_init needs to perform an additional validation, so I removed that code. This will result in route validations in other situations that weren't covered by my first push, such as in homa_xmit_unknown and homa_need_ack_pkt.

    Hi John,

    After checking I think it may be incorrect to remove validation in homa_message_out_init. When server replies the response, it may been already a while since the server rpc allocation (hence homa_route_get) and route status may changed.

    Cheers,
    Tianyi

  7. johnousterhout commented on Oct 7, 2026

    @johnousterhout
    Member
  8. breakertt commented on Oct 7, 2026

    @breakertt
    Contributor
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions