Homa: The end of TCP for AI clusters [video]

(youtube.com)

46 points | by signa11 4 hours ago

10 comments

  • Animats 1 hour ago
    Homa has been around for a while. Here's the 2018 paper.[1]

    The core idea: When a message arrives at the sender’s transport module, Homa divides the message into two parts: an initial unscheduled portion (the first RTTbytes bytes), followed by a scheduled portion. The sender transmits the unscheduled bytes immediately, using one or more DATA packets. The scheduled bytes are not transmitted until requested explicitly by the receiver using GRANT packets.

    So it sends blind for short requests, then needs a go-ahead from the receiver. That's reasonable when the main application is a remote procedure call. It's reminiscent of QNX's networking protocol, which is also single packet message request/response but can also handle arbitrarily long messages.

    What makes this work today is that per-packet processing overhead in hardware switches is low vs. per-byte overhead. In early software driven switches, per-packet overhead tended to dominate, and sending small packets was very inefficient. In modern hardware switches, where FPGAs are doing the processing, the per-packet overhead is low enough that small packets are not inefficient.

    It's amusing that web stuff is so bloated today that any transaction under 1MB is considered "small". So this is not a suitable protocol for open web use.

    [1] https://people.csail.mit.edu/alizadeh/papers/homa-sigcomm18....

  • Veserv 24 minutes ago
    Homa is not a good design. [1]

    1. No way to detect whole RPC loss. Since there is no outer connection state, if every packet in the send-side of a RPC is lost then there is no way for a server to detect that it should issue a resend. The RPC is just lost to the ether. This affects small messages, like messages that fit in a single packet, more since there are fewer packets in the send-side.

    2. Related to the above, there is no builtin encryption support. So, if you want encryption then you need to layer it either above or below.

    3. Benchmarked performance is awful. The 60 kB average message case in [2] Table 4 takes 5(!) hyperthreads to average 20 Gbit/s. That is just 4 Gbit/s per hyperthread. Even a totally naive one-packet per system call network protocol design and implementation should get to ~8 Gbit/s per hyperthread. 30 Gbit/s per hyperthread is easy with just a little focus on performance.

    4. Despite all the performance design problems in QUIC (though still faster than Homa) it already solves basically every problem Homa is trying to solve in a much cleaner way. Stream IDs correspond to RPC IDs. Stream Max corresponds to Grants. Multiple streams under single Client allows prioritization.

    Except you do not randomly lose entire messages. You can compact small messages into packets. You get more precise RTT time allowing more accurate pacing/congestion calculations. You get builtin encryption. It survives ossified middleboxs. It has multiple ack frames/packets reducing ack overhead.

    The only real difference is that Homa uses explicit receiver Resend instead of implicit sender Resend. Except that actually consumes significantly more receiver resources in non-trivial loss scenarios especially due to Resend packets only supporting a single Resend span. It also incurs higher latency and has higher requirements on the entire lossy network path due to all the extra transiting data that you do not want to lose.

    And those are just serious problems off the top of my head after reloading the RFC into my head. I can come up with some more if needed.

    [1] https://github.com/johnousterhout/homa-rfc/blob/main/draft-o...

    [2] https://www.usenix.org/system/files/atc21-ousterhout.pdf

  • giovannibonetti 2 hours ago
    I remember listening to Jane Street’s Ron Minsky on their podcast talking about this a few months ago, how TCP becomes the bottleneck in AI clusters. As an electrical engineer, I remember that circuit switching gave away to packet switching due to very sparse usage of the network when there are many actors going through it. It is not very efficient, but that wide variety of traffic makes it hard to optimize it since the flow patterns are too dynamic. A good analogy with car traffic is that downtown there are so many cars going to a large variety of places, that traffic lights – as inefficient as they are – are a solution that at least works good enough.

    On the other hand, if the traffic follows a very predictable pattern, a custom implementation can be much more efficient. Specially nowadays machine learning can find much better solutions through reinforcement learning. And AI cluster data flow is much more predictable than what goes over the internet as a whole.

  • throw0101c 1 hour ago
    With regards to (~6:01) "congestion control is the responsibility of the sender" and "somehow we have to get the sending nodes to stop sending so fast". Does that not exist in Ethernet/RoCE?

    > Link Level Flow Control: InfiniBand uses a credit-based algorithm to guarantee lossless HCA-to-HCA communication. RoCE runs on top of Ethernet. Implementations may require lossless Ethernet network for reaching to performance characteristics similar to InfiniBand. Lossless Ethernet is typically configured via Ethernet flow control or priority flow control (PFC). Configuring a Data center bridging (DCB) Ethernet network can be more complex than configuring an InfiniBand network.[19]

    * https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet

    > A sending station (computer or network switch) may be transmitting data faster than the other end of the link can accept it. Using flow control, the receiving station can signal the sender requesting suspension of transmissions until the receiver catches up. Flow control on Ethernet can be implemented at the data link layer.

    * https://en.wikipedia.org/wiki/Ethernet_flow_control

    • teraflop 51 minutes ago
      Ethernet flow control doesn't really fix congestion except in very special cases.

      Consider a very simple topology:

          A         C
           \       /
            S1===S2
           /       \
          B         D
      
      Say hosts A and B are both sending data to C, as fast as they can, via switches S1 and S2 (which are connected via a high-speed link). And say the sum of these two flows is more than the capacity of the link to C.

      S2 is receiving packets destined for C faster than it can forward them, but sending an Ethernet pause frame from S2 to S1 is not a very productive way to alleviate the situation, because it also disrupts any traffic that would be bound for D. It just moves the bottleneck elsewhere and causes collateral damage.

    • lokar 58 minutes ago
      The issue is never really with the destination, it’s some link/switch along the way.

      When a link becomes saturated working out how to manage that is a hard problem.

      I have only used RoCE once at scale, it was really finicky. We would get big waves to pause frames that stalled everything.

    • wmf 57 minutes ago
      Flow control is better than nothing but it can cause congestion spreading and bufferbloat. QCN, Falcon, and Ultra Ethernet provide much better congestion control for RoCE but they also require newer hardware compared to Homa.
  • adastra22 2 hours ago
    Please don’t make the primary link a video.
  • wmf 1 hour ago
    I've been hearing about Homa for years and wondered what's new. I found a changelog in the readme of the git repo: https://github.com/PlatformLab/HomaModule
  • jMyles 2 hours ago
    I wonder if, as LLMs get accustomed to using homa or some other optimized protocol for, as the article lists, "chores such as weight gradients, model weights, KV cache entries, and checkpoints", whether we'll start to see TCP as a bottleneck for their post-trained interactions as well, for many of the reasons.
  • almost_usual 3 hours ago
    • dang 2 hours ago
      Thanks, those are great! Added to toptext as well.
  • paradiselord-de 3 hours ago
    [flagged]