New c6620 configuration is problematic

19 views
Skip to first unread message

john.ou...@gmail.com

unread,
Aug 24, 2026, 12:52:43 AMAug 24
to cloudlab-users
The recent change to the c6620 cluster, where it is now divided between two switches, is problematic for anyone trying to do network-intensive experimentation:

* The bandwidth between the two switches is only 1600 Gbps of bandwidth, which is not enough to support uniform traffic patterns that involve both switches. For example, if all of the c6620 nodes on both switches are used for an experiment driving the uplinks at full capacity with uniform traffic patterns, half of the traffic would cross between the switches, which would require 3200 Gbps.

* Even running smaller experiments (such as a 50-node experiment with Homa) I am seeing significant performance degradation because the nodes are split across the two switches.

* There seems to be no effective way for me to get all of my nodes on one switch. The suggested `link.setNoInterSwitchLinks()` approach only works if there are enough free nodes to choose some and not others. But I can typically only get nodes with a reservation, and it doesn't guarantee that the reserved nodes are all on the same switch. Thus to get 50 nodes on one switch I would have to reserve 100 nodes, then only use 50 of them.

Is there any way you could modify the reservation system so that we can request that all of the nodes of our reservation are on the same switch? I suspect that this would be useful on clusters other than c6620s too.

-John-

Aleksander Maricq

unread,
Aug 26, 2026, 12:19:33 AMAug 26
to cloudlab-users
I know that you (probably) figured out the true reason for your current problems over in the other thread, but just to address some of these points at a higher level beyond that:
  • 8x400Gb was acceptable for the original Z9664 switch because we assumed that most experiments using the c6620 nodes wouldn't have a substantial amount of traffic crossing over to other hardware types.  We split the existing uplink into two 4x400Gb trunks when we moved to the two new switches, but yes, a very glaring problem with that is the fact that much more c6620 traffic needs to cross into the core than before.  I agree that this is not a great situation.  We have port capacity on each side for adding to the uplinks, it's just a question of whether we have the money at the moment to do that.
  • Another thing we're looking to do is to wire the 2nd 100Gb interface on each node back to the old Z9664F-ON switch.  This would allow people who want big LANs and don't care about PTP to be able to set those up again.  This is another thing that will depend on availability of money.
  • Yes, the "link.setNoInterSwtichLinks()" approach falls apart with higher node counts.  Unfortunately, due to the way the reservation system works, I suspect that what you're suggesting (requesting everything to be on the same switch) would be impractical at best and outright impossible at worst.  Someone who knows more about the reservation system can chime in and either corroborate or correct me.
Best,
 - Aleks
Reply all
Reply to author
Forward
0 new messages