CloudLab Utah c6620 downtime - August 6th

109 views
Skip to first unread message

Aleksander Maricq

unread,
Jul 21, 2026, 9:04:23 PMJul 21
to cloudlab-users
Hi all,

We are planning an all-day downtime of the c6620 machines at CloudLab Utah on August 6th from 9 AM to 6 PM MDT.  Existing experiments using the c6620 machines will not be able to be extended past August 6th at 9 AM.  None of the other hardware types at CloudLab Utah should be affected by this.

During this time, we will be moving the c6620 100Gb experiment net wires to new switches, which will allow the 100Gb interfaces on the c6620 machines to be synchronized using PTP.  This feature had been planned since the beginning of this expansion, but unforeseen limitations with our current switch prevented us from rolling this out until we got new switches.  For reasons both technical and non-technical, we want to make this move as soon as we are able to.  While we have the nodes out of service, we will also be upgrading the firmware on all of the c6620 machines to their latest available versions.

The side-effect of this change is that the c6620 100Gb interfaces will have to be split between two switches, where they were all on the same switch before.  That means that some pairs of nodes will have more switch hops between them than before, so if you want to avoid interswitch links in your experiment, you can specify this in your profile using `link.setNoInterSwitchLinks()`.

Let us know if you have any questions or concerns.

Best,
 - Aleks

John Ousterhout

unread,
Jul 22, 2026, 11:46:54 AMJul 22
to cloudla...@googlegroups.com
Hi Alexs,

Will the c6620 reconfiguration affect the switch changes we made to
enable priority queues for Homa? For example, are the new switches a
different type where I'll need to re-figure out how to enable the
priority queues?

-John-
> --
> You received this message because you are subscribed to the Google Groups "cloudlab-users" group.
> To unsubscribe from this group and stop receiving emails from it, send an email to cloudlab-user...@googlegroups.com.
> To view this discussion visit https://groups.google.com/d/msgid/cloudlab-users/53f6b0bd-26a8-4e2b-a84b-8a658fb72587n%40googlegroups.com.

Aleksander Maricq

unread,
Jul 22, 2026, 1:42:23 PMJul 22
to cloudlab-users
Hi John,

The new switches are two Z9432F-ONs instead of a Z9664F-ON.  The same configuration should be possible to put on these new switches, but we can work with you in advance of the move to make sure that's the case.

Best,
 - Aleks

John Ousterhout

unread,
Jul 22, 2026, 2:10:57 PMJul 22
to cloudla...@googlegroups.com
Sounds good. If you can apply the current patches to the new switches, then once the new switches are up I'll run a quick experiment to see if the priority queues are working. If not, we can coordinate to make whatever adjustments are needed.

-John-

Aleksander Maricq

unread,
Aug 6, 2026, 5:51:42 PMAug 6
to cloudlab-users
Hi all,

We have concluded our work for this downtime.  Let us know if you run into any issues.

Best,
 - Aleks

John Ousterhout

unread,
Aug 17, 2026, 2:55:43 PMAug 17
to cloudla...@googlegroups.com
Hi Alexs,

I now have an experiment running on the c6620 cluster, and it appears to me that priorities are not working. Can you double-check to make sure you applied the switch commands to enable priority queues?

Thanks.

-John-


John Ousterhout

unread,
Aug 21, 2026, 12:58:51 AMAug 21
to cloudla...@googlegroups.com
Hi Aleks (sorry for misspelling your name in my previous email):

After exploring things further, it appears to me that some of the c6620 nodes are correctly configured for priorities and others are not. For example, in my current experiment (https://www.cloudlab.us/status.php?uuid=79312f89-0e45-4585-ab0d-ef160ca7768f) the TOR(s) do not correctly support priorities for the following nodes:

er006
er013
er073
er027
er002
er074
er016
er003
er012
er030
er004
er011
er017
er009

Priorities seem to be working for all the other nodes except node0 (which I haven't been able to test). All of the nodes that work seem to be attached to switch cl4-expt5, and nodes that don't work seem to be attached to switch cl4-expt4.

-John-

cloudlab-users

unread,
Aug 22, 2026, 4:27:52 PMAug 22
to cloudlab-users
I swore I answered this...but I guess it must still be sitting in a buffer at work.

Aleks was on vacation all last week, and he asked me to answer your question. The two switches have identical settings, so if it works on one it _should_ work on the other. One possibility is that your observation point is on a node on cl4-expt5? Because we did not enable QOS on the link between the switches via the core experiment switch. Is it not working between two nodes that are both on cl4-expt4?

John Ousterhout

unread,
Aug 22, 2026, 5:12:59 PMAug 22
to cloudla...@googlegroups.com
Homa does not need QOS on the links between switches, so I don't think that explains what I'm seeing.

I've run some more experiments, and I see that if I run my benchmarks using only nodes on switch cl4-expt4, or only nodes on switch cl4-expt5, I get identical (and good) performance with both switches. However, if I run benchmarks with all 50 nodes, those on cl4-expt5 perform worse than those on cl4-expt4. This suggests that the switches are configured properly for priorities.

I'm now wondering if this is related to the interconnect between the switches. My particular experiment happens to have 15 nodes under cl4-expt4 and 35 under cl4-expt5. In my benchmarks, each node spreads its traffic evenly across all the other nodes. This means that the nodes under cl4-expt4 must use the switch-to-switch interconnect for most of their requests, whereas most of the requests made by nodes under cl4-expt5 are to other nodes under the same switch.

What is the topology of the c6620 cluster? In particular, how many nodes are under each switch and what is the connectivity between the switches? Is there full bisection bandwidth? If the switch-to-switch links are getting congested, that could explain the performance I'm seeing.

-John-

cloudlab-users

unread,
Aug 22, 2026, 6:32:34 PMAug 22
to cloudlab-users
Ah yes. There are 64 nodes on one switch (expt4) and 68 on the other (expt5).
The interconnect is 4 x 400Gb from each to the core switch, so that is probably
what you are running up against. I don't think we have enough ports on the
core switch to handle more.

John Ousterhout

unread,
Aug 22, 2026, 7:33:59 PMAug 22
to cloudla...@googlegroups.com
After doing some math, I don't think that quite explains what I'm seeing.

Each node in my experiment is generating 80 Gbps of outbound traffic and receiving 80 Gbps of inbound traffic.

On cl4-expt4 there are 15 nodes, so the fraction of traffic that passes through the core switch is 35/50. Thus the total core traffic is 15 * (35/50) * 80, which is 840 Gbps.

On cl5-expt4 there are 35 nodes, and the fraction of traffic that passes through the core switch is 15/50. Thus the total core traffic is 35 * (15/50) * 80, which is 840 Gbps.

According to your email there is 1600 Gbps of bandwidth to the core from each switch, so if I have done my math correctly the uplinks should not be overloaded (assuming I'm the only experiment using them, which I suspect is generally true).

Is it possible that traffic is not being balanced effectively across the 4x400 uplinks, thereby creating hotspots? Is it easy to get data from the switches on bandwidth utilization for each of the uplinks? If so, I'm wondering if it would be possible for me to start a long-running experiment and then have you read out the utilizations to make sure that all the uplinks are being utilized equally (it's still possible that there could be transient congestion that doesn't appear in longer-term readouts).

Also, suppose I want to eliminate this issue by using only nodes on the same switch. Is there a way for me to request this in reservations? Otherwise the `link.setNoInterSwitchLinks()` approach suggested in an earlier email in this thread won't be helpful, because there won't generally be a choice of nodes once the reservation starts.

-John-

cloudlab-users

unread,
Aug 23, 2026, 12:02:46 AMAug 23
to cloudlab-users
Looking at the switch counters on the core switch for the two port channels to expt4 and expt5:

Over 2 weeks 4 days 07:26:05
               in packets            in octets                        out packets        out octets
1/1/45: 413013599311    2542746869856054  405682947271  2541513643917933
1/1/46: 410618566485    2539632901089467  406553610810  2543124939402011
1/1/47: 410685350293    2539243809290657  405731254961  2541998596414487
1/1/48: 412499136259    2541342116842177  403863758759  2539219454219424
errors: 199 corrected FCC on input, 33 discards on output

1/1/49: 407856735784    2543893123763185  410783833406  2541135259014456
1/1/50: 406909861289    2542287067417205  414071256549  2546693121545647
1/1/51: 402812633460    2535040071662637  410501010633  2541226921319665
1/1/52: 404360697785    2537519856192937  418764788399  2551839698816303
errors: 121 corrected FCC on input, 37 discards on output

they look pertty well balanced. The errors are for each port-channel as a whole.

On the experiment status page, there is a "portstats" tab which will give you switch port counters
for every node. Unfortunately, the presentations of the numbers is relative to the previous time
the counters were read. So each time you refresh that tab, you get numbers since the last refresh.
And...unfortunately, I already refreshed them once. But you should be able to: refresh the tab,
run your experiment, refresh the tab again to see what they look like for that run of the experiment.

Here is what the command line version tool showed when I read it:

                                  In  InUnicast InNUnicast        Out OutUnicast  OutNUcast                              
Port                          Octets    Packets    Packets     Octets    Packets    Packets                              
-------------------------------------------------------------------------------------------                              
dbox:4                     373272004 2465159093      18414  953012788 1386590374     595431                              
dbox:4                     373272004 2465159093      18414  953012788 1386590374     595431                              
node0:2                   1880401452  630282326       8374 2093519244  138169723      49850                              
node10:2                  1042378451 2234788372        986  405143283  669634794      51709                              
node11:2                  1542389267 3805561373        638  110717407   59108485      56456                              
node12:2                  1542978546 3770000041        564  909224354  114501497      56530                              
node13:2                  1325139903 3816284396        566 2579485717   52510964      56529                              
node14:2                  1386426425 3165933802        568 1573052805 3902083287      56522                              
node15:2                  3352015565 3252997761        571 3629998240 3852179350      56542                              
node16:2                  2618796525 1641015356        861 2976654464   89960245      51837                              
node17:2                  3525983612 1602307991        567 2007549698   82297948      52127                              
node18:2                    18928238 1613816511        621 2372122449   98265145      52076
node19:2                  1824115688 3345964101        894 1619516470 3934077787      56201
node1:2                   3514187069 2527628411        658  288816334 1216703105      52044
node20:2                   592855735 3269711107        580 1012883713 3839382139      56515
node21:2                  4037264344 2841726953        630 4252518201 3513653392      56482
node22:2                  2310800027 2828898000        926  956329663 3464077494      56166
node23:2                  3456298678 2733345513        889 1689020197 3455666855      56200
node24:2                    22262668 1036130874        947  225346055 3711734585      51747
node25:2                  3120517821 2872953339        592 2199030726 3524974802      56518
node26:2                  2187858236 2850566996        899 3670769276 3494272077      56214
node27:2                  3351682521 2867078655        670  838020326 3530154286      56436
node28:2                   955255757  976923909        993  111127990 3742811251      51702
node29:2                   316486279 2812283148        689 3286905930 3546817734      56419
node2:2                   1049527989  551259647        549  141516298 1372230114      56545
node30:2                  1537108242 2832456412        606 3325247070 3493614136      56508
node31:2                  3994618148 2300149512        895 1176816204 2989080317      56209
node32:2                   330959274  892541662        946  744632725 3615811376      51751
node33:2                  1895909297 2370717025        603 1363562216 3033769294      56490
node34:2                   884796029 2337192272        617 2130247305 3025905629      56475
node35:2                  2817279385 2412874428        602 3422211547 2987755329      56503
node36:2                  2123073331  810067319        607  640006670 3624282352      52090
node37:2                  3177267835 2363348009        599 1320603695 3049884017      56513
node38:2                   506330378 2356585278        977 3997325366 2977369822      56134
node39:2                   265626563 2375534168        600 3792005956 3052136794      56495
node3:2                   2523398462  416926110        834 2711411820 1088195633      56260
node40:2                   889977964 2397505849        613 3347302299 3024843257      56499
node41:2                  2337004395 2266894895        915 1096661038 2907717396      56175
node42:2                   901604193 2264244758        944 2821771879 2884272816      56161
node43:2                  3176192262  783181151        619 4016857162 3523295366      52075
node44:2                   468525360 2266243100        613 2885405102 2885399892      56492
node45:2                   909133525 2230971133        923  863738298 2926217330      56188
node46:2                  2158744842  709569365        681 2875711024 3503176202      51987
node47:2                  3526965827  726709576        702 2499346844 3540874396      51993
node48:2                  2697446094 2352572883        940 1857716579 2905377273      56153
node49:2                  3880472992 2408203284        937 1722441450 2988846758      56175
node4:2                   3895738478 2677393754        683  224652853 1137950758      52014
node50:2                  1131963036 2338379915        701 2127368103 3019898149      56391
node5:2                   1073700261  306726176        945 3798105238 1123502428      56149
node6:2                   2980353276  423154500        943  577857917 1207300308      56152
node7:2                   4162575857   13895273        636  973478245  597736301      56459
node8:2                   1676395719 2286777538        866 1426823679  661518691      51831
node9:2                   2896306664 2256956305        652 3810704426  655677564      52026

Quite a lot of imbalance there. The individual (per-port) error counters were all zero.

John Ousterhout

unread,
Aug 23, 2026, 12:31:52 AMAug 23
to cloudla...@googlegroups.com
Is there a way for me to query the switch counters for the links between the TORs and the core switch? Those are the ones that I'm most interested in (to see if the load is being balanced properly).

-John-

Mike Hibler

unread,
Aug 23, 2026, 9:56:20 AMAug 23
to cloudla...@googlegroups.com
Not at the current time, no. Since those links are shared between all
experiments, we don't give that info to users. Not even sure that we track it
in the database.
> f43f7fb3-efe7-42ee-9a6e-98e31a6b2726n%40googlegroups.com
> .
>
> --
> You received this message because you are
> subscribed to the Google Groups
> "cloudlab-users" group.
> To unsubscribe from this group and stop
> receiving emails from it, send an email to
> cloudlab-user...@googlegroups.com.
> To view this discussion visit https://
> groups.google.com/d/msgid/cloudlab-users/
> 790ec5c7-5d3c-44e1-96c3-a1ef46ea8b6cn%40googlegroups.com
> .
>
> --
> You received this message because you are subscribed to the
> Google Groups "cloudlab-users" group.
> To unsubscribe from this group and stop receiving emails
> from it, send an email to
> cloudlab-user...@googlegroups.com.
>
> To view this discussion visit https://groups.google.com/d/
> msgid/cloudlab-users/
> b73a0727-d605-44e9-8cfd-2d6eeb5d525dn%40googlegroups.com.
>
> --
> You received this message because you are subscribed to the Google
> Groups "cloudlab-users" group.
> To unsubscribe from this group and stop receiving emails from it,
> send an email to cloudlab-user...@googlegroups.com.
>
> To view this discussion visit https://groups.google.com/d/msgid/
> cloudlab-users/
> 37bbea6e-9a0e-4555-bb2a-067954a93b96n%40googlegroups.com.
>
> --
> You received this message because you are subscribed to the Google Groups
> "cloudlab-users" group.
> To unsubscribe from this group and stop receiving emails from it, send an
> email to cloudlab-user...@googlegroups.com.
> To view this discussion visit https://groups.google.com/d/msgid/
> cloudlab-users/65d605ff-dd88-4670-9422-a4465bc295ccn%40googlegroups.com.
>
> --
> You received this message because you are subscribed to the Google Groups
> "cloudlab-users" group.
> To unsubscribe from this group and stop receiving emails from it, send an email
> to cloudlab-user...@googlegroups.com.
> To view this discussion visit https://groups.google.com/d/msgid/cloudlab-users/
> CAGXJAmxsqMe2HH%3DYaJJoKTWC7DGot6hBWWufkMXyF4X-r7mx%2BA%40mail.gmail.com.

John Ousterhout

unread,
Aug 24, 2026, 12:42:54 AMAug 24
to cloudla...@googlegroups.com
That's unfortunate....

-John-

Reply all
Reply to author
Forward
0 new messages