I have a few post-meeting observations:
1) For the RegMixup part, the AAM loss regarding the virtual training examples unpacks into:
2) The best extractor from Table 1 (SSL-AASIST+AAM+RegMixup) is the one used in Table 2, but its zero-shot cosine numbers don't match across the two tables. This is because Table 2 explicitly fixes the fingerprinting condition to "nc-all" (non-common sentences, all available enrollment utterances per attack).