Subsetting catalog loci before sstacks to reduce memory load

13 views
Skip to first unread message

Brian Dorsey

unread,
Aug 29, 2026, 2:40:09 PM (12 days ago) Aug 29
to Stacks
Hello all,

I have a ddRAD data set with over 900 individuals across 33 populations in ~8 species sequenced on 2 lanes of a NovaSeq. I created a catalog with about 420 of these to reduce the number of loci. However, gstacks still runs out of memory and crashes on our server with 768GiB of RAM. There are 6588728 loci in the catalog. I know I could further reduce the individuals used to create the catalog but I am interested in rare and private alleles so would prefer to keep as many as possible. 

My question is whether I can manually subset the loci in the catalog file and then rerun from sstacks on. I was thinking to choose the first X% or to randomly sample the loci and write to a new set of catalog files. Is this possible and is it a reasonable idea? If not, are there any suggestions for dealing with this situation?

Thanks very much for any advice.

Brian D.
Reply all
Reply to author
Forward
0 new messages