Massive upload to Dataverse

89 views
Skip to first unread message

Federico Yemurenko (ANII)

unread,
Sep 21, 2026, 9:35:30 AMSep 21
to Dataverse Users Community
Hi.  

I have launched a load of 70.000+ files to a dataverse, calling DVUploader in batches of 1.000 files with a 5 secs delay after each batch is completed.  All files are very small in size; .gif and .txt.  

The server is a VM with 8GB, 2CPUs, filesystem storage. Payara, Postgres and Solr are running on the same machine. It´s a quite small DV installation: 150 datasets published. 

The process has been running fine so far.  However, it is taking too long: after three three days, only 14 batches (14.000+) were uploaded.  It started by  loading 1 file a second but now it is taking about 30 secs to load a file.  Steady.

Is it reasonable ?
Any tip to improve it for when I have to do something similar next time ?

Regards

Federico Yemurenko

James Myers

unread,
Sep 21, 2026, 10:09:49 AMSep 21
to dataverse...@googlegroups.com

Dataverse does not yet handle that many files per dataset well. The recommended limits are from ~1K to at most ~10K files per dataset (the latter requiring use of the latest Dataverse versions and careful configuration – see info in https://guides.dataverse.org/en/latest/admin/big-data-administration.html ).

 

I would definitely suggest stopping your upload job – it gets difficult/slow just to delete that many data files.

 

The general recommended work-around is to upload zip files and configure the Zip Previewer – this limits the number of separate datafiles in the database while still allowing users to see/download individual files if they want.

 

There are also config options – listed in the guide above, to limit the number of files allowed per dataset to avoid someone adding more than you are prepared to support.

 

Hope that helps.

-- Jim

--
You received this message because you are subscribed to the Google Groups "Dataverse Users Community" group.
To unsubscribe from this group and stop receiving emails from it, send an email to dataverse-commu...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/dataverse-community/0b143121-0960-4e76-a2e0-26015d00228fn%40googlegroups.com.

Federico Yemurenko (ANII)

unread,
Sep 21, 2026, 10:15:53 AMSep 21
to dataverse...@googlegroups.com
Thanks a lot, Jim.

Regards.

FEDERICO YEMURENKO
Av. Italia 6201 - Edificio Los Nogales
Montevideo, Uruguay
T. (598) 2600 4411


Federico Yemurenko (ANII)

unread,
Sep 21, 2026, 10:44:45 AMSep 21
to Dataverse Users Community
Hi,

Following Jim´s recomendation I have stopped the process to devise a different approach.
Do you think that spliting the 70000 files to be published in several linked datasets would be a good idea ? 
But ......... it would mean about 70 linked datasets which is not user friendly; too much complexity for the publisher, don´t you think ?
I feel that it is first the dataset´s owner job to better think in how to group, organize his data in more than one huge, plain dataset.  

Regards  

Andreev, Leonid

unread,
Sep 21, 2026, 11:05:33 AMSep 21
to dataverse...@googlegroups.com
Hi Federico, 
To follow up on the excellent advice from Jim, rather than considering splitting the 70,000 files between several datasets, please do consider (or suggest to the author) his suggestion to repackage the files in fewer, but much larger zip bundles. 
It is difficult to think of a scenario where the author would have a reason to publish this many small files as separate, standalone files. (one possible exception that I can think of is if there is a very specific need to assign individual persistent ids, DOIs or Handles to every file). Will any user ever have a practical need to only download one or two files from a dataset like this? If the answer is no, packaging the files in as large as possible Zip bundles is a way to go. 
Do note that it will still be possible to download a couple of select files without downloading an entire huge zip bundle - the Zip Previewer that Jim mentions will allow that: 

> The general recommended work-around is to upload zip files and configure the Zip Previewer – this limits the number of separate datafiles in the database while still allowing users to see/download individual files if they want.

All the best,
-Leo


--
You received this message because you are subscribed to the Google Groups "Dataverse Users Community" group.
To unsubscribe from this group and stop receiving emails from it, send an email to dataverse-commu...@googlegroups.com.

Federico Yemurenko (ANII)

unread,
Sep 21, 2026, 2:46:41 PMSep 21
to dataverse...@googlegroups.com
Thanks, Leo.

Regards

FEDERICO YEMURENKO
Av. Italia 6201 - Edificio Los Nogales
Montevideo, Uruguay
T. (598) 2600 4411

You received this message because you are subscribed to a topic in the Google Groups "Dataverse Users Community" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/dataverse-community/pzSjF1YPaJw/unsubscribe.
To unsubscribe from this group and all its topics, send an email to dataverse-commu...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/dataverse-community/CAGSsysttx0NJNAqJyDKqcU7My3S2rSYxpomF0yDjVKzm0JOXPQ%40mail.gmail.com.
Reply all
Reply to author
Forward
0 new messages