I am a masters student at TU Delft and I want to replicate the system in "Looking to listen at the Cocktail Party", but there is some information missing about the artificial dataset creation. The authors told that they added non-speech background noise from AudioSet to clean speech videos. But they did not mention what kind of "non-speech background noise" did they use. AudioSet contains millions of videos under 527 labels.
Performance of the system is highly dependent on training data. So, I feel kind of noise used to create training dataset will affect the performance of the system. What are the required type (labels) of noises that should be added to clean speech or should it be random. And what should be the optimum number of noisy audio that should be generated for each clean audio. Is it one noisy audio corresponding to one clean audio or more than that?
Thanks in advance.
Regards,
Riya Maan