Handling incomplete cases

22 views
Skip to first unread message

Sam O'Dell

unread,
Sep 23, 2026, 9:31:15 PMSep 23
to spOccupancy and spAbundance users
All,

I'm just getting started with occupancy modeling and have run into a road block I know I can fix... with some guidance.

What is the best way to handle incomplete cases? I have a dataset with some missing values for detection covariates where observers did not record all of the necessary information (e.g., start time of survey, or temperature) despite conducting a survey.

Error in msPGOcc(), : error: some elements in det.covs have missing values where there is an observed data value in y. Please either replace the NA values in det.covs with non-missing values (e.g., mean imputation) or set the corresponding values in y to NA where the covariate is missing.

These gaps do not perfectly overlap each other across all surveys, detection covariates, or detection histories. So far, I have only attempted working with Julian date and time of day as detection covariates. It strikes me that methods of filling missing data (e.g., mean imputation) would be inappropriate for these variables (maybe you have some other idea?). Any suggestions are much appreciated!

Sincerely,
Sam O'Dell

Jeffrey Doser

unread,
Sep 28, 2026, 5:37:27 AM (11 days ago) Sep 28
to Sam O'Dell, spOccupancy and spAbundance users
Hi Sam, 

Sorry for the delay in response. There is no single best way to handle situations like the one you describe where observers did not collect all the relevant detection covariates during a survey. There are of course limitations to any imputation approach, as it has the potential to bias results if the imputation is not accurate. Mean imputation is arguably the most simple approach, but you aren't necessary restricted to taking the mean across the whole data set and instead you could take the mean for a subset of the data that is perhaps more closely tied to the missing observation. For example, if there is a detection variable measured over time for repeated visits at a site and some of them are missing, you could take the mean of the variable when it is measured at the site to fill in the values for when it is missing. There are a variety of more advanced missing data imputation approaches, the most complex of which would seek to fit some model to the covariate that has missing data to try and predict the missing values, and then you would use those model-predicted values as inputs into the occupancy model. Some Google searches regarding missing data imputation approaches in ecology turns up a variety of useful references. This recent paper by Mike Dumelle gives a broad overview of the statistical challenges with data imputation that could be of interest. 

Jeff

--
You received this message because you are subscribed to the Google Groups "spOccupancy and spAbundance users" group.
To unsubscribe from this group and stop receiving emails from it, send an email to spocc-spabund-u...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/spocc-spabund-users/3d7e3eae-5b94-4417-b08f-bb9d4d69de4cn%40googlegroups.com.


--
Jeffrey W. Doser, Ph.D.
Assistant Professor
Department of Forestry and Environmental Resources
North Carolina State University
Reply all
Reply to author
Forward
0 new messages