Setting population size and distribution of fitness effect in rice evolution simulation

31 views
Skip to first unread message

­박현빈 / 학생 / 식물생산과학부

unread,
Aug 31, 2026, 4:12:07 AM (5 days ago) Aug 31
to slim-discuss
Hello.

I am interested in simulating the evolution of modern rice (Oryza sativa) using SLiM. I am having difficulty setting the DFE and population size for the simulation, so I would like to ask the slim-discuss users for their advice.

1) The simulation will be performed using 17 representative rice varieties. I want to calculate the population size based on nucleotide diversity and mutation rates, so I plan to calculate nucleotide diversity using VCFtools with the VCF files of the 17 varieties as input. I would like to ask if this approach is appropriate, or if I should refer to other papers that have calculated nucleotide diversity for a wider range of rice populations.

2) Since 80% of mutations occurring in the coding region of rice genes are slightly harmful, I plan to set the distribution of fitness based on the gamma distribution. I searched through reference papers to determine the mean and shape parameters of the gamma distribution, but I could not find appropriate values. If anyone has experience performing simulations on rice, I would like to ask about the appropriate method for setting the DFE.

Any help would be really appreciated.

Best wishes,

Hyunbin

­박현빈 / 학생 / 식물생산과학부

unread,
Aug 31, 2026, 7:30:29 AM (5 days ago) Aug 31
to slim-discuss
I have one more question.

The simulation I am currently planning operates by inputting VCF files of actual rice varieties and assigning different fitness effects based on changes in the nucleotide sequences of specific gene regions (e.g., C for survival gain, A for survival disadvantage). However, I do not fully understand the advantages of using a nucleotide-based model compared to using a basic model (e.g., wt for survival gain, m1 mutation for survival disadvantage), as I have found very few relevant examples of model application.

I would like to ask for advice on what advantages my current simulation implementation plan offers over a simpler approach (one that does not reflect actual nucleotide sequences), and if this plan alone does not provide any advantages, what elements I should add.

2026년 8월 31일 월요일 오후 5시 12분 7초 UTC+9에 ­박현빈 / 학생 / 식물생산과학부님이 작성:

Peter Ralph

unread,
Aug 31, 2026, 1:32:56 PM (5 days ago) Aug 31
to slim-discuss
Hi, Hyunbin! This is a big topic, but one that comes up fairly frequently on this list. I'll be brief, others may have more to say. If anyone knows a good review article on the topic of "how do I make best-guess choices for a popgen simulation, and which ones matter", say so? (Or, write one?)

A lot of the answers depends on your goals. You say you're simulating "the evolution of modern rice", so I'll take the goal to be to create a simulation that looks, more or less, like our current understanding of modern rice evolution. And, I'll assume you're just looking at O. sativa.

(1) What should you use for the population size? Well, choosing this based on heterozygosity of your varieties sounds like a good start? For the purpose of simulating evolution of modern rice, one important thing would be to get roughly the right amount of genetic diversity that was available at the time. By taking modern heterozygosity between these 17 breeds, you're assuming that genetic diversity available at the time  is similar to what it is today (and that those 17 breeds aren't particularly closely related). If rice went through a significant bottleneck then you'd probably want a larger ancestral population size (based on wild rice relatives maybe? Or just a wild guess?) and then a smaller size for domestication.

(2) What deleterious distribution of fitness effects (DFE) should you use? The easiest would be to use a Gamma and take parameters estimated for some other species (whatever species you can find that seems least dissimilar?). Another approach here would be to use a program that estimates the DFE, reviewed here: https://academic.oup.com/mbe/article-abstract/42/11/msaf236/8263265

(3) Should I use the actual genome sequence and a nucleotide model? I don't have a strong opinion on this? It really shouldn't matter - AFAIK - at all whether you use the standard mutation model (with stacking, etc) or a nucleotide model. Both are approximations, and the differences are going to be small (since most variant sites have only a single mutation). However, assigning different DFEs in different annotated regions (coding,  introns, etc) is probably a fine idea (see the paper above) and not too hard.

Another question you might have asked is whether you should use the VCFs to try to match allele frequencies. You didn't ask so probably you know, but the short answer here is no:  your simulation you certainly won't have the same set of SNPs as in the real data (since this is a random outcome!).

Happy SliMulating!

 Peter 

From: slim-d...@googlegroups.com <slim-d...@googlegroups.com> on behalf of ­박현빈 / 학생 / 식물생산과학부 <vince...@snu.ac.kr>
Sent: Monday, August 31, 2026 4:30 AM
To: slim-discuss <slim-d...@googlegroups.com>
Subject: Re: Setting population size and distribution of fitness effect in rice evolution simulation
 
I have one more question. The simulation I am currently planning operates by inputting VCF files of actual rice varieties and assigning different fitness effects based on changes in the nucleotide sequences of specific gene regions (e. g. , C
ZjQcmQRYFpfptBannerStart
You haven't previously corresponded with this sender.
Use caution with links and attachments. Learn more about this email warning tag.
 
ZjQcmQRYFpfptBannerEnd
--
SLiM forward genetic simulation: http://messerlab.org/slim/
---
You received this message because you are subscribed to the Google Groups "slim-discuss" group.
To unsubscribe from this group and stop receiving emails from it, send an email to slim-discuss...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/slim-discuss/aae5224c-b906-46ee-af1f-69df713e851fn%40googlegroups.com.

Andrew Kern

unread,
Aug 31, 2026, 1:36:13 PM (5 days ago) Aug 31
to slim-discuss
It's also worth pointing out that stdpopsim could be a good place to start for O. sativa demographic history: https://popsim-consortium.github.io/stdpopsim-docs/stable/catalog.html#sec_catalog_OrySat

­박현빈 / 학생 / 식물생산과학부

unread,
Sep 1, 2026, 8:47:28 AM (4 days ago) Sep 1
to slim-discuss
Thank you Peter and Andrew!

Your advice was a great help in configuring the simulation. I will share the results again if I succeed in obtaining meaningful outcomes later.

2026년 9월 1일 화요일 오전 2시 36분 13초 UTC+9에 Andrew Kern님이 작성:

Gregor Gorjanc

unread,
Sep 1, 2026, 9:03:16 AM (4 days ago) Sep 1
to slim-discuss
Hi,

In addition to other fine comments, I can share that from my experience modelling change in Ne over time is a very important aspect for simulating genomes of species used in agriculture due to domestication and selection (much larger Ne in the past than today with a roughly linear drop on log-Ne-vs-log-time plane).

One challenge here is that domestication induced a lot of selection, probably a mix of selecting for traits driven by a few major effect loci (generating sweeps) and of selecting for traits driven by many minor effect loci, but also drift! Much of this selection will be directional for some traits and unclear to me how much stabilising selection is/was there (surely for some traits!?). Both of these processes show up as a reduction in Ne, so Ne estimates from some demographic models might already account for the effect of both of these processes, though directional selection will generate directional changes in allele frequency while drift does not have a direction. I don't have a good experience with separating these two processes, but there are some demographic model estimation tools that estimate both! I think such models are (yet) implemented in stdpopsim for rice (maybe you would be open to take a stab at this!?).

You did not mention if all your lines are from one subspecies and/or population - do you need to consider multiple populations in your simulation, say indica vs japonica rice?

One option would be to estimate a demographic model from your data and then simulate according to your estimated demographic model! There is now a host of tools to do this, dadi, moments, momi, GADMA, etc. That would be the closest thing you could do to "get close to your samples".

As you see, there is a couple of unknowns you would have to work with, leading to a nice project;)

gg

­박현빈 / 학생 / 식물생산과학부

unread,
Sep 3, 2026, 2:02:29 AM (2 days ago) Sep 3
to slim-discuss
Thank you Gregor!
The variety I will be using contains an equal amount of indica and japonica, and the core of the simulation is actually observing the differences between the two subspecies. I plan to study the tools you mentioned further and find them useful. Thank you again.


2026년 9월 1일 화요일 오후 10시 3분 16초 UTC+9에 Gregor Gorjanc님이 작성:

Gregor Gorjanc

unread,
Sep 3, 2026, 2:37:21 AM (2 days ago) Sep 3
to slim-d...@googlegroups.com
The demographic model mentioned before (https://popsim-consortium.github.io/stdpopsim-docs/stable/catalog.html#sec_catalog_orysat_models_bottleneckmigration_3c07) should be a very good starting point for this - it includes rufipogon, indica, and japonica. Some colleagues say it’s not a perfect model, but simulates reasonable SFS.

With regards,
  Gregor


--
SLiM forward genetic simulation: http://messerlab.org/slim/
---
You received this message because you are subscribed to the Google Groups "slim-discuss" group.
To unsubscribe from this group and stop receiving emails from it, send an email to slim-discuss...@googlegroups.com.

­박현빈 / 학생 / 식물생산과학부

unread,
Sep 3, 2026, 9:24:18 PM (2 days ago) Sep 3
to slim-discuss
Thank you for your advice. The model presented in stdpopsim seems suitable for the initial setup of the simulation I intended.

2026년 9월 3일 목요일 오후 3시 37분 21초 UTC+9에 Gregor Gorjanc님이 작성:
Reply all
Reply to author
Forward
0 new messages