Probabilistic models for CRISPR spacer content evolution

Kupczok, Anne and Bollback, Jonathan P (2013) Probabilistic models for CRISPR spacer content evolution. BMC Evolutionary Biology, 13. ISSN 1471-2148

[img] Text
1471-2148-13-54.pdf - Published Version
Available under License Creative Commons Attribution.
Download (506Kb)
Official URL:


Background: The CRISPR/Cas system is known to act as an adaptive and heritable immune system in Eubacteria and Archaea. Immunity is encoded in an array of spacer sequences. Each spacer can provide specific immunity to invasive elements that carry the same or a similar sequence. Even in closely related strains, spacer content is very dynamic and evolves quickly. Standard models of nucleotide evolutioncannot be applied to quantify its rate of change since processes other than single nucleotide changes determine its evolution.Methods We present probabilistic models that are specific for spacer content evolution. They account for the different processes of insertion and deletion. Insertions can be constrained to occur on one end only or are allowed to occur throughout the array. One deletion event can affect one spacer or a whole fragment of adjacent spacers. Parameters of the underlying models are estimated for a pair of arrays by maximum likelihood using explicit ancestor enumeration.Results Simulations show that parameters are well estimated on average under the models presented here. There is a bias in the rate estimation when including fragment deletions. The models also estimate times between pairs of strains. But with increasing time, spacer overlap goes to zero, and thus there is an upper bound on the distance that can be estimated. Spacer content similarities are displayed in a distance based phylogeny using the estimated times.We use the presented models to analyze different Yersinia pestis data sets and find that the results among them are largely congruent. The models also capture the variation in diversity of spacers among the data sets. A comparison of spacer-based phylogenies and Cas gene phylogenies shows that they resolve very different time scales for this data set.Conclusions The simulations and data analyses show that the presented models are useful for quantifying spacer content evolution and for displaying spacer content similarities of closely related strains in a phylogeny. This allows for comparisons of different CRISPR arrays or for comparisons between CRISPR arrays and nucleotide substitution rates.

Item Type: Article
DOI: 10.1186/1471-2148-13-54
Uncontrolled Keywords: Maximum likelihood, Bacterial immunity, CRISPR/Cas, Microbial genome evolution
Subjects: 500 Science > 570 Life sciences; biology > 576 Genetics and evolution
Research Group: Bollback Group
SWORD Depositor: Sword Import User
Depositing User: Sword Import User
Date Deposited: 23 Dec 2015 11:14
Last Modified: 05 Sep 2017 14:27

Actions (login required)

View Item View Item