A circular permutation is a relationship between proteins whereby the proteins have a changed order of amino acids in their peptide sequence. The result is a protein structure with different connectivity, but overall similar three-dimensional (3D) shape. In 1979, the first pair of circularly permuted proteins – concanavalin A and lectin – were discovered; over 2000 such proteins are now known.
Circular permutation can occur as the result of evolutionary events, posttranslational modifications, or artificially engineered mutations. The two main models proposed to explain the evolution of circularly permuted proteins are permutation by duplication and fission and fusion. Permutation by duplication occurs when a gene undergoes duplication to form a tandem repeat, before redundant sections of the protein are removed; this relationship is found between saposin and swaposin. Fission and fusion occurs when partial proteins fuse to form a single polypeptide, such as in nicotinamide nucleotide transhydrogenases.
Circular permutations are routinely engineered in the laboratory to improve their catalytic activity or thermostability, or to investigate properties of the original protein.
In 1979, Bruce Cunningham and his colleagues discovered the first instance of a circularly permuted protein in nature.[1] After determining the peptide sequence of the lectin protein favin, they noticed its similarity to a known protein – concanavalin A – except that the ends were circularly permuted. Later work confirmed the circular permutation between the pair[2] and showed that concanavalin A is permuted post-translationally[3] through cleavage and an unusual protein ligation.[4]
After the discovery of a natural circularly permuted protein, researchers looked for a way to emulate this process. In 1983, David Goldenberg and Thomas Creighton were able to create a circularly permuted version of a protein by chemically ligating the termini to create a cyclic protein, then introducing new termini elsewhere using trypsin.[5] In 1989, Karolin Luger and her colleagues introduced a genetic method for making circular permutations by carefully fragmenting and ligating DNA.[6] This method allowed for permutations to be introduced at arbitrary sites.[6]
Despite the early discovery of post-translational circular permutations and the suggestion of a possible genetic mechanism for evolving circular permutants, it was not until 1995 that the first circularly permuted pair of genes were discovered. Saposins are a class of proteins involved in sphingolipid catabolism and antigen presentation of lipids in humans. Chris Ponting and Robert Russell identified a circularly permuted version of a saposin inserted into plant aspartic proteinase, which they nicknamed swaposin.[7] Saposin and swaposin were the first known case of two natural genes related by a circular permutation.[7]
Hundreds of examples of protein pairs related by a circular permutation were subsequently discovered in nature or produced in the laboratory. As of February 2012, the Circular Permutation Database[8] contains 2,238 circularly permuted protein pairs with known structures, and many more are known without structures.[9] The CyBase database collects proteins that are cyclic, some of which are permuted variants of cyclic wild-type proteins.[10] SISYPHUS is a database that contains a collection of hand-curated manual alignments of proteins with non-trivial relationships, several of which have circular permutations.[11]
Evolution
There are two main models that are currently being used to explain the evolution of circularly permuted proteins: permutation by duplication and fission and fusion. The two models have compelling examples supporting them, but the relative contribution of each model in evolution is still under debate.[12] Other, less common, mechanisms have been proposed, such as "cut and paste"[13] or "exon shuffling".[14]
Permutation by duplication
The earliest model proposed for the evolution of circular permutations is the permutation by duplication mechanism.[1] In this model, a precursor gene first undergoes a duplication and fusion to form a large tandem repeat. Next, start and stop codons are introduced at corresponding locations in the duplicated gene, removing redundant sections of the protein.
One surprising prediction of the permutation by duplication mechanism is that intermediate permutations can occur. For instance, the duplicated version of the protein should still be functional, since otherwise evolution would quickly select against such proteins. Likewise, partially duplicated intermediates where only one terminus was truncated should be functional. Such intermediates have been extensively documented in protein families such as DNA methyltransferases.[15]
Saposin and swaposin
An example for permutation by duplication is the relationship between saposin and swaposin. Saposins are highly conserved glycoproteins, approximately 80 amino acid residues long and forming a four alpha helical structure. They have a nearly identical placement of cysteine residues and glycosylation sites. The cDNA sequence that codes for saposin is called prosaposin. It is a precursor for four cleavage products, the saposins A, B, C, and D. The four saposin domains most likely arose from two tandem duplications of an ancestral gene.[16] This repeat suggests a mechanism for the evolution of the relationship with the plant-specific insert (PSI). The PSI is a domain exclusively found in plants, consisting of approximately 100 residues and found in plant aspartic proteases.[17] It belongs to the saposin-like protein family (SAPLIP) and has the N- and C- termini "swapped", such that the order of helices is 3-4-1-2 compared with saposin, thus leading to the name "swaposin".[7][18]
Fission and fusion
Another model for the evolution of circular permutations is the fission and fusion model. The process starts with two partial proteins. These may represent two independent polypeptides (such as two parts of a heterodimer), or may have originally been halves of a single protein that underwent a fission event to become two polypeptides.
The two proteins can later fuse together to form a single polypeptide. Regardless of which protein comes first, this fusion protein may show similar function. Thus, if a fusion between two proteins occurs twice in evolution (either between paralogues within the same species or between orthologues in different species) but in a different order, the resulting fusion proteins will be related by a circular permutation.
Evidence for a particular protein having evolved by a fission and fusion mechanism can be provided by observing the halves of the permutation as independent polypeptides in related species, or by demonstrating experimentally that the two halves can function as separate polypeptides.[19]
Other processes that can lead to circular permutations
Post-translational modification
The two evolutionary models mentioned above describe ways in which genes may be circularly permuted, resulting in a circularly permuted mRNA after transcription. Proteins can also be circularly permuted via post-translational modification, without permuting the underlying gene. Circular permutations can happen spontaneously through autocatalysis, as in the case of concanavalin A.[4] Alternately, permutation may require restriction enzymes and ligases.[5]
Role in protein engineering
Many proteins have their termini located close together in 3D space.[21][22] Because of this, it is often possible to design circular permutations of proteins. Today, circular permutations are generated routinely in the lab using standard genetics techniques.[6] Although some permutation sites prevent the protein from folding correctly, many permutants have been created with nearly identical structure and function to the original protein.
The motivation for creating a circular permutant of a protein can vary. Scientists may want to improve some property of the protein, such as:
Reduce proteolytic susceptibility. The rate at which proteins are broken down can have a large impact on their activity in cells. Since termini are often accessible to proteases, designing a circularly permuted protein with less-accessible termini can increase the lifespan of that protein in the cell.[23]
Improve catalytic activity. Circularly permuting a protein can sometimes increase the rate at which it catalyzes a chemical reaction, leading to more efficient proteins.[24]
Alter substrate or ligand binding. Circularly permuting a protein can result in the loss of substrate binding, but can occasionally lead to novel ligand binding activity or altered substrate specificity.[25]
Improve thermostability. Making proteins active over a wider range of temperatures and conditions can improve their utility.[26]
Alternately, scientists may be interested in properties of the original protein, such as:
Fold order. Determining the order in which different parts of a protein fold is challenging due to the extremely fast time scales involved. Circularly permuted versions of proteins will often fold in a different order, providing information about the folding of the original protein.[27][28][29]
Essential structural elements. Artificial circularly permuted proteins can allow parts of a protein to be selectively deleted. This gives insight into which structural elements are essential or not.[30]
Modify quaternary structure. Circularly permuted proteins have been shown to take on different quaternary structure than wild-type proteins.[31]
Find insertion sites for other proteins. Inserting one protein as a domain into another protein can be useful. For instance, inserting calmodulin into green fluorescent protein (GFP) allowed researchers to measure the activity of calmodulin via the fluorescence of the split-GFP.[32] Regions of GFP that tolerate the introduction of circular permutation are more likely to accept the addition of another protein while retaining the function of both proteins.
Design of novel biocatalysts and biosensors. Introducing circular permutations can be used to design proteins to catalyze specific chemical reactions,[24][33] or to detect the presence of certain molecules using proteins. For instance, the GFP-calmodulin fusion described above can be used to detect the level of calcium ions in a sample.[32]
Algorithmic detection
Many sequence alignment and protein structure alignment algorithms have been developed assuming linear data representations and as such are not able to detect circular permutations between proteins.[34] Two examples of frequently used methods that have problems correctly aligning proteins related by circular permutation are dynamic programming and many hidden Markov models.[34] As an alternative to these, a number of algorithms are built on top of non-linear approaches and are able to detect topology-independent similarities, or employ modifications allowing them to circumvent the limitations of dynamic programming.[34][35] The table below is a collection of such methods.
The algorithms are classified according to the type of input they require. Sequence-based algorithms require only the sequence of two proteins in order to create an alignment.[36] Sequence methods are generally fast and suitable for searching whole genomes for circularly permuted pairs of proteins.[36]Structure-based methods require 3D structures of both proteins being considered.[37] They are often slower than sequence-based methods, but are able to detect circular permutations between distantly related proteins with low sequence similarity.[37] Some structural methods are topology independent, meaning that they are also able to detect more complex rearrangements than circular permutation.[38]
Describes protein structures as one-dimensional text strings by using a Ramachandran sequential transformation (RST) algorithm. Detects circular permutations through a duplication of the sequence representation and "double filter-and-refine" strategy.
Works in two stages: Stage one identifies coarse alignments based on secondary structure elements. Stage two refines the alignment on residue level and extends into loop regions.
^ abBowles DJ, Pappin DJ (February 1988). "Traffic and assembly of concanavalin A". Trends in Biochemical Sciences. 13 (2): 60–4. doi:10.1016/0968-0004(88)90030-8. PMID3070848.
^ abGoldenberg DP, Creighton TE (April 1983). "Circular and circularly permuted forms of bovine pancreatic trypsin inhibitor". Journal of Molecular Biology. 165 (2): 407–13. doi:10.1016/S0022-2836(83)80265-4. PMID6188846.
^ abcLuger K, Hommel U, Herold M, Hofsteenge J, Kirschner K (January 1989). "Correct folding of circularly permuted variants of a beta alpha barrel enzyme in vivo". Science. 243 (4888): 206–10. Bibcode:1989Sci...243..206L. doi:10.1126/science.2643160. PMID2643160.
^ abcdPonting CP, Russell RB (May 1995). "Swaposins: circular permutations within genes encoding saposin homologues". Trends in Biochemical Sciences. 20 (5): 179–80. doi:10.1016/S0968-0004(00)89003-9. PMID7610480.
^Hazkani-Covo E, Altman N, Horowitz M, Graur D (January 2002). "The evolutionary history of prosaposin: two successive tandem-duplication events gave rise to the four saposin domains in vertebrates". Journal of Molecular Evolution. 54 (1): 30–4. Bibcode:2002JMolE..54...30H. doi:10.1007/s00239-001-0014-0. PMID11734895. S2CID7402721.
^Thornton JM, Sibanda BL (June 1983). "Amino and carboxy-terminal regions in globular proteins". Journal of Molecular Biology. 167 (2): 443–60. doi:10.1016/S0022-2836(83)80344-1. PMID6864804.
^Yu Y, Lutz S (January 2011). "Circular permutation: a different way to engineer enzyme structure and function". Trends in Biotechnology. 29 (1): 18–25. doi:10.1016/j.tibtech.2010.10.004. PMID21087800.
^Qian Z, Lutz S (October 2005). "Improving the catalytic activity of Candida antarctica lipase B by circular permutation". Journal of the American Chemical Society. 127 (39): 13466–7. doi:10.1021/ja053932h. PMID16190688. (primary source)
^Viguera AR, Serrano L, Wilmanns M (October 1996). "Different folding transition states may result in the same native structure". Nature Structural Biology. 3 (10): 874–80. doi:10.1038/nsb1096-874. PMID8836105. S2CID11542397. (primary source)
^Turner NJ (August 2009). "Directed evolution drives the next generation of biocatalysts". Nature Chemical Biology. 5 (8): 567–73. doi:10.1038/nchembio.203. PMID19620998.
^ abBachar O, Fischer D, Nussinov R, Wolfson H (April 1993). "A computer vision based technique for 3-D sequence-independent structural comparison of proteins". Protein Engineering. 6 (3): 279–88. doi:10.1093/protein/6.3.279. PMID8506262.
^ abShatsky M, Nussinov R, Wolfson HJ (July 2004). "A method for simultaneous alignment of multiple protein structures". Proteins. 56 (1): 143–56. doi:10.1002/prot.10628. PMID15162494. S2CID14665486.
^Zuker M (September 1991). "Suboptimal sequence alignment in molecular biology. Alignment with error analysis". Journal of Molecular Biology. 221 (2): 403–20. doi:10.1016/0022-2836(91)80062-Y. PMID1920426.
^Schmidt-Goenner T, Guerler A, Kolbeck B, Knapp EW (May 2010). "Circular permuted proteins in the universe of protein folds". Proteins. 78 (7): 1618–30. doi:10.1002/prot.22678. PMID20112421. S2CID20673981.
^Wang L, Wu LY, Wang Y, Zhang XS, Chen L (July 2010). "SANA: an algorithm for sequential and non-sequential protein structure alignment". Amino Acids. 39 (2): 417–25. doi:10.1007/s00726-009-0457-y. PMID20127263. S2CID2292831.