The Replicase round 1 puzzle launched today. Players have lots of data from the pilot puzzle to analyze while creating new designs throughout August. Here is a spreadsheet of the score components to see which designs performed better on activity vs. copyability. Please join me in welcoming Deni Szokoli as the new post-doc in the Das Lab who will be leading the research.
I downloaded the data images to a google doc for easier viewing. The google drive achieve however have the advantage that it is searchable. (Eg. search by your name or design ID)
Summary images of individual designs
I have also started to update the Replicase spreadsheet from last round
The spots that was pointed out as problematic areas - like the cross-linking sites that @Denszok pointed me to, the G-4 area and some Joyce paper positions, also looks to be some of those that causes issues when viewing the copyability windows in the data. But they are also spots that are hard in the starter design already. Good thing is that we have already managed to make some of these windows better. So there is plenty of room for improvement.
I have been dreaming about having the option to patch some of those good windows together.
I downloaded the designs dataset that was uploaded in the Replicase Pilot conclusion and have started to add extra data. More is to come.
G-quads (G4) in the replicase
First thing I saw when I viewed our replicase starter in round 1 was that it had a potential G4 (Q-quadruplex). I was wondering if that was a coincidence so I looked up other replicases and some of them had a G4 at the same spot. I did a single experiment where I specifically tried to kill the G4. The replicase wasn’t active. But I had made other changes too. So I will do a new experiment this round.
G4’s hurt copyability
When I sort for best copyability, the designs that score best, tends to not have a G4. Many of these are also not active.
Otherwise they tended not to be predicted to form in Alphafold 3 with 2 K+
Alphafold 3 run on G4 sequences
I have run some of the G4 sequences through Alphafold 3 to get a prediction about them folding with 2 K+ ions. They typically folds when there is K+, where as Mg2+ makes them relax and unfold)
High ribozyme activity
However when I sort for ribozyme activity, it looks like G4’s are well wanted.
Plus not just any kind. Most of the G4’s among the designs with the more activity are from successful clusters. Or they have certain types of G4’s, namely those known already from the starting sequence or sibling that @DigitalEmbrace put up. (The yellow and light purple above)
G4’s seems fine to have under activity. I still suspect they may have a function. So if we are keeping them and want them to not hurt copyability, one could change ion concentration towards them not forming in that phase where they causes issues. This leads me to ask:
What characterises Denszok’s highly copyable designs?
Jill’s lab conclusion on the Replicase Pilot lab states:
“…denszok’s algorithmic DesiRNA series is consistently strong on copyability”
All @denszok’s designs removes the potential G4.
@Jieux also systematically breaks up the G4 in his Steps series. He does not have a lot of ribozyme activity, but lots of copybility.
However what characterizes the designs with high activity, is that most of them keep their G4. Plus when I run that G4 sequence through Alphafold 3 with 2 K+, that most of them looks like they reacts to K+ concentration changes.
When I sort for eterna score to see which did better than the starter with both copyability and activity, there are both designs with and without potential G4’s.
Questions for the experimentalists
Are you using same ion concentration under the two experiment phases? Copyability and activity?
Will it be allowed to not have the same ion concentration?
If we can’t change the concentrations, we can either delete the G4’s or find better ones. ![]()
Hey Eli, you say I “systematically breaks up the G4 in his Steps series. He does not a lot of ribozyme activity, but lots of copability.” — not sure what that means in terms of this lab or where I should focus. The lab description is not terribly accessable to my understanding. Open to suggestions… my instnict is to do a mutation generator on bases 10, 16, 84 & 90 (the nearest neighbors to the small PK on the 3’ end… but the mutation generator does not appear to be active in this lab… also they seem to want 500 designs but don’t have a player limit… so feels a bit “first come first served”
Hi @Jieux!
We don’t have to use all the slots. They are just there if we want them. We both know we want them ![]()
Ok, here is what we are trying to achieve.
-
We want an RNA that can copy itself or copy some other RNA. (Copyability)
-
We wants lots of RNA copies. (Activity)
How do we want it? We want both at the same time.
It doesn’t help if we have an RNA that can copy itself but takes days and only leave 1 copy.
It also doesn’t help if we can get lots of copies of 30 base snippets from the replicase, if we can’t get a full length copy.
A way you can use your nearest neighbors strategy is to focus on designs that score higher than the start sequence on both copyability_score (67.4351) and ribozyme_activity_score (54.0628)
I have isolated the designs for you:
For Jieux - copyability and activity over original
I just noticed your series as scoring okay on being able to being copied and having no G-quardruplexes. (Tiny knot like structures that have 4 strands hydrogen bond around some K ions.)
Using the lab data
I have color coded the 30 base frames directly in the data set in the same manners as the images we received. Because then I can easier compare more designs to each other.
I already mentioned that most of our designs and the starter sequences tends to get stuck at some of the same windows. Here are some ideas for what can be done about it.
Window 3 is a problematic spot in my pilot design Replica 38 #M13. My replicase would pretty much stop here if it were to copy itself full length.
More blue = Low copyability
More yellow = High copyability
Notice that the starter sequence isn’t doing fantastic either. It is an area that many of our designs struggle with. However not all of them. ![]()
So I start looking for another design that does well in window. I have the sheet sorted after Ribozyme activity score.
Window 3 covers the base range 11-40 in the design.
My design Weakening stems 3 #M12 scores 78 which looks promising. However when I check the design, it has no base changes in the 11-40 base areas compared to the design I want to improve.
So instead I take a look at Denszok’s DesiRNA9 that scores 100 for window 3.
I take a stab and guess that the region in my design that isn’t working well is the region 11-15. Because the windows before and after window 3, don’t look too bad.
Now I want a look at the summary images with window walks. I search the image archive with my design ID. The bad region in window 3 pretty much overlaps with region 11-15:
So I only adopt the changes from @Denszok’s design for that region and not for the full stretch of the region 11-40. Originally I didn’t have any changes in the 11-15 area before.
Finding more good copy windows
Another way to find the best scoring windows for a range is to sort by column arrow and use Sort sheet by Z-A.
Response from Rhiju:
We don’t have potassium ions in our reactions, which are typically required to stabilize quadruplexes. So my guess is that we don’t have to worry about G quads, though we are amenable to trying other solution conditions.
The experimental conditions are listed in one of the docs linked in the pilot lab conclusion.
I do not see any 4 g in a row.
@Astromon, when I was talking about G4’s, it is the new fancy name for a G-quadruplex. So not 4 G’s in a row. But rather 2 x 4 GG’s. Earlier they were only counted if they were 3 x 4 G’s. But the smaller ones can also form g-quardruplexes. The potential one present in our replicase even have a 5 wheel - an extra wheel, so if one of the GG islands were deleted at either end, it may still function.
@DigitalEmbrace, thanks to you and Rhiju for your feedback. It is good to know there are no potassium near our RNA’s.
I looked at the documentation and it mentions Tris - a stabilizing buffer for RNA. I haven’t been able to find anything about Tris in relation to G-quadruplexes. Although I may not have found it, because Tris has a lot of synonym names.
I will still be on the lookout for if G4 entirely disappear from our top scorers. As it looks now, Tris isn’t making G-quadruplexes form. Although I can’t know for sure.
I see the following options.
-
The G4 has a function for the replicase itself, Tris can stabilise it, if so we might need to preserve it
-
The G4 has no function, Tris can’t stabilize it, we don’t need to conserve it
-
The G4 would disrupt the replicase from working - in which case it should be deleted, in case Tris can stabilize it
In any case I can as well advice everyone but me to completely ignore the G4’s and eventually the data will spill on what is true. If we can make tons of winners without G4 that are good at copying and activity, then G4’s are irrelevant here. I’ll just keep watching for them for the fun of it.
Hi Eli: Thanks for all the posts. I thought I would “pick your brain”. I have noticed that many of your designs pair the locked bases putting them into stacks. I would “think” that this would led to a reduction of catalytic activity but your designs assume otherwise. Is this an experiment or have you run across evidence that states locked bases in stacks are as good as loose locked bases as far as catalytic activity is concerned?
Hi @JR1!
Thx for asking. I need to ask a bit more to make sure I understand.
When you mention locked bases do you mean the bases in the puzzles with the locks on them? Can you explain a bit more and perhaps show me an image with an example?
Do you mean that I tend to put GC base pairs around the locked bases, when there are pairs around them?
The lab discussion states that “some NTs are locked because they are needed for catalytic activity”
Of the 13 locked bases, 3 are paired, 10 are loose. I’ll go though the designs tomorrow to see if pairing more is a good thing or bad thing, or if it even matters. Thought you might have an opinion.
Ah, ok, now I understand. Thx for your explanation.
Ok, first the bases that appears single in game, aren’t necessarily so viewed as crystal. Our rounds new single base pair (32 - 146) that appears to be flying lonely in loop space.:
It isn’t nearly as lonely when viewed in the crystal:
I was actually quite surprised when I saw. As I told Jill:
that new 32-146 base pair. It isn’t just base pairing in space. It is under what almost looks like a GCA triplet, and at the other end it stacks with 4 bp (correction 6) and then a GA mismatch. So it is stuck into a stack of base pairs. So it basically locks some parts together.
So perhaps it is permissible using a weaker base pair there.
Base 32 in larger context:
I have added images to most of the weird places from the crystal in the Replicase sequences sheet.
Often when there are locked important places in an RNA, when viewed in crystal, there are structural enhancement of that area. Like triplet bases, purine stacking and such. Sometimes a seemingly lonely pair, is locked into another stack.
Eg, here is a nearby enhancement right next to the 146 base. It is 5 A tertiary stacking with each other. I bet they prefer to stay purine.
Purine stacking. A129, A135, A35, A63, A147. Actually also involved a C130 stacking too.
Also I think I can now answer what 1 nt good bulges can be good for in an RNA. What I see in this design is that several of them are used for locking different parts of the structure together. Here is an example:
There is one more 1 nt bulge case in this lab. 31 - 140
These two could just as well be UC, CC or UU. Or GA, AG, GG or AA. (If there is space) As long as they stay in the family of either purine or pyrimidine. Because they lock onto each other. The larger purines (A and G) are stronger for the task.
More copyable than the starter
When I look for potential exchanges for 31 under copyability, it looks like U is a good bet. That will make the two stacking bases into UU. A doesn’t look bad either. But I would expect it to have its partner 140 rather be A or G then.
When I look at 31 for ribozyme activity, U comes out best. However since we didn’t knew that these two were locked onto each other, I bet not many have tried the combo of the two. So I still plan trying putting in two A’s there instead.
Similarly I can check what the partner base 140 says. When I check for activity 140 would love to become A. It isn’t crazy about becoming C. But again I don’t think that combo has been tried a lot and I intend trying because maybe the replicase likes it. ![]()
So we can check bases across each other to see what combos are successful already across base pairs or as here stacks.
Advice on replicase design: residue conservation.
Thank you @DigitalEmbrace! I’m really excited to be here!
Hello players! I am Deni Szokoli, an incoming postdoc in the Das lab, and will be leading the charge on the replicase project. Some of you may already know me from my time as a player during the previous replicase round.
I wanted to provide you all with a list of residues I think are important to the 3D fold or catalytic function of the ribozyme, that I personally would not mutate. Most of these residues were not locked in the lab because for most of them we are not certain that there aren’t mutations that could be tolerated, or perhaps mutations that are only tolerated in the presence of some other mutation (epistatic effects). For that reason we wanted to collect data on all these residues as well, and if you are feeling bold you should feel free to modify all these residues, and explore — it is possible you discover some mutation or set of mutations that are activity enhancing, but we wouldn’t have dared to try them.
The list is in descending order of importance (top more critical residues, and bottom residues assumed to be less important). A lot of these residues are also a part of an interacting group of residues, so you might want to change the whole group all at once, if you want to tinker with them. For that reason I have also provided you with an excel sheet that has the ranked residues grouped and color-coded. Good luck, and have fun!
The list:
C37
C20
A19
A58
G18
A21
A22
G59
R132 (R = A or G)
A133
A134
G36
C104
A105
A33
U143
A35
A63
A147
A129
A135
C73
G96
C106
G107
U108
A109
C110
C111
C112
C30
G34
A145
G146
C32
A60
G61
C70
U71
EDIT2: I missed one mistake: It is supposed to be U143 not U144!
EDIT3: I also missed something at the NTP site. A133 should be conserved, and the R should be at position 132, not 133 as previously.
Any ideas on what is controlling copyability? The delta doesn’t seem to be it. The high copyability designs have very poor activities. Definitely looks like a tradeoff.
Some of your identities don’t match the default starting identities.
With your info on left, in the pentastack the differences are A130C, A136G, A148C. In the loop-loop there’s A146G and G147A. Do we know the cause or source of the differences?
Good eye! I was in a hurry and was afraid something would be off with the numbering. I will fix it immediately. Let me know if I missed something like this. Looks like I was looking at a version with maybe some deletion?
Hi @denszok, thanks a lot for your residue conservation list!
Can you explain what JL/A stands for?
Hi! Yeah, of course! JL/A is shorthand for “single-stranded Junction between Ligase and Accessory domain”. In other words, it is A105-G123.
For those of you who may have wondered about base stacking is. Or what pentastacking (Denszok’s fine phrase) is.
Base stacking can happen between aromatic rings. RNA and DNA bases are aromatic rings. Base stacking (Pi stacking or Pi-interaction) adds extra stability inside stems. But it can also happen in the case of tertiary interactions, where two bases from different places in the RNA sequence stacks onto each other. Like the stacking example with base 31 - 140 I mentioned in one of my posts above. Or the 20 - 37 stacking.
I have dug up a tiny video that explains base stacking. This is the best I could find so far.
What do you mean by delta? The delta G?
It doesn’t suprise me that some high copyability designs would be inactive. After all, a fully unfolded RNA would be very copyable, but wouldn’t be very active.























