Base pairing of 8-letter genetic alphabet in E. coli RNAP transcription
The P:Z and B:S base pairs expand the canonical four-letter genetic alphabet by introducing alternative hydrogen-bonding patterns that remain compatible with Watson-Crick geometry in DNA duplex (Fig. 1a). Prior work has shown that B:S pair can serve as 5th and 6th bases to support efficient E. coli RNAP transcription9. Here, we sought to test whether we can expand to eight-letter genetic alphabet and determine how E. coli RNAP selects and incorporates nucleotides opposite the P:Z pair.
We performed single-nucleotide transcription assays using DNA/RNA hybrid scaffolds containing a site-specific unnatural base at the +1 position of the template strand (dP, dZ, dB, dS) (Fig. 1b). Eight nucleotide triphosphates—including all four natural NTPs and four unnatural triphosphates—were tested for incorporation opposite the templating base by E. coli RNAP. In all scaffolds tested, we found that addition of the cognate partner nucleotide effectively produced a clear band shift from the 9-mer RNA to the n + 1 product within 15 s incubation, demonstrating that both P:Z and S:B are efficiently and selectively incorporated as orthogonal unnatural base pairs during transcription (Fig. 1b; Supplementary Fig. 1). In addition, single-turnover kinetics measurements further showed that kobs values for dZ:PTP and dP:ZTP are only approximately two-fold lower than those of the natural dG:CTP pair in the same conditions, demonstrating that P:Z incorporation is comparable to natural nucleotide incorporation for natural base pairs (Supplementary Fig. 2). Thus, the combined use of P:Z and B:S supports the feasibility of an expanded eight-letter genetic alphabet in which all four unnatural nucleotides pair selectively and independently, preserving informational integrity in transcription processes.
To comprehensively evaluate substrate recognition selectivity, we quantified band intensities across all 32 template-NTP combinations (four scaffolds × eight NTPs (ATP, GTP, CTP, UTP, PTP, ZTP, BTP, and STP)) at the 15-s time point (Supplementary Fig. 1). For all the template-NTP combinations tested, the cognate unnatural substrate incorporations are among the most efficient ones as expected. Substrate discrimination between cognate and mismatch substrates is most pronounced at the early time point, indicating misincorporation is very slow in comparison with cognate substrate incorporation for most cases.
dB:UTP and dZ:GTP misincorporation stand out as two major misincorporations (Fig. 1b, Supplementary Fig. 1). We have previously kinetically characterized dB:UTP misincorporation and revealed the structural basis of dB:UTP transcription recognition. These studies showed that the isoguanine heterocycle in the dB:UTP mismatch adopts rare tautomer to form Watson–Crick like base pair at the active site of E. coli RNAP9.
For the P:Z system, dZ:GTP misincorporation is substantially higher than that of any other non-cognate pair. Similar Z:G mismatch behaviors have previously been observed during DNA replication and PCR. Benner and coworkers reported that deprotonated Z (Z⁻) can pair with G through a complementary hydrogen-bonding pattern, forming a Watson–Crick-like geometry (Fig. 2a)7,11. Consistent with this, the fidelity of Z:P has been shown to be pH-dependent in PCR reactions with Taq polymerase, supporting the involvement of Z deprotonation rather than polymerase-specific interactions7.
Fig. 2: Z* suppresses Z:G mismatch incorporation by E. coli RNA polymerase.
a Chemical mechanism underlying Z deprotonation. The left panel illustrates how Z undergoes deprotonation to form Z-, which can mispair with guanine to generate a Z-:G base pair. The right panel shows the structure of 2′-F-α-carboxamide-Z (Z*), in which substitution of the nitro group with a carboxamide raises the pKa above 10, thereby preventing deprotonation and reducing G mispairing. Hydrogen-bond donors and acceptors are colored blue and red, respectively. Regions highlighted in pink denote functional groups that distinguish these unnatural nucleotides from natural bases. b Single-nucleotide incorporation assays for comparing scaffolds containing dZ or dZ* at the template i + 1 position. The nucleic acid scaffold used in b–d—comprising template-strand DNA (tsDNA, cyan), non-template-strand DNA (ntsDNA, green), and RNA (red)—is shown in the upper panel of (b). “X” marks the template i + 1 position. For the assay in (b), E. coli RNAP elongation complexes were assembled on each scaffold and treated with 1 μM of the indicated NTP substrate, including all four natural NTPs (ATP, GTP, CTP, and UTP) and all four unnatural NTPs (PTP, ZTP, BTP, and STP). Reactions were sampled at 10 s, resolved by 12% denaturing urea-PAGE, and visualized by autoradiography of p32-labeled RNA. These assays were performed in three independent replicates. c, d Single-nucleotide incorporation assays comparing the mismatch tendencies of Z and Z* opposite guanine. The left c shows PTP and GTP incorporation opposite dZ or dZ*. The right d shows ZTP and Z*TP incorporation opposite dP or dG. Each reaction contained 20 μM of the indicated nucleotide triphosphate and was sampled at 0, 30 s, 1 min, 10 min, and 30 min. The “9-mer” band corresponds to the p32-labeled RNA primer, while “+1” through “+3” indicate extended transcription products (10-nt to 12-nt products). These assays were performed in two independent replicates.
By analogy, the Z:G mismatch observed in our transcription assay most likely arises from the intrinsic hydrogen-bonding properties of the Z:P system—specifically, the ability of Z to deprotonate and adopt a complementary pattern that pairs with G—rather than from any unusual or idiosyncratic interaction between P:Z base pair and E. coli RNAP. In addition, we also noted other minor mismatches, including dP:UTP, dP:CTP, dZ:ATP, dS:ZTP, and dS:ATP. However, cognate incorporation proceeds with substantially higher efficiency than misincorporation under the same conditions, enabling kinetic discrimination. Note that as an intrinsic limitation in these single-nucleotide transcription assays, we “force” the misincorporation of single substrate in the absence of competition by the cognate complementary nucleoside triphosphate. Therefore, the misincorporation frequencies observed here are likely overestimates of the actual error rates when the cognate substrate (and all of the other NTPs) are present and compete simultaneously. The correctly complementary cognate substrates would kinetically outcompete these rare and slow mismatches in such complete systems.
A higher selective Z base analog: Z* for transcription
To reduce the intrinsic mismatch propensity of Z, the central strategy is to suppress formation of its deprotonated form, which is Watson-Crick complementary to G. The nitro group of Z, positioned in the major groove, is a strong electron-withdrawing substituent that lowers the pKa of the nucleobase to approximately 7.8, thereby favoring deprotonation under physiological conditions. To mitigate this effect, we employed the recently developed 2′-F-α-carboxamide-Z (Z*)12, which replaced the nitro group with a carboxamide moiety (the complete chemical structures of P, Z, and Z* in their NTP substrate and template DNA forms are shown in Supplementary Fig. 3). The pKa of Z* exceeds 10 and thus greatly disfavors deprotonation at neutral pH (Fig. 2a).
To systematically assess how the Z* modification affects the full substrate selectivity profile of E. coli RNAP, we tested all eight Hachimoji NTPs against dZ and dZ* template scaffolds at 1 μM substrate concentration and a 10-s reaction time (Fig. 2b; Supplementary Fig. 4). Cognate incorporation of PTP was clearly detected opposite both dZ and dZ*, confirming that Z* retains functional P:Z pairing capacity. In contrast, misincorporation of GTP detectable with the dZ scaffold was substantially reduced with dZ*, indicating that the carboxamide substitution substantially improves substrate selectivity. No other major misincorporation were observed for the dZ* template. A more complete characterization using higher substrate concentrations (20 μM) and extended reaction times (up to 30 min), including band intensity quantification, is provided in Fig. 2c, further confirming the improved selectivity for the dZ* template over the dZ template.
We next assessed the selectivity of incoming ZTP and Z*TP opposite dP in the template (Fig. 2d). Both nucleotides supported efficient cognate incorporation opposite P (n + 1 band). However, we also observed higher bands indicating further extension up to the n + 3 position due to ZTP misincorporation against dA and dG. In contrast, we observed the n + 1 product only for Z*TP incorporation opposite dP template. Similarly, we also observed multiple bands of ZTP misincorporation for the dG template (n + 1 and n + 3 bands). In contrast, no misincorporation bands were observed with Z*TP, indicating substantial improvement of selectivity. Notably, we also observed a substantial reduction in Z*TP misincorporation on the dS scaffold compared with ZTP (Supplementary Fig. 4c).
To test whether E. coli RNAP can continue elongation beyond the PZ site, we performed NTP-chase assays (Supplementary Fig. 5). Among all of the tested P:Z and P:Z* configurations, we found that E. coli RNAP can extend from n + 1 position and generate long RNA transcript efficiently. We observed no pausing at n + 2 site, indicating that E. coli RNAP can efficiently translocate from UBP site and add next nucleotide.
Together, these results demonstrate that the carboxamide substituent (Z*) significantly enhances the transcriptional fidelity of the P:Z pair and reduces dZ:GTP and dS:ZTP misincorporation. Importantly, neither P:Z nor P:Z* incorporation impairs subsequent processive elongation by E. coli RNAP.
Cryo-EM structure of E. coli elongation complex with dZ:PTP
To elucidate how a multi-subunit RNA polymerase recognizes and incorporates the P:Z base pair, we determined the cryo-EM structure of an E. coli RNAP elongation complex (EC) containing a dZ template paired with the incoming PTP nucleotide at the i + 1 position (Fig. 3; Supplementary Fig. 6; Supplementary Table 1). To trap a pre-catalytic state, we reconstituted the EC with a 3′-dOH RNA, which allows substrate binding but prevents phosphodiester bond formation, thereby stabilizing pre-incorporation intermediates. The cryo-EM reconstruction reached 2.42 Å resolution, with all five RNAP subunits—αI, αII, β, β′, and ω—clearly resolved (Fig. 3A).
Fig. 3: Cryo-EM structure of the E. coli RNA polymerase elongation complex containing dZ:PTP.
a Overview of the cryo-EM structure of the E. coli RNAP elongation complex (EC) assembled with a dZ template and the incoming PTP substrate. The EC contains five RNAP subunits—αI (light blue), αII (teal), β (yellow), β′ (silver), and ω (salmon)—together with the nucleic-acid scaffold composed of tsDNA (cyan), ntsDNA (green), and RNA (hot pink). The trigger loop (TL) is highlighted in purple. b The incoming PTP (hot pink) forms a canonical Watson–Crick base pair with templating dZ (cyan). The cryo-EM density (contoured at 8-σ) is shown together with the atomic model. Hydrogen bonds are indicated by blue dashed lines. c Architecture of the active site in the dZ:PTP EC. Key structural elements include RNA, tsDNA, the P-loop (silver), TL (purple), two catalytic magnesium ions (green), and the bridge helix (BH; light green). Two water molecules coordinating the metals—W0 for metal A and W1 for metal B—are depicted as red spheres. Critical residues involved in substrate stabilization are shown as sticks: TL residues M932, F935, and H936, and BH residue T790. The hydrogen bond between H936 and the β-phosphate of PTP is shown as a blue dashed line. Hydrophobic interactions between TL/BH residues and the nucleobase moieties are depicted as black bold dashed lines. d Detailed view of the two-metal coordination. Both atomic model and cryo-EM density (8-σ) are shown. The two Mg²+ ions (green) coordinate the three catalytic aspartates of the P-loop (silver), the triphosphate of PTP (hot pink), and two ordered water molecules (red), forming a complete two-metal-ion geometry characteristic of a pre-catalytic RNAP active site.
The resulting structure captures a well-defined PTP substrate positioned in the active site together with a fully folded trigger loop (TL; residues 916–1146 of β′), a hallmark of the catalytically competent, pre-catalytic conformation (Supplementary Fig. 6). The substrate PTP forms a canonical Watson–Crick base pair with the templating dZ (Fig. 3b), and its triphosphate adopts a characteristic chair-like conformation precisely coordinated by two catalytic magnesium ions (Fig. 3c, d).
The folded trigger loop (TL) was captured to sense the substrate PTP loading. In specific, TL residue M932 and bridge helix (BH) residue T790 jointly stabilize the nucleobases of PTP and dZ through hydrophobic interactions, while TL residue H936 forms a hydrogen bond with the β-phosphate of PTP (Fig. 3c). Comparison with published structures of RNAP ECs containing natural base pairs—including E. coli RNAP (PDB: 6RH3)—reveals that TL conformation and substrate positioning in our dZ:PTP complex closely resemble those in the native systems. Key TL residues adopt similar orientations across all structures, with only minor rotamer differences (Supplementary Fig. 7). These observations demonstrate that the P:Z base pair behaves analogously to a natural Watson-Crick pair during substrate loading and TL closure.
The high-resolution map further reveals a fully formed two-metal-ion catalytic configuration. Metal B (Mg²+) exhibits clear octahedral geometry, coordinating the α-, β-, and γ-phosphate groups of PTP in a tridentate arrangement, with a water molecule (W1) completing the coordination sphere (Fig. 3d). Asp460 and Asp462 (β′) bridge metals A and B, while metal A is further coordinated by the substrate α-phosphate, Asp464 (β′), an ordered water molecule (W0), and presumably the RNA 3′-OH (Fig. 3d). In previously reported RNA polymerase structures, the geometry of the catalytic metal coordination is often less well defined, largely because of limited map quality, and in some cases the density for metal B is not observed. To better assess the accuracy of our assignment, we therefore compared our active-site configuration to high-resolution structures of native RNA13 and DNA polymerases14,15. The metal coordination geometry and the chair-like triphosphate conformation in our EC are nearly identical to those observed in native polymerases (Supplementary Fig. 8).
Together, these results demonstrate that the P:Z base pair is accommodated by E. coli RNAP in a manner highly similar to natural substrates, supporting accurate Watson–Crick pairing, proper alignment of catalytic metals, and a fully folded TL conformation characteristic of a bona fide pre-catalytic state.
Cryo-EM structure of E. coli elongation complex with dP:Z*TP
To further understand how the P:Z* base pair is recognized during transcription by a multi-subunit RNA polymerase, we applied the same strategy used for the dZ:PTP complex and reconstituted an E. coli RNAP EC containing a dP template paired with the incoming Z*TP nucleotide (Fig. 4; Supplementary Fig. 9; Supplementary Table 1). Because P is now on the template strand and Z* is the incoming nucleotide, this dataset provides a strand-switched counterpart to the dZ:PTP complex, enabling direct comparison of their substrate-recognition mechanisms.
Fig. 4: Cryo-EM structure of the E. coli RNA polymerase elongation complex containing dP:Z*TP.
a Overview of the cryo-EM structures of the E. coli RNAP EC assembled with a dP template and the incoming Z*TP substrate. Three conformational states were resolved from this dataset based on the position of the SI3 domain and the conformation of the TL: TL-open (a) (steel blue), TL-open (b) (light green), and TL-closed (orange). The upper panel shows a conformational spectrum illustrating SI3 movement from a distal position (steel blue) to a proximal position (orange) relative to the active site, with the three structures mapped onto this trajectory. The lower panel shows the overall structure of the dP:Z*TP EC in the TL-closed state. The five RNAP subunits are rendered as transparent surfaces, colored and labeled as indicated. The three SI3 conformations corresponding to the states above are superimposed for comparison. b Structural comparison of the TL and SI3 domain among the three states. The active site of the TL-closed state is shown as a cartoon and includes RNA (hot pink), tsDNA (cyan), ntsDNA (lime), the P-loop (silver), the catalytic magnesium ion (green), the BH (transparent light green), TL, and SI3. TL and SI3 from all three states are superimposed using the color scheme from panel a. The inset highlights the conformation of the incoming Z*TP. Black arrows indicate the repositioning of SI3, TL folding, and loading of the Z*TP triphosphate moiety. c Base-pairing geometry of dP:Z*TP across all three states. In each structure, the template dP (cyan) forms a canonical Watson–Crick base pair with Z*TP. Atomic models are shown together with cryo-EM densities contoured at 6-σ. The fluorine atom at the C2′ position of Z*TP is colored in lime green. d Architecture of the active site in the TL-closed state of the dP:Z*TP EC. Key structural elements include RNA, tsDNA, the P-loop (silver), TL (purple), the catalytic magnesium ion (metal A, green), and BH (light green). Critical residues involved in substrate stabilization are shown as sticks: TL residues M932, F935, and H936, and BH residue T790. The hydrogen bond between H936 and the β-phosphate of Z*TP is shown as a blue dashed line. Hydrophobic interactions between TL/BH residues and nucleobase moieties are depicted as black bold dashed lines.
In contrast to the relatively homogeneous TL conformation observed in dZ:PTP dataset, the dP:Z*TP dataset displayed substantial higher conformational heterogeneity in the trigger loop (TL) and SI3 domain (residues 943–1130 of β′), allowing us to resolve three distinct EC states (Fig. 4; Supplementary Fig. 9). Based on TL conformation, these structures were classified as TL-open (a) (2.60 Å), TL-open (b) (2.64 Å), and TL-closed (2.75 Å).
The three structures exhibit coordinated movements of the SI3 domain and the TL as substrate loading progresses (Fig. 4a, b). SI3 is a large insertion within the TL, and its repositioning reflects the early steps of TL engagement. SI3 resides in its most distal position relative to the active site in the TL-open (a) state. In the TL-open (b) state, SI3 shifts toward the active site, representing an intermediate position in the repositioning pathway while the TL remains fully open. SI3 then continues the repositioning event that occurs in concert with TL folding into helices. Across this transition—from the TL-open (a) state to the TL-closed state—SI3 rotates by ~25° and shifts by ~26 Å toward the active site. A similar SI3 repositioning pathway has been reported for RNAP ECs containing natural base pairs16, indicating that Z* does not perturb this native conformational mechanism.
Across all three states, Z*TP is clearly resolved in the active site and forms a Watson–Crick base pair with the templating dP (Fig. 4c). Remarkably, correct P:Z pairing is achieved even in the TL-open states, demonstrating that accurate base-pairing geometry forms early—prior to TL folding and independent of assistance from the closed TL conformation. As the TL transitions to the closed state, the triphosphate moiety of Z*TP undergoes accommodation and adopts the characteristic chair-like conformation also observed in the dZ:PTP structure (Fig. 4b, inset). Key TL residues (H936, F935, M932) adopt similar conformations as in natural systems and as in the dZ:PTP pre-catalytic complex, stabilizing the incoming Z*TP (Fig. 4d).
Together with the dZ:PTP dataset, these results demonstrate that both P:Z and P:Z* base pairs are recognized by E. coli RNAP in a manner highly similar to natural Watson–Crick pairs (Supplementary Fig. 7). Z*TP achieves correct pairing geometry even before TL folding, and the subsequent folding of the TL, repositioning of SI3, and accommodation of the triphosphate proceed following the same sequence of structural transitions described for native substrates. Thus, P:Z and P:Z* behave as native-like base pairs during the early steps of substrate loading, with the major distinction occurring later in the catalytic trajectory, as elaborated in the following section.
A nitro-dependent π-hole interaction stabilizes BH bending and shifts the TL-folding equilibrium
Comparison of the dZ:PTP and dP:Z*TP cryo-EM datasets revealed strikingly different particle distributions between TL-open and TL-closed states. In the dZ:PTP dataset, 76.8% of particles adopted the TL-closed conformation, whereas only 23.2% remained TL-open (Fig. 5a). In contrast, the dP:Z*TP dataset displayed the opposite trend: 80.6% of particles were TL-open and only 19.4% were TL-closed (Fig. 5a). These results demonstrate that the TL-open/closed equilibrium is markedly shifted between the two systems, with the dZ:PTP EC strongly favoring the TL-closed state. Single-turnover incorporation assays confirmed that this TL equilibrium shift has direct kinetic consequences: dZ:PTP supports substantially faster PTP incorporation than dP:Z*TP (Supplementary Fig. 2e, f), consistent with the interpretation that a higher proportion of TL-closed particles reflects more productive active-site closure. Furthermore, we found that the dZ scaffold supports measurably faster PTP incorporation than the dZ* scaffold under otherwise identical conditions (Supplementary Fig. 2c, d), indicating that the nitro group specifically contributes to the rate enhancement of PTP incorporation at the i + 1 position. This effect does not propagate to downstream elongation steps as we observed no measurable difference on the kinetics of subsequent nucleotide addition by chase assay (Supplementary Fig. 5b).
Fig. 5: A nitro-dependent π-hole interaction stabilizes BH bending and shifts the TL-folding equilibrium.
a Particle distributions of TL-open (steel blue) and TL-closed (orange red) states in the dZ:PTP and dP:Z*TP cryo-EM datasets. The y-axis indicates the fraction of particles in each TL conformation relative to the total particle set, and the x-axis corresponds to the two datasets. b Structural comparison of the bridge helix (BH) and trigger loop (TL) among three cryo-EM structures: the dZ:PTP EC (salmon), the TL-closed state of the dP:Z*TP EC (orange), and the post-translocated E. coli RNAP EC (gray; PDB: 8FVR17). In the absence of an incoming substrate in the post-translocated state, the BH adopts a straight conformation, whereas both TL-closed states exhibit a substantially bent BH. The maximal deviation of the bent BH—measured at residue A787—is ~2 Å relative to the straight BH (indicated). Surrounding active-site elements from the dZ:PTP EC are shown, including RNA (hot pink), tsDNA (cyan), the P-loop (silver), and two catalytic magnesium ions (green). The water molecule (chain T, residue 104 in the dZ:PTP EC structure determined in this study; PDB 9ZP4), which mediates the interaction between the dZ nitro group and BH, is shown as a red sphere. Hydrogen bonds between the water molecule and BH residues are shown as blue dashed lines. The π-hole interaction between the dZ nitro group and the water molecule is indicated by black dashed arrows, pointing from the water lone pair toward the electron-deficient nitro nitrogen. Black arrows depict the directions of TL folding and BH bending. c Detailed view of the water-mediated interaction between BH and the nitro group of dZ. The π-hole-interacting water molecule is highlighted as a red sphere. Hydrogen bonds and the π-hole interaction are shown as blue dashed lines and black dashed arrows, respectively. Distances for each interaction are labeled. The lateral displacement of the water molecule relative to the normal axis of the nitro plane is 0.64 Å. The color scheme for tsDNA and BH follows (b). Cryo-EM density (contoured at 6-σ) is shown and the density for the π-hole-interacting water molecule is highlighted in slate blue. d Chemical schematic illustrating the water-mediated network between the BH and the dZ nitro group, corresponding to the structural features shown in (c).
Furthermore, examination of the active sites across both TL-closed structures revealed that the dZ:PTP EC adopts a conformation much closer to the catalytic transition than any state captured in the dP:Z*TP dataset. The TL density in dP:Z*TP is noticeably weaker (Supplementary Figs. 6d, 9i), suggesting incomplete TL folding or increased conformational flexibility or heterogeneity. In addition, the catalytic metal coordination differs: in dP:Z*TP, only metal A is observed; the density corresponding to metal B is too weak to model reliably (Fig. 4c; Supplementary Figs. 6, 9). Moreover, the distance between metal A and the α-phosphate of Z*TP ( ~ 3.4 Å) is incompatible with coordination (Fig. 4d). These features indicate that, despite adopting a TL-closed conformation, the dP:Z*TP EC is not as catalytically advanced as the dZ:PTP EC. Together, these observations suggest that an additional stabilizing interaction is present in the dZ:PTP complex that promotes a more mature TL-closed configuration.
Closer inspection of the UBP environment at the i + 1 position revealed a previously undiscovered interaction present only in the dZ:PTP EC, mediated by a specifically positioned water molecule (chain T, residue 104 in the dZ:PTP EC structure determined in this study; PDB 9ZP4) (Fig. 5b, c). This water molecule bridges the bridge-helix (BH) residues A787 and A791 to the nitro group of dZ in the template strand, forming a three-way interaction network (Fig. 5b). This interaction helps stabilize the bent conformation of the BH, a conformation required to accommodate TL folding into the helical closed state.
Relative to the straight BH conformation seen in the post-translocated EC (PDB: 8FVR17), the BH in the dZ:PTP TL-closed structure is substantially bent, with the greatest deviation ( ~ 2 Å) occurring at A787 (Fig. 5b). Notably, the water molecule is positioned precisely at this hinge point, supporting the idea that it contributes to maintaining BH bending and thereby facilitates TL folding.
Detailed geometric analysis shows that this water molecule does not hydrogen-bond directly to the nitro group. Instead, it donates a lone pair toward the electron-deficient π-hole located above the nitro nitrogen (Fig. 5c, d). Nitro groups are known to generate a positive electrostatic region along the normal axis of the nitro plane, created by electron withdrawal through its two oxygen atoms18. In our structure, this π-hole-interacting water molecule lies 2.62 Å from the nitro nitrogen, with a lateral displacement of only 0.64 Å from the normal axis of the nitro plane, a geometry characteristic of π-hole interactions (Fig. 5c). In addition, the water molecule forms two hydrogen bonds with the backbone carbonyl of A787 and the backbone amide of A791, producing a highly stable three-directional network (Fig. 5c, d). In contrast to the dP:Z*TP EC, this interaction is entirely absent. At the position corresponding to the dZ nitro group, the C7 atom of dP lacks an electron-deficient substituent and does not engage in direct or water-mediated interactions.
It is important to note that the nitro-dependent π-hole interaction is unlikely to be the sole determinant of the TL-equilibrium shift. Other factors—such as differences in nucleobase size (e.g., the purine PTP in dZ:PTP potentially stacking more strongly with TL residue M932 than the pyrimidine Z*TP)—may also influence the equilibrium. However, the water-mediated π-hole interaction provides a compelling and structurally explicit mechanism by which template-base chemistry can stabilize BH bending and thereby promote TL folding.
These observations highlight that modifications at the C5 position of pyrimidines or the N7 position of purines on the template strand can modulate the TL-open/closed equilibrium of the RNAP EC. This provides a conceptual basis for designing nucleotide analogues to tune RNAP conformational dynamics and potentially regulate transcriptional outcomes.