An eight-letter genetic alphabet already works in DNA copying. Putting it through the transcription machinery that cellular life actually uses shows it passes the enzyme's checks, and reveals why the obvious fix for its main error backfires.

Every organism on Earth writes with four letters. Whether four is necessary or merely what happened is a question you can only answer by building alternatives, and chemists have built one: the hachimoji system, Japanese for eight letters, which adds two more base pairs to the usual set by rearranging where the hydrogen bonds sit while keeping the overall geometry the same.
The extra letters already work in DNA copying, and in transcription by a simple viral enzyme. The enzyme that matters for cellular life is a different proposition: a multi-subunit machine with elaborate internal checks on what it lets through. A collaboration between UC San Diego, the Foundation for Applied Molecular Evolution and SLAC has now put the eight-letter alphabet through the bacterial version and solved structures of it happening.
Why it matters: Whether a cell could ever run on an expanded alphabet depends on whether its own transcription machinery accepts the new letters. Working with a stripped-down viral enzyme does not answer that.
Given a template containing all eight letters, the bacterial polymerase read it and built the matching RNA. Four cryo-electron microscopy structures, at resolutions between 2.42 and 2.75 angstroms, show why it worked.
The unnatural pairs adopt the same geometry as natural ones in the active site, and, more tellingly, they trigger the same internal response. The enzyme has a mobile element called the trigger loop that folds into a catalytically active shape when a correct nucleotide arrives, which is part of how it checks its own work. The new letters induce that folding. They are not slipping past the checks; they are passing them.
Speed and continuity held up too. Incorporation rates were comparable to natural building blocks, and the enzyme resumed normal elongation afterwards rather than stalling at the modification.
One of the new letters, Z, has a persistent habit of pairing with G when it should not. The cause is chemical rather than enzymatic. Z carries a nitro group that pulls electrons away, which lowers the pKa of the ring to around 7.8. At the pH inside a cell that means a substantial fraction of Z molecules have lost a proton, and in that deprotonated form Z looks like cytosine, which pairs with G perfectly well.
This was known from DNA replication. What the structures establish is that the problem is not specific to any one enzyme. As Li and colleagues write in Nature Communications, the enzyme's stringent active site checkpoints are insufficient to overcome this fundamental chemical vulnerability. Both molecules being presented are, at that moment, genuinely the right shape. No amount of proofreading distinguishes them.
The fix therefore has to be chemical. Replacing the nitro group with a carboxamide raises the pKa above 10, so the deprotonated form barely exists at neutral pH, and misincorporation falls sharply.
The structures then reveal that the offending nitro group was doing something useful. A water molecule sits bridging two residues of the bridge helix and the nitro group, held by two hydrogen bonds and by a third contact to the electron-poor region above the nitro plane. That water sits at the hinge of the helix and helps hold the bend that encourages the trigger loop to fold.
Which means the group causing the errors was also assisting the catalysis, and the kinetics bear it out: the original version supports faster incorporation than the corrected one. Fixing the fidelity costs speed. The authors are direct that the improved letter is a first step and that further work on the scaffold is needed.
They also float a different route: rather than perfecting the chemistry, build the ambiguity into the code. The natural genetic code already spells several amino acids more than one way, and an expanded code could do the same, letting two versions of a codon mean the same thing so that a mispairing stops mattering. Those experiments are described as underway.
This is transcription in a tube, not a cell. The enzyme is purified, the templates are supplied, and the unnatural building blocks are provided at working concentrations. A living cell would have to import or manufacture them and keep them available, which is a separate and unsolved problem.
Nor does making RNA mean making protein. Getting an eight-letter message translated requires ribosomes and transfer RNAs that recognise the new codons, none of which is addressed here.
And the fidelity gain has been demonstrated as a reduction in one specific error, not as a solution. The paper is explicit that a fully optimised version of this pair does not yet exist, and the trade-off it uncovers suggests optimisation will not be a matter of simply removing the problematic group.
Why would anyone want more letters? More letters allow more information per unit length and, in principle, more amino acids than the twenty life uses. The immediate applications are in engineered systems rather than organisms.
What is pKa doing in this story? It measures how readily a molecule gives up a proton. Z gives one up easily enough that at cellular pH many copies are in a form that mimics a different letter, which is the whole source of the error.
What's the one-line takeaway? Bacterial RNA polymerase transcribes an eight-letter alphabet at natural speed and passes it through its own quality checks, but the chemical group responsible for the main error also helps the enzyme work, so improving accuracy currently costs speed.
Li et al. "Structural basis of transcription of the hachimoji eight-letter alphabet by E. coli RNA polymerase." Nature Communications, 2026;17(1). doi.org/10.1038/s41467-026-76668-0
PubMed PMID: 42686745.
Image: Bacterial RNA polymerase structure (PDB 1HQM), litvinanna, CC BY-SA 4.0, via Wikimedia Commons.
Weekly research updates, breakthrough summaries, and new articles — straight to your inbox. Free, always.
Comments