AI

Biological Language Models: AI That Reads & Writes the Genome

TL;DR The central dogma of molecular biology: DNA → RNA → Protein. The exciting frontier is that AI models can now operate at each layer and across layers. Models like ESM-2 and ESMFold operate at the protein level. Nucleotide Transformer and DNABERT operate at the DNA level. EVO (Arambam, Hie et al., 2024) operates across the full DNA-to-protein span

Read this article as text (accessible version)
Biological AI · Synthetic Biology · 2026 · 4,000 words · 4 interactive labs · April 2026

Biological Language
Models: AI That
Reads & Writes
the Genome

For a century, drug discovery meant waiting for nature to invent something we could steal — screening molecules found in soil bacteria, rainforest plants, and deep-sea organisms. For the last decade, it meant using AI to predict the shapes of proteins that already existed. Now, something fundamentally different is happening. AI is generating novel biological sequences that have never existed on Earth — and some of them work.

Read the Deep Dive ↓ Open Genomics Lab 🧬 ATGC GATC AGCT TAGC ... 3.2 billion base pairs Table of Contents
  1. What Are Biological Language Models?
  2. Genome Tokenization & Training
  3. EVO & the Genomic Foundation Models
  4. AI-Designed CRISPR & Anti-CRISPR Tools
  5. Synthetic Bacteriophages Against Superbugs
  6. Biology as a Software Design Space

01What Are Biological Language Models — and Why They're Different

In 2017, a paper called "Attention Is All You Need" introduced the Transformer architecture that would go on to power GPT-4, Claude, Gemini, and every large language model you've heard of. Those models learned language by predicting the next word in a sequence — trained on billions of sentences until they developed a rich internal model of how human language works. The insight driving biological language models is deceptively simple: DNA, RNA, and proteins are also sequences. They have their own alphabet (A, T, G, C for DNA; twenty amino acids for proteins), their own grammar (codon reading frames, regulatory motifs, structural constraints), and their own semantics (sequences that fold into functional proteins vs. ones that don't, sequences that bind specific targets vs. ones that don't).

If you train a Transformer-class model on enough biological sequences — millions of genomes from bacteria, archaea, viruses, plants, animals — it learns the statistical regularities of biological "language" the same way GPT models learn human language. It learns that certain nucleotide patterns reliably precede protein-coding regions. It learns which amino acid sequences fold into stable structures and which don't. It learns the "vocabulary" of regulatory signals that switch genes on and off across diverse organisms. This learned knowledge isn't explicitly programmed — it emerges from the training data, just like how a language model learns grammar without ever being taught grammatical rules explicitly.

The critical difference from earlier AI in biology: AlphaFold predicts the structure of proteins that already exist in nature, given their amino acid sequence. That's an extraordinary achievement — solving a 50-year-old problem. But it's still fundamentally a lookup: given a known input, find the corresponding output. Biological language models operate in generative mode. Given a partial sequence, a function description, or a structural constraint, they can generate novel sequences that have never existed in any organism — and that exhibit the desired biological properties. This is the difference between using a dictionary to look up words and being able to compose poetry in a language you've mastered.

The counterintuitive insight about why this works at all: evolution has already explored an enormous fraction of functional biological space over 3.8 billion years. The sequences in the genomes of millions of sequenced organisms represent a massive implicit dataset of "what works biologically." A model trained on this dataset learns the deep structural principles that make sequences functional — and can then extrapolate to generate sequences that follow those principles even if they've never appeared in any sequenced organism.

💡 Biological Models Work Across the Central Dogma

The central dogma of molecular biology: DNA → RNA → Protein. The exciting frontier is that AI models can now operate at each layer and across layers. Models like ESM-2 and ESMFold operate at the protein level. Nucleotide Transformer and DNABERT operate at the DNA level. EVO (Arambam, Hie et al., 2024) operates across the full DNA-to-protein span — it can model the relationship between a gene sequence and its encoded protein without being explicitly told which is which. This cross-layer capability means EVO can generate a DNA sequence that encodes a protein with specified structural properties — working backward from desired protein function to the DNA sequence that produces it.

biological_llm_analogy.py — DNA as language
# DNA ↔ Language analogy
# Nucleotides (A, T, G, C) → Characters (a-z)
# Codons (ATG, TAA, ...) → Words (3-char "words" encoding amino acids)
# Genes → Sentences (sequences encoding one protein)
# Genomes → Books (complete instruction set for one organism)

# The codon table: DNA "words" → amino acid "meaning"
codon_table = {
 "ATG": "Met (Start)", # Every protein starts here
 "TGG": "Trp", # Only codon for tryptophan
 "TAA": "Stop", # End of protein
 "GCT": "Ala",
 "CGT": "Arg",
 # ... 64 total codons encode 20 amino acids + 3 stop signals
}

# Biological LLMs learn these mappings implicitly from sequence data
# and discover higher-order structure: motifs, regulatory signals,
# protein fold patterns — without being explicitly programmed

# Using HuggingFace for a pre-trained biological model (ESM-2)
from transformers import EsmModel, EsmTokenizer

tokenizer = EsmTokenizer.from_pretrained("facebook/esm2_t6_8M_UR50D")
model = EsmModel.from_pretrained("facebook/esm2_t6_8M_UR50D")

# Tokenize a protein sequence (single letter amino acid codes)
protein = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEKAVQVKVKALPDAQFEVVHSLAKWKRQTLGQHDFSAGEGLYTHMKALRPDEDRLSPLHSVYVDQWDWERVMGDGERQFSTLKSTVEAIWAGIKATEAAVSEEFGLAPFLPDQIHFVHSQELLSRYPDLDAKGRERAIAKDLGAVFLVGIGGKLSDGHRHDVRAPDYDDWSTPSELGHAGLNGDILVWNPVLEDAFELSSMGIRVDADTLKHQLALTGDEDRLELEWHQALLRGEMPQTIGGGIGQSRLTMLLLQLPHIGQVQAGVWPAAVRESVPSLL"

inputs = tokenizer(protein, return_tensors="pt")
outputs = model(**inputs)
# outputs.last_hidden_state: shape [1, seq_len, 480]
# Each position = 480-dimensional embedding encoding local + global context
print(ff"Sequence embedding shape: {outputs.last_hidden_state.shape}")

02Genome Tokenization: Teaching the Model to Read DNA

Before a biological language model can learn anything, you have to solve the tokenization problem: how do you convert a DNA sequence into discrete tokens that a Transformer can process? This turns out to be more nuanced than simply treating each nucleotide (A, T, G, C) as a character token. Character-level tokenization works but creates extremely long sequences — the human genome is 3.2 billion characters, making context window management a fundamental challenge. The alternative approaches each embody different assumptions about biological structure.

The three main approaches: k-mer tokenization (treating overlapping subsequences of length k as tokens — 6-mers give 4⁶ = 4,096 possible tokens, roughly matching the vocabulary size of protein models) is the most common and was used in models like DNABERT. It's interpretable (each token represents a specific sequence motif) but creates a vocabulary explosion for larger k. Byte-pair encoding (BPE) — the same algorithm used in GPT models — learns frequently occurring subsequences from the training data and compresses them into tokens. This naturally captures biologically meaningful motifs: common regulatory sequences, codon patterns, and repetitive elements emerge as high-frequency tokens. Single-token models (treating each nucleotide as one token) maximize resolution at the cost of extremely long sequence lengths that challenge even the longest context window architectures.

EVO, the genomic foundation model from Arc Institute and Stanford (2024), trained on 2.7 million prokaryotic and viral genomes using a single-nucleotide tokenization with a context window of 131,072 tokens — enough to model entire bacterial genomes and many viral genomes in a single pass. This long-context capability is crucial: biological functions often depend on long-range interactions between sequence elements separated by thousands of base pairs. A model that can only see local context misses these relationships. EVO's 7-billion parameter scale and 300 billion training tokens make it orders of magnitude larger than previous biological sequence models.

⚡ The Conservation Signal in Biological Training Data

Here's the thing most biology-AI explainers miss: the most informative training signal in biological sequence models isn't prediction accuracy on single sequences — it's the pattern of conservation across species. When the same sequence motif appears in bacteria, archaea, and eukaryotes separated by billions of years of evolution, the model learns that this motif is functionally important — strong selection pressure has kept it intact. When sequences vary between species but maintain similar predicted function, the model learns which positions are functionally constrained (can't change without losing function) vs. which are neutral (can change freely). This implicit evolutionary information, encoded in the patterns of the training data, is what gives these models genuine biological insight rather than superficial pattern matching.


03EVO and the Emergence of Genomic Foundation Models

In early 2024, researchers at Arc Institute, Stanford, and collaborating institutions published a paper describing EVO — a 7-billion-parameter language model trained on the nucleotide sequences of 2.7 million prokaryotic and phage genomes. The results were striking not just for their scale, but for what the model learned to do without being explicitly taught. EVO could generate novel CRISPR-Cas protein sequences with plausible function. It could predict the essentiality of genes in unseen organisms. It could design novel transposable elements. And in zero-shot settings (no task-specific fine-tuning), it performed competitively with models explicitly trained for these specialized tasks.

Imagine you're reading a dense technical manual in a language you've never formally studied, but you've read millions of pages of text in that language. You've absorbed its grammar, vocabulary, and even some idioms through sheer exposure — without anyone explaining the rules to you. This is essentially what EVO did with genomic sequences. The model internalized the "grammar" of DNA — what sequences tend to precede protein-coding regions, what patterns characterize functional RNA structures, what local sequence contexts are associated with protein-binding sites — all from raw sequence data, without any explicit biological annotations.

The most provocative demonstration in the EVO paper: the model could generate novel CRISPR systems — complete with Cas proteins and guide RNA components — that were functional when tested experimentally. These weren't sequences found in any known organism; they were genuinely novel inventions that EVO generated by applying its learned model of "what biological sequences look like and how they work." The success rate wasn't 100% — biology is complex and even the best generative models produce many non-functional sequences — but the fraction of viable novel sequences was far higher than random chance, demonstrating that EVO had truly learned meaningful biological structure.

⚠️ Foundation Models ≠ Perfect Biological Designers

Biological language models generate sequences that are plausible by learned statistical criteria, but biology has many failure modes that don't show up in sequence data alone. A generated sequence might look statistically reasonable but fail because: the protein misfolds under cellular conditions (not just in silico), the sequence triggers immune responses that weren't selected against in training data, expression levels are inadequate in the target organism, or the functional activity works in isolation but not in the complex cellular environment. Always treat computationally generated sequences as hypotheses requiring wet-lab validation, not finished products. The current generation of models dramatically accelerates hypothesis generation — it doesn't eliminate experimental verification.

LLM next token prediction probability distribution diagram showing vocabulary tokens with probability bars and sampling mechanism

04AI-Designed CRISPR Tools: Writing New Gene Editors from Scratch

CRISPR-Cas9 became the dominant gene editing tool after 2012 because it's both precise and programmable: the Cas9 protein cuts DNA at a specific location guided by a short RNA sequence that you design. But the CRISPR systems found in nature have limitations — the Cas9 protein is large (making it hard to deliver into cells), it only works at specific DNA sequences adjacent to a PAM motif, and it has off-target activity that's acceptable in research but concerning for therapeutic use. The field has spent years searching for better CRISPR variants in the genomes of obscure bacteria — smaller Cas proteins, different PAM requirements, novel mechanisms.

Biological language models reframe this search problem. Instead of waiting for nature to have already evolved a better CRISPR system and waiting for biologists to find the organism that contains it, you can ask: given EVO's learned model of what functional CRISPR-Cas systems look like, can it generate novel ones that have never existed? The 2024 EVO paper demonstrated exactly this: the model generated novel CRISPR-Cas protein sequences with predicted structural features consistent with DNA-editing function. Experimental validation of a subset showed that some of these generated proteins bound DNA and exhibited editing activity. The critical insight: EVO wasn't copying or combining known CRISPR sequences — it was generating genuinely novel sequences that followed the learned grammar of CRISPR biology.

Anti-CRISPR proteins represent the other side of the equation. Just as bacteria have evolved CRISPR systems to defend against viruses, viruses have evolved anti-CRISPR proteins (Acrs) that block CRISPR activity — nature's arms race at the molecular level. Anti-CRISPR proteins are valuable therapeutic tools: they can turn off CRISPR editing in cells where you want to stop editing activity, providing an "off switch" for gene therapies. Finding natural anti-CRISPR proteins is laborious; generating designed ones with specific inhibitory properties is a frontier that biological language models are beginning to enable. Models can generate candidate Acr sequences and computational docking simulations can filter for those likely to interact with specific Cas protein targets.

💡 The PAM Problem and Why AI-Designed Cas Variants Matter

One of CRISPR-Cas9's limitations is its requirement for a specific PAM sequence (NGG for SpCas9) adjacent to the target site. This constrains which genomic locations can be edited — about one-third of possible positions are inaccessible. Smaller Cas variants discovered in nature (like CasΦ from huge phages) have different or looser PAM requirements, expanding targetable positions. AI-designed Cas variants could theoretically be optimized for any PAM requirement, any size constraint (for cell delivery), or any specificity requirement — effectively removing the natural constraint that only sequences evolution happened to create are available as tools. This is what "turning biology into a software design space" concretely means for gene editing.


05Synthetic Bacteriophages: Programming Viruses to Kill Superbugs

Antibiotic-resistant bacteria kill approximately 700,000 people per year globally — a number projected to reach 10 million by 2050 if trends continue. The pipeline of new antibiotics has essentially run dry: developing a new antibiotic class takes 10-15 years and costs billions, and the economics are terrible (patients take antibiotics for 7-14 days, making recoupment of development costs difficult). Bacteriophages — viruses that specifically infect and kill bacteria — have been used therapeutically in Eastern Europe and the former Soviet Union for decades, but their adoption in Western medicine has been hampered by specificity challenges (each phage typically kills only a narrow range of bacterial strains) and regulatory complexity.

Biological language models are beginning to address the phage specificity problem. A phage infects a specific bacterium by recognizing specific proteins on the bacterial surface using its tail fiber proteins. The sequence of these tail fiber proteins determines host range. EVO and related models, trained on the genomes of thousands of phages, learn the relationship between tail fiber sequences and host specificity — and can generate novel tail fiber sequences predicted to target specific bacterial strains, including antibiotic-resistant pathogens that natural phage collections don't cover. This is genuinely exciting: computationally designing a phage to kill a specific superbug in a patient resistant to all available antibiotics.

The practical pipeline is still experimental but the proof of concept is there. You would: sequence the pathogenic bacterial strain to identify its surface proteins, use a biological language model to generate candidate tail fiber sequences predicted to bind those surface proteins, synthesize these sequences and insert them into a phage chassis (a phage backbone whose tail fiber gene has been removed), test activity against the target bacterium, and iterate. This cycle — currently taking months per iteration — is what AI-accelerated biology is beginning to compress. The target is weeks or even days for emergency compassionate use in antibiotic-resistant infections with no other treatment options.

🔬 The Regulatory Challenge Is as Hard as the Science

Generating a synthetic phage computationally is increasingly tractable. Getting it approved for use in patients is not. Each synthetic phage is essentially a novel biologic product requiring safety evaluation that current regulatory frameworks weren't designed for. The most promising near-term path is compassionate use authorization for patients with life-threatening antibiotic-resistant infections and no alternatives — a pathway that has been used successfully for natural phage therapy in individual cases. Regulatory frameworks for AI-designed biological therapies are actively being developed, but the mismatch between the pace of the technology and the pace of regulatory adaptation is a genuine bottleneck. The technology is advancing faster than the institutional structures to deploy it safely.


06Biology as Software: What the Design Space Actually Looks Like

The framing of "biology as software" is both useful and potentially misleading, and it's worth unpacking both dimensions. The useful part: biological sequences really do function like code. They can be read, edited, compiled (expressed into proteins), and tested. DNA sequence changes (mutations) can be designed with the predictability of code changes. Biological systems have modular components (regulatory elements, protein domains) that can be mixed and matched. And increasingly, you can describe desired biological functionality in high-level terms and use AI to generate the low-level sequence implementation — the biological analogue of high-level programming languages compiling to machine code.

The misleading part: biological systems are dramatically more context-dependent than software. A gene that functions perfectly in one organism may be toxic in another. A protein sequence that folds correctly in vitro may misfold in the specific cellular environment of a target tissue. Biological systems have evolved in wet, noisy, variable environments — and their robustness to noise means they also have complex error-correction mechanisms that make them harder to engineer predictably than clean computational systems. The "biology as software" analogy is a useful design framework but shouldn't be taken to mean that biological engineering is as predictable as software engineering. Not yet.

Picture this: a future design workflow for a new biologic therapy. A researcher specifies: "Design a protein that binds to this cancer-cell surface marker with high affinity, can be expressed in CHO cells at therapeutic levels, has a half-life of approximately 48 hours in serum, and minimal immunogenicity in human patients." A biological AI design system — combining a generative model like EVO, structural prediction like ESMFold, immunogenicity prediction, and expression optimization models — generates 10,000 candidate sequences ranked by predicted satisfaction of all criteria. The top 100 are synthesized and tested. The top 10 from testing enter a second design-synthesize-test iteration with feedback from experimental results incorporated. This closed-loop AI-accelerated design cycle is what "biology as software" means in practice — not that biology works like software, but that you can apply software-like design iteration processes to biological engineering problems.

⚡ The Wet Lab Is Not Going Away — It Gets Smarter

A common misconception in reporting on biological AI: that AI will replace wet-lab biology. The more accurate picture is that AI dramatically changes the ratio of in silico to in vitro work — and changes which experiments are worth doing. Instead of screening thousands of random variants hoping to find improved properties, you computationally generate hundreds of high-probability candidates and experimentally validate dozens. Instead of iterating blindly, you use experimental results to update the model and generate better candidates for the next round. The experimental work is still essential — models make predictions, experiments validate reality. But the ratio of useful experiments to total experiments improves dramatically.


synthesisThe Convergence: How It All Connects

Biological language models represent a convergence of three previously separate threads: the maturation of large-scale genomic sequencing (providing the training data), the development of Transformer architectures that can model long-range dependencies in sequences (providing the model architecture), and the accumulation of wet-lab validation infrastructure that can test computationally generated sequences at sufficient throughput to close the design loop. None of these three threads alone would have been sufficient — the combination is what's enabling the current moment.

The near-term trajectory: increasingly capable models trained on larger, more diverse biological datasets (eukaryotic genomes, epigenomic data, proteomes, metabolomes), with increasingly sophisticated multi-modal architectures that jointly model DNA, RNA, protein, and cellular context. The design tools become increasingly accessible — not just to well-resourced research institutions but potentially to specialized teams at biotechnology companies who can fine-tune foundation models for specific application domains. The long-term trajectory is less certain but genuinely transformative: biology transitioning from a fundamentally discovery-driven discipline (finding what nature has already made) to an engineering-driven discipline (designing what nature hasn't — but what works).


FAQFrequently Asked Questions

What is a biological language model and how does it differ from AlphaFold? + AlphaFold is a discriminative model: you give it a known protein sequence (amino acids), and it predicts the 3D structure that sequence folds into. It's an extraordinary prediction tool but doesn't generate novel sequences. Biological language models like EVO, ESM-2, and Nucleotide Transformer are generative models: trained on millions of biological sequences (DNA, RNA, proteins), they learn the statistical patterns — the "grammar" — of biological sequences. Given a partial sequence or a functional description, they can generate novel, complete biological sequences that follow these learned patterns. Think of AlphaFold as a very sophisticated function evaluator: "given this sequence, what shape is it?" Biological language models are more like sequence composers: "given these constraints, generate a novel sequence that should have this property." AlphaFold and biological LLMs complement each other — you can use a biological LLM to generate candidate sequences and AlphaFold/ESMFold to predict their structures. What is the EVO model and what makes it significant? + EVO is a 7-billion-parameter biological language model developed by Arc Institute and Stanford, published in 2024. It was trained on 2.7 million prokaryotic and phage genomes using a single-nucleotide tokenization scheme with a 131,072-token context window — large enough to model complete bacterial genomes. What makes EVO significant: (1) Scale — it's substantially larger than previous biological sequence models. (2) Cross-layer modeling — EVO models the relationship between DNA and protein sequences, capturing the full central dogma relationship. (3) Demonstrated generative capability — EVO generated novel CRISPR systems and transposable elements that showed experimental activity. (4) Zero-shot competitiveness — without task-specific fine-tuning, EVO performed competitively with specialized models on tasks like gene essentiality prediction, demonstrating genuine biological understanding rather than task-specific memorization. How can AI be used to design new CRISPR tools? + CRISPR-Cas systems consist of a Cas protein (the molecular scissors) and a guide RNA (the GPS that directs cutting to a specific DNA location). Natural CRISPR systems have limitations: Cas9 is large (hard to deliver into cells), requires specific PAM sequences adjacent to target sites, and has off-target activity. AI approaches to CRISPR design: (1) Generative models like EVO, trained on thousands of natural CRISPR systems, can generate novel Cas protein sequences with predicted structural features consistent with DNA-editing function. (2) Directed evolution with AI guidance — generate many variants of existing Cas proteins, predict activity computationally, test the best candidates experimentally, iterate. (3) Guide RNA optimization — use sequence models to design guide RNAs with optimal target specificity and minimal off-target activity. The EVO paper showed that computationally generated CRISPR-Cas proteins can exhibit editing activity when tested experimentally — the field is moving from discovery (finding CRISPR systems in nature) to design (generating custom CRISPR systems for specific applications). What are synthetic bacteriophages and why are they important for antibiotic resistance? + Bacteriophages (phages) are viruses that specifically infect and kill bacteria. They're highly specific — a phage that kills E. coli typically doesn't kill Staphylococcus. This specificity is both an advantage (no disruption of beneficial bacteria) and a limitation (you need a phage matched to the specific bacterial strain causing infection). Natural phage therapy has been used in Eastern Europe for decades but hasn't been widely adopted in Western medicine due to specificity challenges and regulatory hurdles. Synthetic bacteriophages address the specificity problem: by understanding (and now AI-modeling) the relationship between phage tail fiber proteins and their bacterial host receptors, researchers can potentially design phages to target specific antibiotic-resistant bacterial strains. The goal is a computer-aided design process: sequence the pathogen → use AI to design tail fibers targeting that pathogen's surface proteins → synthesize and test → deploy. This would be particularly valuable for extensively drug-resistant (XDR) bacterial infections with no effective antibiotic treatment options. What does "biology as a software design space" actually mean in practice? + The phrase refers to the transition from biology as a discovery science (finding things in nature) to biology as an engineering science (designing things with specific properties from scratch). In practice: instead of screening natural compound libraries hoping to find a drug with desired properties, you computationally generate candidate molecules and filter computationally before any wet-lab work. Instead of searching diverse environments for bacteria with useful enzymes, you generate novel enzyme sequences with the desired catalytic properties. The "software" analogy captures the ability to specify desired functionality in high-level terms, generate implementation (sequences) computationally, test (experimentally), and iterate. The closed loop of AI design → synthesis → experimental validation → model update → improved design is the core of what makes this feel software-like. The important caveat: biology is noisier and more context-dependent than software, so the iteration cycles are longer and more uncertain. But the trajectory is toward increasingly predictable biological engineering. How is genome tokenization different from text tokenization in LLMs? + Text LLMs (like GPT) typically use byte-pair encoding (BPE): frequently occurring character sequences are compressed into single tokens, creating a vocabulary of ~50,000-100,000 tokens. For DNA, the "alphabet" is just 4 characters (A, T, G, C), but sequence lengths are enormous (the human genome is 3.2 billion characters). The main tokenization approaches for DNA: (1) Single-nucleotide tokenization — each A, T, G, C is one token. Most precise but creates very long sequences. (2) k-mer tokenization — overlapping subsequences of length k are tokens. 6-mers give 4,096 possible tokens. Biologically interpretable (each token is a specific sequence motif). (3) BPE applied to DNA — learns frequently occurring biological sequences as tokens. Naturally captures biologically meaningful patterns like common regulatory elements. EVO uses single-nucleotide tokenization with a 131,072-token context window, trading computational cost for maximum sequence resolution. Protein models like ESM use single-amino-acid tokenization over an alphabet of 20 amino acids plus special tokens. What are the biosecurity and ethical concerns around biological language models? + Biological language models raise genuine biosecurity concerns that the research community is actively grappling with. The core tension: the same capabilities that enable designing novel CRISPR tools and bacteriophages to fight antibiotic resistance could, in theory, enable designing pathogens with enhanced properties. This "dual-use" problem is not new in biology — the knowledge underlying vaccine development is the same knowledge that could be misused — but AI capabilities could lower the barrier to misuse. Current responses: (1) Screening — many biological AI labs (including Evo's developers) screen generated sequences against databases of known dangerous pathogens and refuse to generate certain categories of sequences. (2) Access controls — restricting access to the most capable biological AI tools to vetted researchers and institutions. (3) Governance — the biosecurity community is actively developing frameworks for oversight of biological AI development and deployment. (4) Model-level safety training — similar to RLHF for language models, training biological AI to refuse generation of harmful sequences. The scientific community broadly agrees that these tools require active governance rather than unrestricted access.

🧬 Genomics Lab

Four interactive experiments: DNA tokenizer, genome analysis, sequence fitness predictor, and codon optimizer.

DNA Sequence Tokenizer Enter DNA sequence (A, T, G, C) Tokenization method 3-mer 0 Sequence length 0 Token count 4 Vocab size — Compression ratio Token Output Amino acid translation (3-mer mode)

Nucleotide frequency & sequence complexity across the sequence window

Sequence Complexity Analyzer Sequence (auto-generated or paste your own) — GC content — Shannon entropy — Repeat fraction — Complexity

Predicted sequence fitness landscape — mutations explored computationally

Sequence Fitness Predictor

Simulates how biological AI models score generated sequences — higher fitness = more likely to be functional.

Seed sequence Mutations to explore 50 — Best fitness — Avg fitness — High fitness (>0.7) — vs. seed Codon Usage Optimizer

Different organisms prefer different codons for the same amino acid. Optimizing codon usage increases protein expression.

Protein sequence (single-letter AA codes) Target organism E. coli Optimized DNA Sequence — CAI score (0-1) — GC content — DNA length (bp) — Predicted expression
Tags
biological-AIEVO-modelCRISPR-AIsynthetic-biologybacteriophage-designgenome-foundation-modelantibiotic-resistanceDNA-language-modelprotein-design-AIsynthetic-biology-AI
Share this article