Gene Expression: From DNA to Protein
How a gene is transcribed into RNA and translated into protein, and how cells switch genes on and off to become different cell types.
🎯 By the end of this lesson
- State the central dogma and identify where each step occurs in a eukaryotic cell.
- Compare DNA and RNA in structure and function.
- Transcribe a DNA template strand into mRNA.
- Use codons to determine a sequence of amino acids, including start and stop codons.
- Describe the roles of RNA polymerase, ribosomes and tRNA.
- Explain why cells with identical DNA can have different structures and functions.
- Describe how transcription factors and epigenetic marks regulate gene expression.
- Define differentiation and stem cell.
1Overview
Every cell in a person’s body, from a neuron to a toenail cell, carries essentially the same DNA. Yet neurons conduct signals and toenail cells do not. The difference lies in gene expression: which genes are used, when, and how much. This lesson follows a gene from DNA sequence to working protein, and then explains how cells control that process.
2From gene to protein: the central dogma
Most genes contain the instructions for building a protein. Proteins do much of the work of the cell: they speed up chemical reactions (enzymes), build structures, carry signals and transport molecules. Because DNA stays in the nucleus of a eukaryotic cell while protein factories (ribosomes) sit in the cytoplasm, a messenger is needed. The messenger is RNA.
The flow of information is summarized in the central dogma: DNA is transcribed into messenger RNA (mRNA), and the mRNA is translated by a ribosome into a chain of amino acids that makes up a protein. The protein then folds into a three-dimensional shape and carries out its job, which contributes to the organism’s traits.
A helpful analogy is a cookbook in a library. The master copy of every recipe (DNA) never leaves the reference room. Instead, a photocopy of one recipe (mRNA) is carried to the kitchen (ribosome), where the dish (protein) is prepared. The master copy stays protected, and many photocopies can be made from the same recipe.
| Feature | DNA | RNA (mRNA) |
|---|---|---|
| Strands | Two, forming a double helix | One |
| Sugar | Deoxyribose | Ribose |
| Bases | A, T, G, C | A, U, G, C (uracil replaces thymine) |
| Main job | Long-term storage of genetic information | Short-lived working copy of one gene |
3Step 1: transcription
Transcription is the copying of a gene from DNA into RNA. It is carried out by an enzyme called RNA polymerase. The steps are:
- Initiation. RNA polymerase attaches to a region at the start of the gene, and the DNA helix opens locally.
- Elongation. The enzyme moves along one DNA strand, the template strand, and adds RNA nucleotides that pair with the template bases (A pairs with U, T pairs with A, G with C, C with G). The helix closes up again behind it.
- Termination. At the end of the gene the enzyme releases the new RNA strand.
A template strand reads 3′-TACGGATTC-5′. The mRNA is built with matching bases: A for T, U for A, G for C, C for G, C for G, U for A, A for T, A for T, G for C. The mRNA reads 5′-AUGCCUAAG-3′. Note that uracil (U) appears wherever the DNA template has adenine.
Processing in eukaryotes
In eukaryotes the first RNA copy (pre-mRNA) is processed in the nucleus before it is exported. Genes contain coding regions called exons that are interrupted by non-coding introns. Introns are cut out and the exons are joined together in a step called splicing. A protective cap is added at one end and a tail of adenine nucleotides at the other. These changes protect the mRNA from breakdown and help the ribosome recognize it. As a result, eukaryotic mRNAs last for hours, while a typical bacterial mRNA lasts only seconds.
4Step 2: the genetic code
RNA has four letters, but proteins are built from 20 kinds of amino acids. The bridge between them is the genetic code. Reading mRNA three bases at a time, each triplet is a codon. There are 64 possible codons (4 × 4 × 4). Of these, 61 specify amino acids and 3 are stop codons (UAA, UAG and UGA) that end translation. The codon AUG is the start codon, and it also codes for the amino acid methionine.
The code is redundant (also called degenerate): several different codons can specify the same amino acid. This redundancy can soften the effect of some mutations, which is explored in the mutation lesson. The code is also nearly universal. Almost all organisms use the same codon dictionary, which is strong evidence that all life shares a common origin, and which makes it possible for a human gene to function inside a bacterium in genetic engineering.
The order of bases in a gene determines the order of codons in mRNA. The order of codons determines the order of amino acids. The order of amino acids determines how the protein folds and what it can do. A change in the DNA sequence can therefore change the protein and, through it, the trait.
5Step 3: translation
Translation takes place at a ribosome, a structure made of ribosomal RNA and proteins with a small subunit and a large subunit. The small subunit binds the mRNA, and the large subunit holds the transfer RNAs and joins amino acids together.
Transfer RNA (tRNA) molecules are the adaptors. Each tRNA carries one specific amino acid at one end and has a three-base anticodon at the other end. The anticodon pairs with a matching codon on the mRNA. For example, a tRNA with the anticodon GAU pairs with the codon CUA, which codes for the amino acid leucine.
- Initiation. The ribosome assembles on the mRNA at the start codon AUG, and the first tRNA (carrying methionine) pairs with it.
- Elongation. The ribosome moves along the mRNA one codon at a time, reading it from the 5′ end to the 3′ end. At each codon the matching tRNA arrives, and a peptide bond joins its amino acid to the end of the growing chain. The ribosome advances three bases per cycle.
- Termination. When a stop codon reaches the ribosome, no tRNA matches it. Release factors enter instead, the finished polypeptide is released, and the ribosome subunits separate.
Using the mRNA 5′-AUGCCUAAG-3′ from the transcription example: the codons are AUG, CCU and AAG. AUG codes for methionine (start). CCU codes for proline and AAG for lysine, so the beginning of the chain is methionine, proline, lysine. Each amino acid is carried by a tRNA whose anticodon pairs with the matching codon. A mutation that changed CCU to a stop codon would end the chain early.
Translation does not read DNA directly, and ribosomes do not make RNA. Transcription (DNA to RNA) and translation (RNA to protein) are two different processes with different machinery in different places: transcription happens in the nucleus of a eukaryotic cell and translation happens at ribosomes in the cytoplasm. In bacteria, which have no nucleus, the two processes occur almost at the same time.
| Feature | Transcription | Translation |
|---|---|---|
| Template | DNA (one strand) | mRNA |
| Product | RNA | Polypeptide (protein) |
| Key molecules | RNA polymerase | Ribosome, tRNA |
| Location in eukaryotes | Nucleus | Cytoplasm |
| Unit read | One base at a time | One codon (three bases) at a time |
6Controlling gene expression
If every gene were active all the time, cells would waste huge amounts of energy and space making proteins they do not need. Regulation “conserves energy and space,” and when it malfunctions it can contribute to disease, including cancer. Each cell therefore expresses only a subset of its genes.
Bacteria have no nucleus, so transcription and translation happen nearly simultaneously and control is mostly at the level of transcription: a gene is switched on or off. Eukaryotes can regulate at many stages: how tightly the DNA is packed, whether transcription starts, how the RNA is processed and exported, whether translation occurs, and how the finished protein is modified.
A central role is played by transcription factors, proteins that bind to DNA near a gene and help or block RNA polymerase. A gene can only be transcribed if its DNA is accessible. When DNA is wound tightly around histones, transcription factors cannot reach it.
Epigenetic marks
Epigenetics is the study of changes in gene activity that do not involve a change in the DNA sequence. Chemical tags can be added to the DNA itself (a process called DNA methylation, which usually silences the gene) or to histone proteins (which can loosen or tighten the packing). These tags are not permanent: they can be added or removed, and they often persist through cell division so that a skin cell divides into more skin cells. A well-known example in females is X inactivation, in which one of the two X chromosomes in each cell is silenced during embryonic development.
7Differentiation and stem cells
The development of an embryo is the best demonstration of gene regulation. A single fertilized egg divides into trillions of cells that become muscle, nerve, skin and blood. The process by which cells become specialized is called differentiation, and it happens because different cells switch on different sets of genes.
A stem cell is a cell with the potential to form many of the body’s cell types. When a stem cell divides, it can produce more stem cells or specialized cells. Embryonic stem cells, taken from very early embryos, have the greatest developmental potential. Adult stem cells, found in tissues such as bone marrow, can form only certain cell types: bone marrow stem cells, for example, give rise to blood cells. Stem cells are discussed again in the applied genetics lessons.
Lactose digestion shows regulation at work. The enzyme that digests milk sugar is made in the gut of infants, and in many people the gene is switched down after childhood. Whether a gene is on or off, as well as which allele is present, affects a trait. Medicines and diets that change which genes are active are an active area of research.
8When expression goes wrong
Because proteins carry out the functions of cells, changes in genes or in their control can have large effects. A change to one gene can alter a protein enough to cause a disease: variants in the HBB gene, for example, produce an abnormal form of the hemoglobin protein (hemoglobin S) that can distort red blood cells into a sickle shape. Similarly, variants in the PAH gene reduce the activity of an enzyme that breaks down the amino acid phenylalanine. In both cases a change in DNA changes a protein, and the protein change produces the trait or condition. Genes can also be switched on at the wrong time or in the wrong cell, and this plays a role in cancer.
The next lesson shows that the story does not end with the gene. The environment, from diet to sunlight, also shapes how genes are expressed and what traits result. More background on the molecule is in What is DNA?.
9Practice problems with solutions
Problem 1: transcription and translation together
A section of a template DNA strand reads 3′-TACAAAGCTATT-5′. Find the mRNA, the codons and the amino acids if AUG is methionine, UUU is phenylalanine, CGA is arginine and UAA is a stop codon.
Solution. Pairing each template base gives mRNA 5′-AUGUUUCGAUAA-3′. Grouped into codons this is AUG, UUU, CGA, UAA. The protein is methionine, phenylalanine, arginine, and then translation stops. Three amino acids are made from twelve bases, because each codon uses three.
Problem 2: why a protein can change without a gene changing
Two cells contain identical genes, yet one makes a large amount of a protein and the other makes none. Possible explanations include a transcription factor that is present in one cell and absent in the other, methyl tags that silence the gene in one cell, or tightly packed chromatin around the gene that blocks RNA polymerase. The DNA sequence does not need to differ.
Quick summary
- Transcription copies a gene into RNA, and translation reads the RNA in codons to build a protein.
- The genetic code is shared by nearly all organisms, which is evidence of common ancestry and the reason genes can be moved between species.
- Cells with the same DNA differ because they express different genes, controlled by transcription factors and epigenetic marks.
🔑Key terms
?Quick check
Try each question first, then reveal the answer.
1. State the three main steps of the central dogma in order.
DNA is transcribed into mRNA, the mRNA is translated at a ribosome into a polypeptide, and the protein carries out a function that contributes to a trait.
2. Write the mRNA made from the DNA template 3'-TACCGA-5'.
The mRNA is 5'-AUGGCU-3', because each template base is paired with A for T, U for A, G for C and C for G.
3. A section of mRNA reads AUG-CUA-UAA. How many amino acids are added, and why?
Two amino acids are added (methionine and leucine). UAA is a stop codon, so translation ends instead of adding an amino acid.
4. Why is the genetic code described as redundant, and why does it matter?
Several codons can specify the same amino acid, so some changes in DNA do not change the protein. Redundancy therefore reduces the effect of some mutations.
5. Describe the role of tRNA in translation.
A tRNA carries one specific amino acid and has an anticodon that pairs with a matching mRNA codon, so it places the correct amino acid at each position in the chain.
6. Why can a muscle cell and a nerve cell have the same DNA but different functions?
Each cell type expresses a different subset of its genes, producing different proteins. Regulation of gene expression, not a difference in the DNA itself, accounts for the specialization.
7. Explain how DNA methylation can change a trait without changing the DNA sequence.
Methyl tags added to DNA usually silence the gene, so the protein is not made. The sequence is unchanged, but the gene is switched off, and the tags can be added or removed.
8. Why do eukaryotic cells process pre-mRNA before translation, and what does splicing remove?
Processing prepares the RNA for export and protects it from breakdown. Splicing removes the non-coding introns and joins the coding exons so the ribosome reads a continuous coding sequence.
BC curriculum content covered in this lesson
- DNA structure and function: gene expression
References
- BC Ministry of Education and Child Care. Science 10 (curriculum, Content and Elaborations). Accessed October 7, 2026.
- OpenStax. Biology 2e, 15.1 The Genetic Code. Accessed October 7, 2026.
- OpenStax. Biology 2e, 15.4 RNA Processing in Eukaryotes. Accessed October 7, 2026.
- OpenStax. Biology 2e, 15.5 Ribosomes and Protein Synthesis. Accessed October 7, 2026.
- OpenStax. Biology 2e, 16.1 Regulation of Gene Expression. Accessed October 7, 2026.
- OpenStax. Biology 2e, 16.3 Eukaryotic Epigenetic Gene Regulation. Accessed October 7, 2026.
- NHGRI. Epigenetics (Talking Glossary). Accessed October 7, 2026.
- NHGRI. Stem Cell (Talking Glossary). Accessed October 7, 2026.
- MedlinePlus (NIH). Sickle cell disease. Accessed October 7, 2026.
- MedlinePlus (NIH). Phenylketonuria. Accessed October 7, 2026.
These lessons follow the content areas listed in the British Columbia curriculum. They are study material written for this site and are not an official document. The official curriculum is the authority on what each course requires. Lessons are general education, not medical advice.