Log in Sign up
Back to Discover
🧬

Whole genome sequencing

life science Maturity 9-11

All living things have a code.

Haemophilus influenzae 01.jpg
Haemophilus influenzae 01.jpg
This code is inside you. It tells your body how to grow. It is like a tiny book. We can read the whole book at once. This helps us stay healthy. Can you imagine reading your own code?

47 words

All living things have a code.

Haemophilus influenzae 01.jpg
Haemophilus influenzae 01.jpg
This code is inside your cells. It looks like a long, curly staircase.
Human karyotype with bands and sub-bands.png
Human karyotype with bands and sub-bands.png
This staircase has many tiny steps. Each step is a part of the code. Scientists can read every single step at once. This is called whole genome sequencing. It helps us learn about plants and animals. It can even help people stay healthy. We can find this code in hair or spit.
C. elegans.jpg
C. elegans.jpg
Reading the code is a big job!

89 words

All living things have a code called DNA.

Human karyotype with bands and sub-bands.png
Human karyotype with bands and sub-bands.png
This code is shaped like a spiral staircase. It is made of many tiny parts called bases. These bases form the steps of the staircase.
Haemophilus influenzae 01.jpg
Haemophilus influenzae 01.jpg
Whole genome sequencing is a way to read every single base in an organism. It looks at all the DNA in the cell. This includes DNA in the nucleus and the mitochondria.
C. elegans.jpg
C. elegans.jpg
Scientists use this tool to study life. They can find DNA in hair, spit, or even old bones. In 1995, scientists finished the first full genome of a bacterium. This bacterium was called Haemophilus influenzae. Later, they sequenced the genome of a worm named Caenorhabditis elegans. They also sequenced the genome of a fruit fly.
Drosophila melanogaster - front (aka).jpg
Drosophila melanogaster - front (aka).jpg
Some parts of the code are hard to read. These parts have many repeats. New ways of reading, called long-read sequencing, help fix this. This can create a complete map from end to end. This is called telomere-to-telomere sequencing. A full human version was published in 2022.

183 words

Every living thing has a special code called DNA.

Human karyotype with bands and sub-bands.png
Human karyotype with bands and sub-bands.png
This code is kept inside the cells of your body. Most of it sits in the nucleus, but some is in the mitochondria. DNA looks like a long, coiled spiral staircase. The steps of this staircase are made of four tiny molecules called bases. These bases always pair up in a specific way. Adenine always pairs with thymine, and guanine always pairs with cytosine. Whole genome sequencing is the way scientists read this entire code at once.
Haemophilus influenzae 01.jpg
Haemophilus influenzae 01.jpg
This process helps us understand how life works.

To read the code, scientists often use a method called shotgun sequencing. This starts by breaking the long DNA strands into many small fragments. Scientists then sequence these small pieces one by one. They use computer programs to find where the pieces overlap. These overlapping parts act like a puzzle. By matching the ends, the computer can build the original long sequence.

ABI PRISM 3100 Genetic Analyzer 3.jpg
ABI PRISM 3100 Genetic Analyzer 3.jpg
Some newer ways, like long-read sequencing, read much larger chunks of DNA. This helps scientists read through parts of the code that repeat many times. Reading these repeating parts is a hard job for older methods.

Scientists have been working on this for many years. In the 1970s and 1980s, the work was done by hand. They used manual methods like Sanger sequencing. By the 1990s, machines made the work much faster and automated.

Chromatogram.jpg
Chromatogram.jpg
In 1992, scientists fully sequenced a single chromosome from yeast. The first time a whole organism was sequenced was in 1995. That organism was a bacterium called Haemophilus influenzae. Since then, many different living things have been mapped by researchers.

There are many important milestones in this history. The first animal to have its whole genome sequenced was the worm Caenorhabditis elegans in 1998.

C. elegans.jpg
C. elegans.jpg
In the year 2000, scientists sequenced the fruit fly, Drosophila melanogaster. They also finished the first plant genome, Arabidopsis thaliana, in 2000.
Drosophila melanogaster - front (aka).jpg
Drosophila melanogaster - front (aka).jpg
The genome for the lab mouse, Mus musculus, was published in 2002.
54986main mouse med.jpg
54986main mouse med.jpg
Humans have about 3.2 billion nucleotide pairs in their cells. This is much larger than the 1,830,140 base pairs found in the first bacterium.

Today, this science is used in many ways. Scientists can find DNA in saliva, hair, or even old bones.

Arabidopsis thaliana inflorescencias.jpg
Arabidopsis thaliana inflorescencias.jpg
They use this data to study how life changes over time. In the future, it might help doctors choose the best medicine for you. This is called personalized medicine. We are even reaching a goal called telomere-to-telomere sequencing. This means reading the DNA from one end to the very other end. A complete human version of this was published in 2022.

461 words

Whole genome sequencing (WGS) is the process of determining the entire DNA sequence of an organism's genome at a single time. The genome is the complete set of genetic instructions for a living thing. This process includes sequencing all chromosomal DNA found in the cell nucleus. It also includes DNA located in the mitochondria. For plants, WGS involves sequencing the DNA found in chloroplasts as well.

Human karyotype with bands and sub-bands.png
Human karyotype with bands and sub-bands.png
Understanding the full sequence is vital for modern biology. It helps researchers study how life evolves and how different species are related.

To understand the mechanism, we must look at the structure of DNA. DNA, or deoxyribonucleic acid, is a long, coiled double helix that looks like a spiral staircase. The sides of this staircase are made of sugar and phosphate molecules. The steps are made of four chemical bases: adenine, thymine, guanine, and cytosine. These bases always pair together using hydrogen bonds. Adenine always pairs with thymine, while guanine pairs with cytosine.

Chromatogram.jpg
Chromatogram.jpg
A gene is a specific segment of this DNA that provides the code to build proteins or RNA molecules. The specific order of these bases acts as the code for life.

Scientists use several different methods to read these sequences. One common approach is shotgun sequencing. This method involves breaking the long DNA strands into many small fragments. Scientists sequence these fragments and then use computer programs to piece them back together. This is similar to solving a massive puzzle where the overlapping edges of the pieces show how they fit.

ABI PRISM 3100 Genetic Analyzer 3.jpg
ABI PRISM 3100 Genetic Analyzer 3.jpg
Another method involves sequencing larger DNA clones from libraries, such as bacterial artificial chromosomes (BACs) or yeast artificial chromosomes (YACs). Newer technologies, called high-throughput or next-generation sequencing, allow for much faster work. These include methods like Illumina dye sequencing, pyrosequencing, and SMRT sequencing.

History shows how much this technology has changed. In the 1970s and 1980s, sequencing was a manual process. Scientists used methods like Maxam–Gilbert and Sanger sequencing to read small parts of genomes. The shift to automated capillary sequencers in the 1990s allowed for much larger projects.

ABI PRISM 3100 Genetic Analyzer 3.jpg
ABI PRISM 3100 Genetic Analyzer 3.jpg
In 1976, the first virus genome, Bacteriophage MS2, was sequenced. By 1992, scientists fully sequenced the third chromosome of yeast. In 1995, the bacterium Haemophilus influenzae became the first entire organism to have its genome fully sequenced.
Haemophilus influenzae 01.jpg
Haemophilus influenzae 01.jpg

Many important organisms have been mapped since those early days. The nematode worm, Caenorhabditis elegans, was the first animal to be sequenced in 1998.

C. elegans.jpg
C. elegans.jpg
In 2000, the fruit fly Drosophila melanogaster and the plant Arabidopsis thaliana were both sequenced.
Drosophila melanogaster - front (aka).jpg
Drosophila melanogaster - front (aka).jpg
Arabidopsis thaliana inflorescencias.jpg
Arabidopsis thaliana inflorescencias.jpg
The laboratory mouse Mus musculus genome was published in 2002.
54986main mouse med.jpg
54986main mouse med.jpg
Genome sizes vary wildly between species. The H. influenzae genome has 1,830,140 base pairs. Humans have much more, with about 3.2 billion nucleotide pairs in each germ cell. Some organisms, like Amoeba dubia, have massive genomes containing 700 billion nucleotide pairs.

One major challenge in sequencing is dealing with repetitive regions. Traditional methods often produced short "reads" that were difficult to assemble in these areas. This left gaps in the genetic map, known as scaffolds.

Human karyotype with bands and sub-bands.png
Human karyotype with bands and sub-bands.png
To solve this, scientists developed long-read sequencing, such as nanopore technology. While nanopore sequencing can be less accurate, its reads are much longer. This allows researchers to span entire repetitive regions. When a genome is sequenced perfectly from end to end, it is called telomere-to-telomere (T2T). A human T2T genome was finally published in 2022.

Whole genome sequencing has massive significance for the future of medicine. It is different from DNA profiling, which only looks at the likelihood of an individual's identity. WGS can pinpoint functional variants that help predict disease susceptibility or how a person might respond to a drug. This is a key part of personalized medicine. Scientists can extract the necessary DNA from many sources, including saliva, hair follicles, bone marrow, or even ancient bones. This technology connects biology, medicine, and computer science to help us understand the very blueprint of life.

686 words
🖼️ Images & Media (10)
File:Chromatogram.jpg
Chromatogram.jpg
File:Human karyotype with bands and sub-bands.png
Human karyotype with bands and sub-bands.png
File:Haemophilus influenzae 01.jpg
Haemophilus influenzae 01.jpg
File:C. elegans.jpg
C. elegans.jpg
File:Drosophila melanogaster - front (aka).jpg
Drosophila melanogaster - front (aka).jpg
File:Arabidopsis thaliana inflorescencias.jpg
Arabidopsis thaliana inflorescencias.jpg
File:54986main mouse med.jpg
54986main mouse med.jpg
File:Elaeis guineensis MS 3467.jpg
Elaeis guineensis MS 3467.jpg
File:ABI PRISM 3100 Genetic Analyzer 3.jpg
ABI PRISM 3100 Genetic Analyzer 3.jpg
File:Historic cost of sequencing a human genome.svg
Historic cost of sequencing a human genome.svg
Up Next
🧬
DNA sequencing
Life Science
More to explore

🔬 Go deeper

More advanced topics to explore

🪜 Step back

Simpler topics to build understanding

What is Nepedia?

A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.