From genome to proteome
A genome is the complete set of genes in a cell. The proteome is the full range of proteins that a cell is able to produce.
In simpler organisms, information from a sequenced genome can be used to establish the amino acid sequences of proteins encoded by that genetic information and therefore investigate the organism's proteome.
Databases can link information about gene sequences with corresponding amino acid or protein sequences. Bioinformatics allows scientists to analyse these data after an organism's genome has been determined.
Genome information can be used to:
- locate likely genes in a DNA sequence;
- predict the amino acid sequences of proteins;
- compare genetic and protein information held in databases;
- investigate which proteins are expressed under different conditions.
Why simpler genomes are easier to interpret
Relating genome sequence to protein sequence is generally more straightforward in simpler organisms than in complex eukaryotes.
- The genomes are often smaller.
- Prokaryotic genes generally do not contain introns.
- The control of gene expression is less complex.
Bacteria and viruses are therefore useful examples when genome information is used to investigate protein sequences.
Exam Tip: Keep the definitions distinct: the genome is the complete set of genes in a cell, whereas the proteome is the full range of proteins that a cell is able to produce.