Showing posts with label genome. Show all posts
Showing posts with label genome. Show all posts

Saturday, April 5, 2008

Genome

In biology the genome of an organism is its whole hereditary information and is encoded in the DNA (or, for some viruses, RNA). This includes both the genes and the non-coding sequences of the DNA. The term was coined in 1920 by Hans Winkler, Professor of Botany at the University of Hamburg, Germany. The Oxford English Dictionary suggests the name to be a portmanteau of the words gene and chromosome, however many related -ome words already existed, such as biome and rhizome, forming a vocabulary into which genome fit all too well.[1]

More precisely, the genome of an organism is a complete genetic sequence on one set of chromosomes; for example, one of the two sets that a diploid individual carries in every somatic cell. The term genome can be applied specifically to mean that stored on a complete set of nuclear DNA (i.e., the "nuclear genome") but can also be applied to that stored within organelles that contain their own DNA, as with the mitochondrial genome or the chloroplast genome. When people say that the genome of a sexually reproducing species has been "sequenced," typically they are referring to a determination of the sequences of one set of autosomes and one of each type of sex chromosome, which together represent both of the possible sexes. Even in species that exist in only one sex, what is described as "a genome sequence" may be a composite read from the chromosomes of various individuals. In general use, the phrase "genetic makeup" is sometimes used conversationally to mean the genome of a particular individual or organism. The study of the global properties of genomes of related organisms is usually referred to as genomics, which distinguishes it from genetics which generally studies the properties of single genes or groups of genes.

Both the number of base pairs and the number of genes vary widely from one species to another, and there is little connection between the two. At present, the highest known number of genes is around 60,000, for the protozoan causing trichomoniasis (see List of sequenced eukaryotic genomes), almost three times as many as in the human genome.

An analogy to the human genome stored on DNA is that of instructions stored in a book:

  • The book over one billion words long.
  • The book is bound 5000 volumes, each 300 pages long.
  • The book fits into a cell nucleus the size of a pinpoint.
  • A copy of the book (all 5000 volumes) is contained in every cell (except red blood cells) as a strand of DNA over two metres in length.

Types

Most biological entities are more complex than a virus sometimes or always carry additional genetic material besides that which resides in their chromosomes. In some contexts, such as sequencing the genome of a pathogenic microbe, "genome" is meant to include information stored on this auxiliary material, which is carried in plasmids. In such circumstances then, "genome" describes all of the genes and information on non-coding DNA that have the potential to be present.

In eukaryotes such as plants, protozoa and animals, however, "genome" carries the typical connotation of only information on chromosomal DNA. So although these organisms contain mitochondria that have their own DNA, the genes in this mitochondrial DNA are not considered part of the genome. In fact, mitochondria are sometimes said to have their own genome, often referred to as the "mitochondrial genome".


Genomes and genetic variation

Note that a genome does not capture the genetic diversity or the genetic polymorphism of a species. For example, the human genome sequence in principle could be determined from just half the information on the DNA of one cell from one individual. To learn what variations in genetic information underlie particular traits or diseases requires comparisons across individuals. This point explains the common usage of "genome" (which parallels a common usage of "gene") to refer not to the information in any particular DNA sequence, but to a whole family of sequences that share a biological context.

Although this concept may seem counter intuitive, it is the same concept that says there is no particular shape that is the shape of a cheetah. Cheetahs vary, and so do the sequences of their genomes. Yet both the individual animals and their sequences share commonalities, so one can learn something about cheetahs and "cheetah-ness" from a single example of either.

Genome projects

For more details on this topic, see Genome project.

The Human Genome Project was organized to map and to sequence the human genome. Other genome projects include mouse, rice, the plant Arabidopsis thaliana, the puffer fish, bacteria like E. coli, etc. In 1976, Walter Fiers at the University of Ghent (Belgium) was the first to establish the complete nucleotide sequence of a viral RNA-genome (bacteriophage MS2). The first DNA-genome project to be completed was the Phage Φ-X174, with only 5368 base pairs, which was sequenced by Fred Sanger in 1977 . The first bacterial genome to be completed was that of Haemophilus influenzae, completed by a team at The Institute for Genomic Research in 1995.

In May 2007, the New York Times announced that the full genome of DNA pioneer James D. Watson had been recorded.[1] The article noted that some scientists believe this to be the gateway to upcoming personalized genomic medicine.

Many genomes have been sequenced by various genome projects. The cost of sequencing continues to drop.

Comparison of different genome sizes

Main article: Genome size
Organism Genome size (base pairs) Note
Virus, Bacteriophage MS2 3569 b First sequenced RNA-genome[2]
Virus, SV40 5224 b [3]
Virus, Phage Φ-X174; 5386 b First sequenced DNA-genome[4]
Virus, Phage λ 50 kb
Bacterium, Haemophilus influenzae 1.83 Mb First genome of living organism, July 1995[5]
Bacterium, Carsonella ruddii 160 kb Smallest non-viral genome, Feb 2007
Bacterium, Buchnera aphidicola 600 kb
Bacterium, Wigglesworthia glossinidia 700 kb
Bacterium, Escherichia coli 4 Mb [6]
Amoeba, Amoeba dubia 670 Gb Largest known genome, Dec 2005
Plant, Arabidopsis thaliana 157 Mb First plant genome sequenced, Dec 2000.[7]
Plant, Genlisea margaretae 63.4 Mb Smallest recorded flowering plant genome, 2006.[7]
Plant, Fritillaria assyrica 130 Gb
Plant, Populus trichocarpa 480 Mb First tree genome, Sept 2006
Yeast,Saccharomyces cerevisiae 20 Mb [8]
Fungus, Aspergillus nidulans 30 Mb
Nematode, Caenorhabditis elegans 98 Mb First multicellular animal genome, December 1998[9]
Insect, Drosophila melanogaster aka Fruit Fly 130 Mb [10]
Insect, Bombyx mori aka Silk Moth 530 Mb
Insect, Apis mellifera aka Honey Bee 1.77 Gb
Fish, Tetraodon nigroviridis, type of Puffer fish 385 Mb Smallest vertebrate genome known
Mammal, Homo sapiens 3.2 Gb
Fish, Protopterus aethiopicus aka Marbled lungfish 130 Gb Largest vertebrate genome known

Note: The DNA from a single human cell has a length of ~1.8 m (but at a width of ~2.4 nanometers).

Since genomes and their organisms are very complex, one research strategy is to reduce the number of genes in a genome to the bare minimum and still have the organism in question survive. There is experimental work being done on minimal genomes for single cell organisms as well as minimal genomes for multicellular organisms (see Developmental biology). The work is both in vivo and in silico.

Genome evolution

Genomes are more than the sum of an organism's genes and have traits that may be measured and studied without reference to the details of any particular genes and their products. Researchers compare traits such as chromosome number (karyotype), genome size, gene order, codon usage bias, and GC-content to determine what mechanisms could have produced the great variety of genomes that exist today (for recent overviews, see Brown 2002; Saccone and Pesole 2003; Benfey and Protopapas 2004; Gibson and Muse 2004; Reese 2004; Gregory 2005).

Duplications play a major role in shaping the genome. Duplications may range from extension of short tandem repeats, to duplication of a cluster of genes, and all the way to duplications of entire chromosomes or even entire genomes. Such duplications are probably fundamental to the creation of genetic novelty.

Horizontal gene transfer is invoked to explain how there is often extreme similarity between small portions of the genomes of two organisms that are otherwise very distantly related. Horizontal gene transfer seems to be common among many microbes. Also, eukaryotic cells seem to have experienced a transfer of some genetic material from their chloroplast and mitochondrial genomes to their nuclear chromosomes.

http://en.wikipedia.org/wiki/Genome

Thursday, March 27, 2008

Gene Therapy for Chronic Pain

Researchers use gene therapy to stop pain signals before they reach the brain.

The pain gate: When we suffer pain--whether from a stubbed toe or a metastasized tumor--pain signals are transmitted to the brain from around the body through these groups of sensory neurons, called dorsal root ganglia (DRG). A new gene-therapy technique intercepts pain signals at the DRG using a gene for a naturally produced opiate-like chemical. On the right, the cells of a rat's DRG glow green with a marker for the opiate-like gene one month after it was injected into the rat's spinal fluid. On the left are DRG cells from a control rat injected with saline solution.
Credit: PNAS

A new kind of gene therapy could bring relief to patients suffering from chronic pain while bypassing many of the debilitating side effects associated with traditional painkillers.

Researchers at Mount Sinai School of Medicine injected a virus carrying the gene for an endogenous opioid--a chemical naturally produced by the body that has the same effect as opiate painkillers such as morphine--directly into the spinal fluid of rats. The injections were targeted to regions of the spinal cord called the dorsal root ganglia, which act as a "pain gate" by intercepting pain signals from the body on their way to the brain. "You can stop pain transmission at the spinal level so that pain impulses never reach the brain," says project leader Andreas Beutler, an assistant professor of hematology and medical oncology at Mount Sinai.

The injection technique is equivalent to a spinal tap, a routine procedure that can be performed quickly at a patient's bedside without general anesthesia.

Because it targets the spinal cord directly, this technique limits the opiate-like substance, and hence any side effects, to a contained area. Normally, when opiate drugs are administered orally or by injection, their effects are spread throughout the body and brain, where they cause unwanted side effects such as constipation, nausea, sedation, and decreased mental acuity.

Side effects are a major hurdle in treating chronic pain, which costs the United States around $100 billion annually in treatment and lost wages. While opiate drugs can be very effective, the doses required to successfully control pain are often too high for the patient to tolerate.

"The side effects can be as bad as the pain," says Doris Cope, director of the University of Pittsburgh Medical Center's Pain Medicine Program. Achieving the benefits of opiate treatment without their accompanying side effects, Cope says, would be a "huge step forward."

Beutler hopes to do just that. "Our strategy was to harness the strength of opioids but target it to the pain gate, and thereby create pain relief without the side effects that you always get when you have systemic distribution of opioids," he says.

Several groups have previously attempted to administer gene therapy for pain through spinal injections, but they failed to achieve powerful, long-lasting pain relief. The new technique produced results that lasted as long as three months from a single injection, and unpublished follow-up studies suggest that the effect could persist for a year or more.

Beutler credits his team's success to the development of an improved virus for delivering the gene. The team uses a specially adapted version of adeno-associated virus, or AAV--a tiny virus whose genome is an unpaired strand of DNA. All the virus's own genes are removed, and the human endogenous opioid gene is inserted in their place. Beutler's team also mixed and matched components from various naturally occurring AAV strains and modified the genome into a double-stranded form. These tweaks likely allow the virus to infect nerve cells more easily and stick around longer.

http://www.technologyreview.com/Biotech/20118/


Friday, March 14, 2008

Synthesizing a Genome from Scratch

Scientists say the results represent a new stage in synthetic biology.

Synthetic genomes: Shown here is a circular piece of DNA synthesized from scratch--the first bacterial genome to be created this way.
Credit: J. Craig Venter Institute

In a technical tour de force, scientists at the J. Craig Venter Institute, in Rockville, MD, have synthesized the genome of the bacterium Mycoplasma genitalium entirely from scratch. The feat is a stepping stone in creating precisely engineered microbial machines capable of generating biofuels and performing other useful functions.

"It really is groundbreaking that you can synthetically build a genome for a bacterium," says Chris Voigt, a synthetic biologist at the University of California, San Francisco, who was not involved in the project. "It's bigger by orders of magnitude than what's been done before."

Biologists creating genetically engineered organisms now routinely order pieces of DNA that are 10,000 to 20,000 base pairs long--big enough to incorporate the genes for a single metabolic pathway. That allows researchers to engineer microbes that can perform specific tasks, but the ability to synthesize entire genomes could grant a whole new level of control over biological design. (See "Tumor-KillingBacteria.")

In the new study, scientists ordered 101 DNA fragments, encompassing the entire Mycoplasma genome, from commercial DNA synthesis companies. These fragments were designed so that each overlapped its neighboring sequence by a small amount; these overlapping stretches stick together, thanks to the chemical properties of DNA. Researchers then bound the fragments piece by piece, eventually generating the full 582,970 base pair Mycoplasma sequence. The findings were published Thursday in the online edition of Science.

"We consider this a second and significant step in a three-step process of our attempt to create the first synthetic organism," says Craig Venter, president of the Venter Institute. Venter and his colleagues ultimately want to create a minimal genome--one with the least number of genes needed to sustain life. Pinpointing the minimal genome will both shed light on key cellular processes and provide a base for designing sophisticated synthetic organisms. "We ultimately want to design cells that could function in a robust fashion to make unique biofuels," says Venter.

The researchers' next step will be to show that the synthetic genome functions as it should. "We have the whole genome assembled in a tube, but we need to transplant it into the cell of a different species to show that it can reboot the cell," says Hamilton Smith, a Nobel laureate who oversaw the project at the Venter Institute. Last year, Smith's group transplanted the genome of one species of Mycoplasma into another, demonstrating that this type of transplant is possible. (See "Transplanting a Genome.")

While the synthesis of a genome might be impressive from a scientific perspective, it is not yet a practical way to engineer microbes to make biofuels. Instead, several companies, including Synthetic Genomics, a biotech company founded by Venter to engineer microbes for energy, are using more traditional metabolic engineering techniques to generate fuel-producing bacteria. (See "Building Better Biofuels.") "What we're doing with synthetic chromosomes will be the design process for the future," says Venter.

Others in the field are excited about that prospect. "Being able to synthesize genomes opens up a new world," says Voigt. "You can build things on the scale of the genome." For example, he says, scientists are now engineering bacteria to perform different steps in the conversion of biomass into ethanol--one strain to break down the biomass, another to make ethanol. But ideally, scientists could put those processes together to create one organism that could eat biomass and spit out fuel. (See "The Price of Biofuels.") "That would require genome-scale design," Voigt says.

He likens the current project, which required multiple steps to glue the fragments together, to the last computers designed before automated manufacturing and microfabrication techniques were introduced. Similar advances are needed for more ambitious genome-synthesis projects. "We still need to develop 'one step' genome construction methods in order to reduce the costs and turn time of genome construction," says Drew Endy, a synthetic biologist at MIT.

http://www.technologyreview.com/Biotech/20112/

1,000 Genomes

Gene-sequencing projects keep getting bigger.

In a testament to the steady plummet in sequencing costs, today the National Human Genome Research Institute (NHGRI) announced a massive international collaboration to sequence the genomes of 1,000 people from around the world.

According to the NHGRI statement,

"The 1000 Genomes Project will examine the human genome at a level of detail that no one has done before," said Richard Durbin, Ph.D., of the Wellcome Trust Sanger Institute, who is co-chair of the consortium. "Such a project would have been unthinkable only two years ago. Today, thanks to amazing strides in sequencing technology, bioinformatics and population genomics, it is now within our grasp. So we are moving forward to build a tool that will greatly expand and further accelerate efforts to find more of the genetic factors involved in human health and disease."

The findings should give added power to the recent wave of studies identifying specific genetic risk factors for common health problems, such as diabetes, heart disease, lupus, and others. (See "Genes for Several Common Diseases Found.")

According to NHGRI director Francis Collins,

"This new project will increase the sensitivity of disease discovery efforts across the genome five-fold and within gene regions at least 10-fold. Our existing databases do a reasonably good job of cataloging variations found in at least 10 percent of a population. By harnessing the power of new sequencing technologies and novel computational methods, we hope to give biomedical researchers a genome-wide map of variation down to the 1 percent level. This will change the way we carry out studies of genetic disease."

Like previous international sequencing projects, the data will be made available for analysis in free public databases. Once scientists identify part of the genome associated with a particular disease, they will be able to look up that area of the genome in the database to find a list of gene variants in that region.

The project will be a huge technological feat; to date, only three human genomes have been sequenced.

From NHGRI:

The project depends on large-scale implementation of several new sequencing platforms. Using standard DNA sequencing technologies, the effort would likely cost more than $500 million. However, leaders of the 1000 Genomes Project expect the costs to be far lower--in the range of $30 million to $50 million--because of the project's pioneering efforts to use new sequencing technologies in the most efficient and cost-effective manner.

In the first phase of the 1000 Genomes Project, lasting about a year, researchers will conduct three pilots. The results of the pilots will be used to decide how to most efficiently and cost effectively produce the project's detailed map of human genetic variation.

The first pilot will involve sequencing the genomes of two nuclear families (both parents and an adult child) at deep coverage that averages 20 passes of each genome. This will provide a comprehensive dataset from six people that will help the project figure out how to identify variants using the new sequencing platforms, and serve as a basis for comparison for other parts of the effort.

The second pilot will involve sequencing the genomes of 180 people at low coverage that averages two passes of each genome. This will test the ability to use low-coverage data from new sequencing platforms to identify sequence variants and to put them in their genomic context.

The third pilot will involve sequencing the coding regions, called exons, of about 1,000 genes in about 1,000 people. This is aimed at exploring how best to obtain an even more detailed catalog in the approximately 2 percent of the genome that is comprised of protein-coding genes.

During its two-year production phase, the 1000 Genomes Project will deliver sequence data at an average rate of about 8.2 billion bases per day, the equivalent of more than two human genomes every 24 hours. The volume of data--and the interpretation of those data--will pose a major challenge for leading experts in the fields of bioinformatics and statistical genetics.

The 1,000 volunteers will be selected from those who participated in the HapMap project, a map of common genetic variation (see "A New Map for Health"), and will include:

Yoruba in Ibadan, Nigeria; Japanese in Tokyo; Chinese in Beijing; Utah residents with ancestry from northern and western Europe; Luhya in Webuye, Kenya; Maasai in Kinyawa, Kenya; Toscani in Italy; Gujarati Indians in Houston; Chinese in metropolitan Denver; people of Mexican ancestry in Los Angeles; and people of African ancestry in the southwestern United States.

http://www.technologyreview.com/blog/editors/22007/