News Menu

Re-learning how to read a genome

K562 cell line data set
A UCSC genome browser shot of the globin locus near the LCR using K562 cell line data sets generated or used in this study.

Study suggests a unified model for how DNA is read, offering insight into how genes evolve

Cold Spring Harbor, NY — There are roughly 20,000 genes and thousands of other regulatory 鈥渆lements鈥 stored within the three billion letters of the human genome. Genes encode information that is used to create proteins, while other genomic elements help regulate the activation of genes, among other tasks. Somehow all of this coded information within our DNA needs to be read by complex molecular machinery and transcribed into messages that can be used by our cells.

New research has revealed that the initial steps of reading DNA are actually remarkably similar at both the genes that encode proteins (here, on the right) and regulatory elements (on the left). The main differences seem to occur after this initial step. Gene messages are long and stable enough to ensure that genes become proteins, whereas regulatory messages are short and unstable, and are rapidly ‘cleaned up’ by the cell.

Usually, reading a gene is thought to be a lot like reading a sentence. The reading machinery is guided to the start of the gene by various sequences in the DNA—the equivalent of a capital letter—and proceeds from left to right, DNA letter by DNA letter, until it reaches a sequence that forms a punctuation mark at the end. The capital letter and punctuation marks that tell the cell where, when, and how to read a gene are known as regulatory elements.

But scientists have recently discovered that genes aren鈥檛 the only messages read by the cell. In fact, many regulatory elements themselves are also read and transcribed into messages, the equivalent of pronouncing the words 鈥渃apital letter,鈥 鈥渃omma,鈥 or 鈥減eriod.鈥 Even more surprising, genes are read bi-directionally from so-called 鈥渟tart sites鈥—in effect, generating messages in both forward and backward directions.

With all these messages, how does the cell know which one encodes the information needed to make a protein? Is there something different about the reading process at genes and regulatory elements that helps avoid confusion? New research, published today in Nature Genetics, has revealed that the initial steps of the reading process itself are actually remarkably similar at both genes and regulatory elements. The main differences seem to occur after this initial step, in the length and stability of the messages. Gene messages are long and stable enough to ensure that genes becomes proteins, whereas regulatory messages are short and unstable, and are rapidly 鈥渃leaned up鈥 by the cell.

To make the distinction, the team, which was co-led by CSHL Professor Adam Siepel and Cornell University Professor John Lis, looked for differences between the initial reading processes at genes and a set of regulatory elements called enhancers. 鈥淲e took advantage of highly sensitive experimental techniques developed in the Lis lab to measure newly made messages in the cell,鈥 says Siepel. 鈥淚t鈥檚 like having a new, more powerful microscope for observing the process of transcription as it occurs in living cells.鈥

Remarkably, the team found that the reading patterns for enhancer and gene messages are highly similar in many respects, sharing a common architecture. 鈥淥ur data suggests that the same basic reading process is happening at genes and these non-genic regulatory elements,鈥 explains Siepel. 鈥淭his points to a unified model for how DNA transcription is initiated throughout the genome.鈥

Working together, the biochemists from Lis鈥檚 laboratory and the computer jockeys from Siepel鈥檚 group carefully compared the patterns at enhancers and genes, combining their own data with vast public data sets from the NIH鈥檚 Encyclopedia of DNA Elements (ENCODE) project. 鈥淏y many different measures, we found that the patterns of transcription initiation are essentially the same at enhancers and genes,鈥 says Siepel. 鈥淢ost RNA messages are rapidly targeted for destruction, but the messages at genes that are read in the right direction—those destined to be a protein—are spared from destruction.鈥 The team was able to devise a model to mathematically explain the difference between stable and unstable transcripts, offering insight into what defines a gene. According to Siepel, 鈥淥ur analysis shows that the 鈥榗ode鈥 for stability is, in large part, written in the DNA, at enhancers and genes alike.鈥

This work has important implications for the evolutionary origins of new genes, according to Siepel. 鈥淏ecause DNA is read in both directions from any start site, every one of these sites has the potential to generate two protein-coding genes with just a few subtle changes. The genome is full of potential new genes.鈥

Written by: Jaclyn Jansen, Science Writer | [email protected] | 516-367-8455


Funding

This work was supported by the .

Citation

鈥淎nalysis of transcription start sites from nascent RNA identifies a unified architecture of initiation regions at mammalian promoters and enhancers.鈥 appears online in Nature Genetics on November 10, 2014. The authors are: Leighton Core, Andr茅 Martins, Charles Danko, Colin Waters, Adam Siepel, and John Lis. The paper can be obtained online at:

Stay informed

Sign up for our newsletter to get the latest discoveries, upcoming events, videos, podcasts, and a news roundup delivered straight to your inbox every month.

  Newsletter Signup

About 黑料吃瓜资源

Founded in 1890, 黑料吃瓜资源 has shaped contemporary biomedical research and education with programs in cancer, neuroscience, plant biology and quantitative biology. Home to eight Nobel Prize winners, the private, not-for-profit Laboratory employs 1,000 people including 600 scientists, students and technicians. The Meetings & Courses Program annually hosts more than 12,000 scientists. The Laboratory鈥檚 education arm also includes an academic publishing house, a graduate school and the DNA Learning Center with programs for middle, high school, and undergraduate students and teachers. For more information, visit www.cshl.edu

Principal Investigator

Adam Siepel

Adam Siepel

Professor
Cancer Center Member
Ph.D., University of California, Santa Cruz, 2005

Tags