News Menu

Mathematical technique de-clutters cancer-cell data, revealing tumor evolution, treatment leads

Krasnitz Wigler heatmap
When genome sequence data from 100 cells sampled from a single human tumor is analyzed, and the mathematical algorithm devised by Krasnitz and Wigler is applied, the rich structure of the data emerges. This is a 'heat map' in which each horizontal row contains data from 1 of the 100 sampled cells; and each vertical column contains information about the presence (black) or absence (no mark) of a 'CORE.' Each core represents a place in the genome where a particular cell either has amplified DNA (blue bar, top) or deleted DNA (red bar, top). From the mass of data underlying these phenomena, signatures of 4 subpopulations of tumor cells now become visible. The four groups and their evolutionary relation is shown along the left vertical axis: about half are 'green,' and their genomes are the least changed by the disease; cells in the remaining three subpopulations harbor multiple mutations in their genomes.

Cold Spring Harbor, NY — In our daily lives, clutter is something that gets in our way, something that makes it harder for us to accomplish things.聽For doctors and scientists trying to parse mountains of raw biological data, clutter is more than a nuisance; it can stand in the way of figuring out how best to treat someone who is very sick.

Using increasingly cheap and rapid methods to read the billions of 鈥渓etters鈥 that comprise human genomes—including the genomes of individual cells sampled from cancerous tumors—scientists are generating far more data than they can easily interpret.

Today, two scientists from 黑料吃瓜资源 (CSHL) publish a mathematical method of simplifying and interpreting genome data bearing evidence of mutations, such as those that characterize specific cancers.聽Not only is the technique highly accurate; it has immediate utility in efforts to parse tumor cells, in order to determine a patient鈥檚 prognosis and the best approach to treatment.

CSHL Assistant Professor Alexander Krasnitz, who developed the new technique jointly with American Cancer Society Professor Michael Wigler, explains that it reduces the burden of interpretation by identifying what he and Wigler call COREs, an acronym for 鈥渃ores of recurrent events.鈥

Consider the example of a cancerous breast tumor. Central to the CORE concept is what Krasnitz and Wigler refer to as 鈥渋ntervals.鈥澛 An example of an interval would be a segment of DNA that is missing in the genetic sequence of one or more cells sampled from the tumor. Tumor cells are often missing DNA that should normally be present; or conversely, they often have genome intervals in which the normal DNA sequence is amplified—it appears in multiple copies.聽 Such deletions and amplifications are called copy-number variations, or CNVs.

鈥淚n cancer,鈥 says Krasnitz, 鈥渨e find intervals in the genome that are hit again and again. You might see this in many cells coming from a single patient鈥檚 tumor; or you may see these repeating patterns in cells sampled from many patients with a similar cancer type.鈥

In either case, if you superimpose the location of each 鈥渉it鈥—whether a deletion or an amplification of DNA—against a map of the full human genome, 鈥測ou end up with these wobbly pile-ups, stacks of 鈥榟its鈥 at the same locations in the genome.鈥

Due to the vagaries of collecting genome data and a certain amount of small-scale variation in the precise boundaries of the deleted or amplified DNA intervals, the stacks don鈥檛 line up straight; as Krasnitz says, they look 鈥渨obbly.鈥澛 This makes them very hard to accurately interpret.

The CORE method he and Wigler describe in a paper appearing in Proceedings of the National Academy of Sciences 鈥渋s a mathematical way of cleaning up this mess and untangling these stacks of data, which often overlap.鈥澛燱hen data from 100 cells from a single tumor are analyzed, for example, and the mathematical algorithm devised by Krasnitz and Wigler is applied, the regularity of the stacks is revealed, and the rich structure of the data emerges.

In the example of analyzing 100 cells from one tumor, the net result is that populations and subpopulations of cancer cells can be distinguished; and if the cancer has already become metastatic, CORE will be useful in discerning the relations among cancer cell subpopulations in various parts of the body. Such analysis is a potentially valuable guide to prognosis and can also help to make important treatment decisions.

Written by: Peter Tarr, Senior Science Writer | [email protected] | 516-367-8455

Citation

鈥淭arget inference from collections of genomic intervals鈥 appears online today ahead of print in Proceedings of the National Academy of Sciences.聽The authors are: Alexander Krasnitz, Guoli Sun, Peter Andrews and Michael Wigler. The paper can be obtained online at:

Stay informed

Sign up for our newsletter to get the latest discoveries, upcoming events, videos, podcasts, and a news roundup delivered straight to your inbox every month.

  Newsletter Signup

About 黑料吃瓜资源

Founded in 1890, 黑料吃瓜资源 has shaped contemporary biomedical research and education with programs in cancer, neuroscience, plant biology and quantitative biology. Home to eight Nobel Prize winners, the private, not-for-profit Laboratory employs 1,000 people including 600 scientists, students and technicians. The Meetings & Courses Program annually hosts more than 12,000 scientists. The Laboratory鈥檚 education arm also includes an academic publishing house, a graduate school and the DNA Learning Center with programs for middle, high school, and undergraduate students and teachers. For more information, visit www.cshl.edu