Base Pairs podcast
Did you know? If unwound and tied together, the strands of DNA in one cell would stretch past the entire length of you body (~ 6 ft). Now imagine that among all that, only enough genetic information to run the length of your thumb-nail is of any importance!
Soon after the Human Genome Project unveiled the sequence of our DNA, many scientists hypothesized that only a small portion—about two percent—of the human genome actually encodes the proteins that make our bodies work. The rest, they said, was 鈥渏unk.鈥 This stunning assertion became a widely known bit of trivia—one that led many experts to focus only on what they referred to as 鈥済enes,鈥 the protein-coding portion of the genome.
In this new episode of Base Pairs, co-hosts Brian and Andrea talk to David Spector and Thomas Gingeras, two CSHL professors who have realized—in their own unique ways—just how wrong this assumption has been.
AA: And I鈥檓 Andrea
BS: And welcome to season two of Base Pairs!
AA: This episode is actually the second half of a two-part season premiere, so if you鈥檙e interested in hearing the whole story, we highly recommend you check out our episode from March of 2017. It鈥檚 titled Dark Matter of the Genome, Part 1.
BS: And in that episode, we talked to CSHL Assistant Professor Molly Hamell 鈥 who coincidently used to be an astrophysicist 鈥 about the 鈥淒ark Matter鈥 of the genome. The chunk of the genome鈥
AA: About 98% actually!
BS: 鈥攖he HUGE chunk that in fact does not code for proteins.
BS: Now before we go any further, I wanted to invite our listeners to join me for a step back into the past 鈥 when I was just an awkward high school student trying to pass his biology class.
(classroom chatter)
AA: Oh, my. I鈥檓 not sure you鈥檙e going to keep an audience for this one, Brian. Are we going to have to hear all about your high school angst?
BS: No鈥 no, no. Nothing like that. But like most of our listeners, I loved science, even then鈥 even if I was sleeping through some of the lectures. And a favorite wacky science fact my teachers 鈥 and even the textbooks! 鈥 tended to repeat was that more than 90% of the genome was鈥 junk. 鈥淛unk DNA,鈥 is what everyone called it. And it stuck. I remembered that. Even when I turned my attention away from biology to space science, and animal science, and climate change, and medical reporting鈥 that crazy statistic stuck with me.
AA: Well, that鈥檚 why we prefer the term 鈥淕enomic dark matter鈥 so much more. Instead of seeing 鈥渏unk,鈥 the world is finally learning to look the non-coding part of the genome 鈥 the part that doesn鈥檛 contain the recipe for proteins 鈥 as the exciting unknown, an area worth investigating.
BS: Right, right鈥 but imagine being a scientist interested in that unknown back when I was in high school. When everyone, even your esteemed colleagues, also pushed the 鈥渏unk DNA鈥 concept. At the time, people researching that ignored part of the genome must have looked a little like鈥 well鈥 garbage pickers, or maybe trivia freaks.
TG: There was sort of a negative connotation associated with any region of the genome that was not engaged in making proteins. And, it acquired a label of Junk DNA, in the literature, and therefore, this pejorative concept really stayed with us and focused the attention of most people on these protein coding regions, without the need for studying these other regions.
BS: That鈥檚 Tom Gingeras, a CSHL professor who can be found carefully scrutinizing genomes at the aptly named Genome Research Center.
TG: Even when I was an undergraduate, one of the major topics that was being pursued by a lot of different labs, and this will sound somewhat strange given where we are today, was that where were genes in DNA? Obviously, we had an understanding that DNA was the genetic molecule that鈥檚 passed from one cell type to another, and from one generation to another. But we didn鈥檛 realize that a unit called a gene 鈥 What was it? How was its structure identifiable in gene, in a genome? How was its expression, how was its information turned on at any one particular time, in one cell type versus another? These are questions that continue to be pursued in science. But, in the days when I was a student, we knew virtually nothing about it. That鈥檚 been sort of a probing question for me for many, many years.
AA: Oh wow. I can see why he鈥檇 want to be a part of that. It鈥檚 the one drive A LOT of scientists share: a desire to break new ground in the search for information about鈥 everything!
BS: That鈥檚 it, and to make those new discoveries, Tom found himself leading projects were he and his lab wadded and sifted through the genome鈥檚 so-called junk, one nucleotide at a time. It was all part of the ENCODE project-
AA: That鈥檚 E-N-C-O-D-E 鈥 an international project consortium first established by the U.S. government in 2003 to examine the recently decoded human genome sequence in greater depth.
BS: Yup. And they were spending most of their time studying the areas occupied by what鈥檚 considered 鈥済enes鈥 鈥 that 2% of the genome that leads the production of proteins.
TG: So, we鈥 looked at two chromosomes. Chromosomes 21 and 22. And we walked along the genome roughly every five bases. Is there an RNA that contains that base as encoded in the genome? And then you鈥檇 walked to the next five and ask, 鈥淚s there an RNA coming from there?鈥
AA: To remind our listeners: RNA is the short-lived cousin of DNA. Scientists first learned about RNA because it carries the instructions for making proteins from the DNA to the cell鈥檚 protein factories.
BS: Right, and with that understanding, most scientists looked for the presence of functioning RNA as a sign of activity coming from the protein-coding parts of a genome. In this way, many hoped to figure out where exactly each gene is on a chromosome, and what it鈥檚 up to. Tom figured that since only 2 percent of the genome was genes, only about 2 percent of each chromosome would be written in an RNA鈥攐r transcribed, as scientists say.
TG: My post-doc at the time, Phil Kapranov, came back with a result, which seemed to me entirely wrong because it went against the very simple principle that the only place that we probably should be seeing things are where known genes are鈥 Now, you might see new known genes, and so, all right, so you would increase the number of spots that would be functional in that fashion. But, instead, we saw almost a totality of those two chromosomes being transcribed. So, I got very upset with Phil, because he obviously did the wrong experiment and what he did was wrong. I asked him, 鈥淧lease, go back and do this again.鈥
BS: So. Phil did it again, and again, and again 鈥 each time with the same result.
TG: The exact same result鈥
AA: Ok. So NOW I鈥檝e got to know. What鈥檚 going on here?
BS: You know? That鈥檚 the thing. Tom had no idea. As far as the scientific community was concerned, Tom鈥檚 lab had been sifting through junk, looking for treasures. But now it seemed that EVERYTHING was treasure! The genome in those 2 chromosomes was expressing massive numbers of RNA messages, or transcripts, as they are called by scientists.
TG: To add to the complexity鈥 this incredibly rich output of RNA from every region of the genome indicated that there was a lot of energy being 鈥渟pent鈥 by cells being put into making these RNAs, which apparently had no function since many of them鈥 most of them鈥 were well outside the 2% of the genome that encodes protein.
AA: To blow so much energy on making useless RNA鈥 that doesn鈥檛 sound like the efficiency we鈥檝e come to associate with cells. Nature, as many experts will tell you, generally has a 鈥渨aste not鈥 kind of attitude.
BS: Exactly. After seeing examples of this phenomena again and again, Tom and his team decided to change how they perceived the genome. Almost overnight, they became those so-called 鈥済arbage picking鈥 scientists I mentioned earlier. And in doing so, they discovered something that really abolishes the 鈥渏unk DNA鈥 mythos.
TG: What we determined is that roughly 80% of the genome is transcribed鈥 Not every cell makes 80% of its genome, but if you look at a totality of many, many different cell types, what is capable of being transcribed is about 80% of the genome. Only 2% of the genome is actually embedded within a protein coding region. So, that meant that a majority of the genome was transcribed and that most of it was making RNAs that were not intended to make a protein.
AA: 80%鈥 Wow. So鈥 we know today that there鈥檚 around 20,000 protein-coding gene regions. Scientists even today are still trying to identify what each gene does. Tom鈥檚 work is exciting stuff 鈥 don鈥檛 get me wrong 鈥 but can鈥檛 we just鈥 I don鈥檛 know鈥 shelve it until we get the protein coding parts figured out?
BS: (laughs) Well鈥nitially, that鈥檚 exactly what happened. A lot of scientists simply didn鈥檛 accept Tom鈥檚 results, saying that Tom, Phil, and rest must have been mistaken.
TG: 鈥樷ither your technology doesn鈥檛 work, or B, it doesn鈥檛 make any difference because it鈥檚 all noise anyway because it鈥檚 not a perfect system.鈥
BS: And then, even among the people who did believe it, they argued that 鈥 like you said 鈥 it鈥檚 just too much. They preferred to focus on what the Human Genome Project had made known, and suggested saving the unknown for a rainy day.
AA: But then鈥 something happened, didn鈥檛 it?
TG: Slowly but surely, what we saw was that these RNAs were seen to be important regions, that if they were mutated, deleted, you would see an effect. A phenotypic effect.
BS: By 鈥減henotypic effect,鈥 Tom means, simply, a physical manifestation in living creatures 鈥 things that made them look different, or in some cases, made them sick. Our first sense of the scope of these changes came about 5 years ago, when teams from 32 institutes in five countries 鈥 all contributing to the ENCODE project 鈥 released 30 papers at once.
AA: Not exactly light reading.
BS: Hehe. Well, it was important stuff! Previously, it had always been assumed the only way to affect the physical expression of traits was to mess with a gene, or 鈥 as we described in the previous episode 鈥 inhibit the regulation of transposable elements. But a lot of the ENCODE data showed that even if you messed with only the non-protein coding parts of the genome, it could affect cells, and whole organisms, in a big way.
AA: Oh wow鈥 that is important. I take it the scientific community is a bit more amenable to Tom鈥檚 results now, right?
BS: Well鈥
TG: I鈥檒l answer that with a personal vignette. So, I would go to meetings and we would present our data. That almost invariably ignited a discussion, no matter how large the audience was. I鈥檝e been in places where there鈥檝e been more than a thousand people, and then people would have no problem jumping up and saying, 鈥淭his is all nonsense. And even if it is all true, we are totally swamped with just learning about the things we know about, instead of dealing with all this other nonsense.鈥 You would have this back and forth. I would go back home feeling very upset. Feeling, 鈥渨hy can鈥檛 they see? Why can鈥檛 they understand?鈥 And I would say, 鈥淲e鈥檙e not incompetent. So why is that we鈥檙e having this discussion?鈥 And I would get very personally worked up. Until one day I honestly had what might be called, in sort of a religious sense, this enlightenment. Seriously. Sitting in my office. Which, basically, led me to understand that you need to let them discover it鈥
I鈥檓 sure that people will do these experiments in their own way in their own systems. And, if they see the same thing then perhaps you鈥檒l understand that there鈥檚 a lot of work to be done. And, that the things we thought were solved, in fact, are not solved, and they offer an opportunity to learn even more.
DS: 鈥 Yes, 7SK. Yeah. That kind of started our interest and then it just expanded from there. Now it鈥檚 pretty much the main focus of my entire lab. We do very little imaging at this point. Mostly we鈥檙e doing RNAseq and try to figure out the function of these long non-coding RNA鈥檚.
AA: That鈥檚 Professor David Spector 鈥 and he鈥檒l explain 鈥7SK鈥 in a minute鈥.
BS: David was once one of 鈥渢hem.鈥 Those doubting scientists Tom was talking about. Not necessarily a critic of ENCODE鈥檚 work. Just a scientist so focused on his own work that he didn鈥檛 pay much attention to the non-coding part of the genome.
AA: You see, David is the Director of Research here at CSHL, so we know a lot about him鈥
BS: And we know he鈥檚 REALLY adamant about speckles.
(Clip from the holiday party) (laughter fades)
DS: My lab is probably best known for the early work that we did on nuclear organization and function. I started that work probably back in 1981. We were very intrigued by how the nucleus might be organized, given the fact there are no membrane-bound structures in the nucleus. Yet if you stain cells with either dyes or with antibodies there are clearly concentrations of specific proteins in different places in the nucleus鈥 So, I got really interested in studying these structures and spent a significant amount of my scientific research career on a particular nuclear structure called nuclear speckles鈥
鈥n fact, some people call them Spector speckles as a joke. We spent a lot of time working on them. We purified them biochemically. Did proteomic analysis, identified 146 proteins in them. We did live cell imaging of them and just an enormous amount of work.
BS: So much work. Back then, the majority of David鈥檚 lab was dedicated to this subject, but that started to change when he had his own run-in with the unknown.
DS: 鈥7SK. It was a known RNA at the time鈥
BS: Ok. So not THAT unknown. But it was still a non-coding RNA, and the fact that it was jumbled in there 鈥 hiding in Spector鈥檚 speckles 鈥 that caught his attention.
DS: that kind of tipped the bucket in a way from proteins to RNA because we were intrigued by the fact that there was such a high concentration of this non-coding RNA in nuclear speckles.
AA: I see. Here鈥檚 this important region of the cell, and David found himself wondering, 鈥渨hy is it wasting energy making so much of this supposedly useless RNA?鈥
BS: That鈥檚 the ticket. But David鈥檚 foray into the Dark Matter of the genome didn鈥檛 stop with just one RNA.
DS: One day we got a call in the lab from a student in a lab in France
BS: 鈥 her name is Delphine Bernard 鈥
DS: She was doing a study in neurons and鈥 She came upon a long non-coding RNA and her adviser was telling her to drop it and forget about it and focus on what his lab was interested in. She persisted and she was curious about it and because my lab had been working on long non-coding RNA鈥檚 she decided to call us and see what do we think? So, she started to tell me about this RNA on the phone and it sounded really exciting so we told her absolutely don鈥檛 drop it. We started a collaboration with her and then we ended up having a paper together in EMBO Journal. Her PI bought into it very strongly, so he was a convert. (laughter) Anyway, so that kind of got us interested. The RNA happened to be MALAT1.
AA: Ok. That sounds VERY familiar. MALAT 1 is a puzzle, even among genomic dark matters.
DS: MALAT 1 in essence breaks the rules for all long non-coating RNA鈥檚. First rule, most long non-coding RNA鈥檚 are present in very few copies per cell. MALAT 1 is widely abundant in cells. In fact, in many cell culture lines it rivals鈥 one of the most highly expressed transcript in tissue culture cells
AA: And that鈥 that would indicate that MALAT 1 is important.
DS: It鈥檚 got to have a really amazing function.
AA: VERY important.
BS: And yet, when they studied mice that couldn鈥檛 produce MALAT 1鈥
DS: At the end of the day, the mouse was perfectly happy without MALAT 1.
AA: What?!
BS: Yeah. Perfectly healthy. David鈥檚 lab even sent a number of tissue samples to folks over at Yale.
DS: They looked at all the tissues, couldn鈥檛 find anything abnormal about them. The mice have been in my lab now breeding perfectly well for probably close to seven years.
AA: Somehow, I doubt this was a complete bust. It seems highly unlikely that cells mass-produce MALAT1 for fun.
BS: that鈥檚 what David thought too. He figured that MALAT1鈥檚 function may in-fact be more important than he and his lab presumed 鈥 SO important that healthy cells have back-up systems in place if the RNA鈥檚 production fails in one part of the genome.
DS: Given that, we said 鈥渙kay we need to take a different approach鈥︹
AA: So, David and his team needed to study cells that were omnivorous 鈥 hungry for all things good for them鈥 (Pause) that sounds like鈥
BS: Yup. Cancer. Aggressive tumor cells are like regular cells gone rogue, multiplying rapidly and sapping up important resources in the otherwise organized world of a living thing. Tumor cells that travel from one place to another 鈥 such as breast cancers that invade the lung 鈥 are called metastatic 鈥 and you can often blame their bad behavior, at least in part, on a mangled genome.
AA: So mangled that the failsafe in place of MALAT1 might not work. I see. Take MALAT1 away from a tumor, and something might happen.
BS: Something dramatic.
CBSNY: 鈥淎 possible breakthrough tonight in the war on Breast cancer鈥︹ CG: 鈥淚t was a eureka moment when a new drug he and colleagues invented, chewed up and destroyed aggressive metastic breast cancer cells.鈥 DS: 鈥淭he effect was so dramatic that it was not something I could have actually predicted鈥︹
AA: That was David talking to Carolyn Gusoff of CBS2 New York just last year 鈥 and his words鈥 they鈥檙e not the kind that scientists say lightly. The effect REALLY was dramatic. So much so, that David has microscopy photos of the experiment framed on his desk.
(in interview) AA: so, I鈥檓 looking at this picture that鈥檚 right behind you. So that kind of captures the difference that you saw between鈥 So can you describe that a little bit, like what it looks like? What you saw that was so different about the chamber without MALAT 1?
DS: When we looked at the tumors in this mouse model that had MALAT 1, the tumors were very aggressive tumors which means that the tumor was just filled with cancer cells.
BS: I think the best way to describe this is to imagine a white sponge, except鈥 there are LOTS of wide holes in this sponge and each of them is PACKED full of this pink鈥 stuff.
AA: Those are the cancer cells. The pink stuff.
DS: When we got rid of MALAT 1, the tumor totally changed. A lot of the cancer cells died and were released and the tumor changed its whole gene expression pattern and it started to form these cystic structures which were formed because cells died, leaving these cysts.
BS: Interestingly, these large cavities, once rife with aggressive pink cancer cells, they鈥檙e now full of something much more benign: Milk.
DS: these cysts are filled with liquid 鈥 milk proteins and so what does that mean? If you think about it, this is a tumor in a mammary gland and at a particular stage a mammary gland can produce milk. What we鈥檝e done is by taking away MALAT 1 we鈥檝e changed that tumor from being this aggressive tumor, to 鈥
AA: To one that has switched its focus to milk production. You can see these images for yourself at our LabDish blog to really feel that 鈥渨ow factor鈥 for yourself. And David and his colleague think it鈥檚 safe to assume that if they turn off MALAT1 in tumors in other parts of the body, the cancer cells will be replaced with other types on harmless proteins.
BS: The coolest thing about all of this is that MALAT1 is just one of over 17,000 known non-coding RNAs. Few of these are as prevalent as MALAT1, and even fewer are seemingly as crucial, but there is an awful lot about non-coding RNAs that we still don鈥檛 know.
AA: And that 17,000 is just the ones recorded so far. (pause) According to David, by targeting MALAT1 and parts of the genome like it, experts might be able to open up new avenues for personalized treatment options in cancer and possibly other illnesses.
DS: The long-term goal of this over the next ten years, let鈥檚 say, is to develop a precision medicine based approach whereby we could get a small piece of the patient鈥檚 tumor. We could put it through a battery of tests鈥 and screen them for these non-coding RNAs and identify which patients would benefit most from knocking down these three RNAs versus another three. That鈥檚 kind of what we hope to do because every patient鈥檚 tumor will be different.
BS: Since David鈥檚 work has taken off, MALAT1 has really become a poster child for the importance of what was once thought of as 鈥渏unk DNA.鈥 And according to Tom Gingeras, the opinion of the scientific community is changing too.
TG: Most of the scientific community, at least I come in contact with, have an appreciation for the fact that these regions we thought were relatively inactive, that is to say, nonfunctional, probably contain some proportion of things which are of biological note. There鈥檚 an argument about how much that is, and then there鈥檚 an argument exactly how important. But, there is, I think, a general appreciation that it鈥檚 not something we just slide under the carpet.
AA: So鈥 that鈥檚 it! While research into the non-coding part of the genome is just getting started, this is where our story is going to have to stop.
Extras for Episode 9
Explaining ENCODE
Professor Thomas Gingeras is not only a leader of the ENCODE project, but also the mouseENCODE and modENCODE (model genome ENCODE) projects of the National Institutes of Health. His research has altered our understanding of the traditional boundaries of genes, but what is an ENCODE project all about anyway? Below is a pair of videos, produced by Nature in 2012, that help explain:
Ten years on from the 鈥檚 , 鈥渢he next colossal chapter in your story鈥 begins: an ENCyclopedia of DNA Element, or ENCODE for short.
ENCODE’s lead coordinator, Ewan Birney, and Nature editor Magdalena Skipper talk about the challenges of managing this colossal project and what we’ve learnt about our genomes.
You can view all the ENCODE results, including the 30+ papers we mention in this episode by visiting Nature鈥檚 .
Seeing white
And remember how we promised we鈥檇 show you that photo of Dave鈥檚 breathtaking results from when he investigated Malat1? Well, here they are:

You can learn more about these incredible results at: /unusual-drug-target-and-drug-generate-exciting-preclinical-results-in-mouse-models-of-metastatic-breast-cancer/
鈥淒ifferentiation of Mammary Tumors and Reduction in Metastasis Upon Malat1 LncRNA Loss鈥 appeared online in Genes & Development on December 23, 2015. It was authored by聽Gayatri聽Arun, Sarah Diermeier, Martin Akerman, Kung-Chi Chang, J. Erby Wilkinson, Stephen Hearn, Youngsoo Kim, A. Robert MacLeod, Adrian R. Krainer, Larry Norton, Edi Brogi, Mikala Egeblad and David L Spector.
Written by: Brian Stallard, Content Developer/Communicator | [email protected] | 516-367-8455
