A 鈥渂ehind the screens鈥 look at how biology is addressing its 鈥渕ost wonderful problem鈥濃攖oo much data. Associate Professor Gurinder S. 鈥淢ickey鈥 Atwal joins us to explain the essential enigma that is quantitative biology.
Read the related story: Biology, behind the screens
Transcript
(Chatter sounds trickle in, growing steadily louder)
BS: Hey guys鈥 I鈥檓 Brian
AA: And I鈥檓 Andrea鈥
BS: And we鈥檙e needing to almost shout here because well鈥 this noise. This raucous crowd you鈥檙e hearing behind us鈥 that鈥檚 hundreds of high schoolers.
AA: 312 high school students to be exact.
BS: And despite this happening at the cusp of summer, this isn鈥檛 a clip from day camp. What you鈥檙e hearing is a scientific poster session where these high schoolers 鈥 many of them freshmen 鈥 got to present their own experimental findings. These are ambitious kids, so I expected to be impressed. But what really blew me away wasn鈥檛 the amount of work they did鈥攊t was how little it resembled the kind of biology I learned in high school.
BS: So what do we have going on here? What am I looking at?
GP: For me鈥 scientifically鈥 I had to learn how to code in Python for that which was REALLY INTERESTING to say the least. But also how to analyze the data we obtained. Which wasn鈥檛 just simple graphs. We had thousands and thousands of sequences that we had to go through!
BS: That was Giovanna Prucia, a 17-year-old Junior from Connetquot High School, and the project she and Chris Paciello showed me was SUPER different from what you鈥檇 see at your average grade-school science fair.
AA: Yea! I鈥檝e heard of Python before鈥 and that鈥檚 a programming code, isn鈥檛 it? That鈥檚 not exactly something you normally learn in highschool, right?
BS: I definitely didn鈥檛. And Victoria DeAmbrosia, a science teacher from William Floyd High School, was also pretty shocked about what she鈥檚 got kids learning these days.
VD: And a lot of the students go to the next level. Those who did barcoding last year, they got to do microbiomes this year, in which they鈥檙e doing these really complex statistical analyses 鈥 both types of projects get to use bioinformatics tools鈥
AA: This is crazy! Brian, not too long ago, I too was an aspiring biologist鈥 and computer coding鈥 statistical analyses?! These were NOT the kinds of things I focused on.
BS: Well鈥 biology is changing! It鈥檚 looking that more and more scientists are going to spend as much time behind a computer as at a lab bench. And that鈥檚 for a REALLY good reason鈥
FC: 鈥淲e鈥檙e all STRUGGLING with the wonderful problem of having too much data. Big Data, as it鈥檚 featured on the cover of Nature magazine. Big Data as we talk about around the table at institute director meetings on Thursday mornings. Big Data as I am even now being asked by people in the White House 鈥榳hat are you going to do about this?鈥 as now everybody recognizes that we are in a circumstance about needing to be very thoughtful and creative about how we handle the very large quantities of biological data that are pouring out of many different approaches鈥 toward understanding how life works and how disease occurs.鈥
BS: That鈥檚 Francis Collins 鈥搕he man who has been directing the National Institutes of Health for nearly a decade. The clip I just played was from 2012, when he was serving under President Obama, but Collins鈥 goals and concerns have hardly changed since.
AA: I can understand why. Collins was also the head of the Human Genome Project 鈥 that massive scientific undertaking that, once accomplished, left the world with a WHOLE lot of data and very little idea of what it all meant. I mean, we just wrapped up a two-part series that was all about how we鈥檙e still sifting through the sequenced human genome, slowly but surely making sense of it all. One could even hazard the guess that Collins feels responsible for all the new work that needs to be done!
BS: Sure! But the other big issue is that for centuries, traditional biologists have been able to get by on their own 鈥 mostly. In this way, the cycle of observe-hypothesize-experiment-observe has been a closed system. And that鈥檚 not just for biology, but for lots of sciences! Here鈥檚 Kafui Dzirasa of the NIH鈥檚 BRAIN Initiative really bringing that point home while chatting with Collins at a meeting called 鈥淔aster Cures鈥 held last year.
K: 鈥淏ig data has gotten bigger right? So 500 years ago big data was staring out at the galaxy and mapping out the planets and how they were orbiting around the Sun. So then there was a role of an individual investigator sitting and observing and framing things. The problems today are SO complex that one person CAN鈥橳 handle all of that at all!鈥
AA: Ok. I can see that. For last year鈥檚 season finale of Base Pairs, we talked about how the problem of mapping the brain is so complex that it can鈥檛 be done by hand 鈥
BS: Or by eye, er鈥 microscopy, so to speak
AA: Right, so instead Neuroscientist Tony Zador is recruiting RNA sequencing to map the brain computationally, and THEN neuroscientists can pick specific neurons and circuits to investigate more traditionally.
BS: It鈥檚 an elegant solution. But what they have yet to really work out is how to identify which neurons are significant for any one problem. Likewise in genomics, biologists have countless genes to choose from when investigating biological function or disease
AA: 鈥 Thanks to the Human Genome project 鈥
BS: but they struggle to select key targets for study鈥 And why is this? CSHL Associate Professor Mickey Atwal suggest that it may just have to do with the fact that people are really bad at making predictions.
MA: In a way I feel like we鈥檙e almost hard-wired to get things wrong. There鈥檚 one example I can remember. Where, I went to a roulette table. And the most common thing to do at a roulette table is to bet either red or black and it鈥檚 roughly 50% (except for the one or 2 green ones)鈥
And these roulette table managers are smart. So what they figured out is if you show the customers what other players have bet in the past, that鈥檚 gonna offset how people think about random. Meaning, that if a person sees that there were 6 reds played in the past, there is a compulsion, within them, to bet that the next one is gonna be black!
And you know, there is almost a primitive part of me that felt that urge! 鈥渙f course it鈥檚 gonna be black, there have been 6 reds in a row!鈥 but if you鈥檙e grounded in understanding probability theory you understand that it makes no difference. Each one is an individual instance! And yet, I saw this time and time again. So, we鈥檙e really bad at understanding whether something is a statistical fluke or a real signal.
BS: What鈥檚 wild is that even in this roulette example, the data we鈥檙e dealing with is very small. Just 36 numbers, 3 colors, and some results. And YET, even Mickey 鈥 who is a trained physicist and quantitative biologist 鈥 even he feels the urge to bet irrationally 鈥 to feel that those results actually influence the next outcome 鈥 even when he KNOWS that in reality, they are what academics call statistical noise. It鈥檚 no wonder we can鈥檛 make heads-or-tails of Big Data!
AA: But Brian. Isn鈥檛 that what quantitative biologists DO? Help biologists filter out that noise from big data sets?
BS: In part. Yeah.
(interview clip) BS: So what IS quantitative biology?
MA: Yeah鈥. Really good question (laughter) so I think there are as many answers to that question as there are quantitative biologists. It鈥檚 not really answered to anyone鈥檚 satisfaction. And I think there鈥檚 a really good reason for that!
In most areas of science鈥攃hemistry, geography) you are defined by your object of study. And this is especially true in biology. So neuroscience is defined by the neural system. And plant biology is studying plants, right? Quantitative biology isn鈥檛. The role of a quantitative biologist is to ask certain kind of questions and to ask for certain kinds of solutions鈥 so what that means is that our domain of study can cut across many fields of biology.
And the kind of questions we ask鈥 are different form the usual question a biologist would ask. Can we simplify what we see into something abstract so that we can make a predictive model of that? Can we make a model of the phenomenon we observe? And can we test those predictions? And to do this you really have to formulate the problem differently than how a traditional biologist would.
(interview clip) BS: would this fall in the lines of say, Punnet squares?
BS: You might remember these little charts from high school biology called Punnet Squares, and even today, students use them to predict the outcome of a basic genetic cross.
MA: So that鈥檚 a really good example! I think that鈥檚 arguably the first time most biologists experience an equation鈥 And that鈥檚 a really simple example because it gives you a prediction of what is the expected observation given a set of hypotheses.
BS: You see, Andrea. Quantitative biologists are often the reinforcements that traditional biologists need in this age of big data. They鈥檙e essentially that outside help that Dr. Dzirasa was talking about in his discussion with Director Collins.
AA: So, they provide a new perspective, allowing predictions and observations to be made on a concrete statistical level. That way biologists can then formulate new hypotheses and experiments based on what is learned.
BS: Right! Right now, Mickey is working in collaboration with a number of specialists in trying to better understand breast cancer, and his lab here at CSHL is bringing that essential QB perspective to the table.
MA: Now, immuno-therapy is a buzzword and you may have read about this in popular press, and there鈥檚 been some really exciting developments in the treatment of lung cancer and skin cancer, melanoma, but it hasn鈥檛 fared so well in breast cancer. So one of the things that keeps me up at nighttime is trying to understand why not? Why aren鈥檛 the immune cells, which we know are found in breast cancers, why aren鈥檛 they doing their job and killing the cancer cells? What is it about the cancer cells that somehow tricks the immune cells into not attacking them?
So, the research team that we鈥檝e built and grouped together is really focused on understanding the communication between the different kinds of cells that you find in a growing tumor. So we have actual biopsies from patients in a clinic based in Los Angeles, actually shipped here to Cold Spring Harbor. And with our DNA sequencing facilities, we are able to measure the activity of thousands of genes in individual cells.
AA: Oh my. THAT鈥橲 a lot of data. And how each cell expresses those genes can differ wildly. One could even think of the environment around a tumor as a neighborhood. You鈥檝e got your behaving cells expressing their genes in one way 鈥 they鈥檙e good citizens. And then there鈥檚 cancer cells acting badly. But there鈥檚 also lots of other cell 鈥減ersonalities,鈥 if you will, who also might act strangely.
BS: That chatter alone creates a lot of statistical noise
AA: Right, and when everyone is talking with everyone鈥
MA: it鈥檚 a bit like a needle-in-a-haystack problem. So we have to develop algorithms that can sift through mountains of data and try to find out which genes are really important, and more importantly for this project, which genes are really important for the cells to bypass the immune system and actually allow the cancer cells to grow without the immune system killing them off.
BS: Essentially, breast cancer cells are really good at conning their neighborhood. Those 鈥済ood-citizen鈥 cells Andrea mentioned can鈥檛 tell that their nasty neighbors are ruining the neighborhood and are happy to communicate with them. And because the cancer cells are acting so darn 鈥渘eighborly,鈥 the immune system 鈥 or the local police in this metaphor 鈥 don鈥檛 realize that they鈥檙e criminals.
AA: But if Mickey and his collaborators are successful, the hope is that they can identify ways to quiet those problematic cell-to-cell conversations, putting a stop to cancer鈥檚 neighborly act.
BS: Mickey鈥檚 project is one of MANY so-called 鈥淏ig Science鈥 collabs 鈥 this is one funded by the group 鈥淪tand Up to Cancer鈥 鈥 and it shows the power of various scientific disciplines all aiming their efforts at one objective. However, Mickey argues that for these projects to truly move science forward in this age of Big Data, everyone needs to become a little more familiar with QB.
MA: You certainly don鈥檛 want to be in a position where you鈥檙e shuttling off your data to somebody else and you鈥檙e treating them like a black box. They somehow perform their magic. And they say 鈥渢hese are probably the targets for your disease.鈥 Right? You want to have some sort of conversation. You want to be more intelligent than that.
Even if you鈥檙e not going to do the experiments themselves, I still think it鈥檚 really important for the experimentalists and the classical biologist to at least be able to understand what are the state-of-the-art techniques that are required to sift through mountains of statistical data.
I really do think that it鈥檚 going to be the next generation of biologists 鈥 undergraduates, graduates, and postdocs 鈥 all need to be trained in computational and quantitative skills.
AA: Hmm, well鈥 there is good news. Here at 黑料吃瓜资源, the students of our Watson School of Biological Sciences
BS: 鈥 that鈥檚 our Ph.D. Program 鈥
AA: every student admitted to the program is required to take what can be best described as a computational 鈥渂oot camp鈥 where they learn PYTHON 鈥 that premiere programming language for MANY important scientific databases.
BS: And amazingly, it鈥檚 mostly being taught to students who have only ever known text books and a lab bench.
MA: We say 鈥渉ey! This a computer. This is what a computer does. This is what it doesn鈥檛 do! This is how we can write code to perform basic commands. And by the second day they鈥檙e actually analyzing next generation sequencing data by themselves!
AA: Mickey teaches this boot camp and as you might expect, he (and his subject) are not exactly popular with new students.
BS: Can you blame them? They came here to do science鈥 and instead Mickey鈥檚 got them sitting behind a computer! Doing code!
AA: And yet鈥 we know that this is exactly how lots of science gets done.
MA: I鈥檓 sure there鈥檚 a whole bunch of them who hate me here when they first arrive鈥 (laughter)
MA: And what鈥檚 interesting is that, because it鈥檚 such a new skillset, such a new concept, and it鈥檚 so immersive on the first day that by the end of the first day they usually end up鈥 dreaming鈥 well, thinking about Python obsessively. And you can get to this state where you can just spend hours in front of a computer coding away! It can become quite addictive.
BS: According to Mickey, it鈥檚 rare to have a true convert 鈥 a student who actually leaves their well-paved bio-science path to wade into the unknown of quantitative biology. However, he did explain that by the end, the idea that you can frame a theoretical scientific question mathematically becomes pretty popular among the students.
AA: So popular in fact, that for three years now, Mickey has been elected to receive the Winnship Herr Teaching Award by the Watson School鈥檚 freshman class.
BS: It鈥檚 basically a teacher popularity contest, and each year, Mickey acts like he has no idea why he鈥檚 won it.
(graduation ceremony clip) MA: I don鈥檛 know how and who decides these awards uhh鈥 but whoever you are鈥 Russian hackers included鈥 your check is in the mail (laughter).
AA: (laughing) I spoke with Mickey about this for a story on our LabDish blog, but it鈥檚 easy to see why his course is actually popular. While other instructors are simply reinforcing old skillsets and scientific methods, Mickey is teaching these students something fresh! It鈥檚 practically a new way to think for many of these young scientists.
BS: And that鈥檚 what I actually found most surprising during my chat with Mickey. While this strategy 鈥 this way of approaching theoretical problems seems rather new to many scientists 鈥 it鈥檚 actually been around for decades.
MA: What鈥檚 not appreciated enough in biology I think is the Watson and Crick paper, their famous paper, THAT鈥橲 A THEORY PAPER! There鈥檚 no new data that鈥檚 reported. It鈥檚 basically a theoretical conjecture on solving some equations based off crystallography.
BS: And yet, in that paper 鈥 in the 1953 Nature paper in which Watson and Crick proposed that the DNA molecule was shaped like a double-helix 鈥 there is a single line that many scientists can recite.
AA: 鈥淚t has not escaped our notice,鈥 it begins, 鈥渢hat the specific pairing we have postulated immediately suggests a possible copying mechanism for the genetic material.鈥
BS: This bon mot is famous probably as much for its coyness and understatement as for the significance of its prediction鈥 and that鈥檚 a good embodiment of what quantitative biology really is. It鈥檚 not a field of a study, but a strategy and even a way of thinking (pause) with predictive prowess that are too-often under appreciated.
AA: Mickey is but one of our scientists employing quantitative biology in a quest to answer some really important questions, so like always, this won鈥檛 be the last you hear of this subject.
BS: Cancer, autism, neuroscience, the evolution of humanity, and SO much more 鈥 all are subjects that employ QB, and all are things we have or will talk about in episodes of Base Pairs.
AA: So stay with us! And as always鈥 鈥渕ore science stories soon!鈥
