Wednesday, August 11, 2021

The Perturbome

When focusing on cell types, we could make a list of the most abundant transcripts in particular cell types. We could also focus on the proteome. We could even ask, “what are the common transcripts/proteins that are rarely seen in a particular cell type?” We could search for cell type markers that are rarely or never seen in other cell types, even if these markers are not particularly abundant in the cell type of interest. Our database is chock-full of the above sorts of lists.

There’s another sort of list we are able to prepare, largely because the sheer size of our database affords the opportunity. The largest portion of the database falls into the category of “perturbation studies” wherein cells are perturbed via drug, knockout, heat, whatever. We can thus ask the question, “what transcripts/proteins are most commonly perturbed in a particular cell type?” We can also ask which entities are least frequently perturbed in a particular cell type. This is not a question of abundance or of “uniqueness” to a particular cell type. Rather, we’re focusing on the entities that fluctuate when you tweak a particular sort of cell.

Pulling data from about 10,000 studies, we’ve constructed these lists for 20 different cell types: brain, liver, skin, muscle, lymphocytes, stem, kidney, breast, colon, prostate, heart, lung, intestines, glands, pancreas, dendritic, ovaries, adipose, fibroblasts, epithelial. Why not other cell types in our database? That’s primarily because the above 20 designations are the most common in our database; we required at least 100 studies for each cell type. We could have also included “blood” as a category, but we chose to break it down further into two common subtypes: lymphocytes and dendritic cells. Some other choices were somewhat arbitrary (we have a lot of macrophage studies…why didn’t we include them?) Note also that some of these cell types can overlap….skin and liver are different tissues, but skin can contain stem cells, and breast cells can be epithelial. For this initial stab at the “perturbome”, this isn’t a problem.

With the 20 cell types, we generated 40 lists, as each cell contains entities that are frequently perturbed, as well as entities that are rarely/never perturbed; two lists per cell type. Entities that are “rarely perturbed” most likely are simply never expressed in the particular cell type, though it is possible that they are indeed expressed, but it’s difficult to tweak them; we don’t discriminate between these two cases.

If you’re interested, the dirty details are as follows: We first generated a list of all genes found in the above studies. We then simply counted their occurrences in the above 20 cell types. We then used the binomial distribution to calculate how significantly a particular gene may be over/under-represented in a particular cell type. The “probability” input for the binomial distribution (which is .5 if you’re talking about coin flips) is calculated by dividing the total genes perturbed in a tissue (e.g. brain) by the total genes perturbed over all 20 tissues. Liver, for example, constitutes 7% (.07) of all genes in our database’s perturbome. Thus, if you know that gene ABC is found 100 times in our liver studies, and 500 times over all studies, you’re equipped to perform a probability calculation. In the final step, we simply rank genes according to these probabilities, making sure to discriminate between significance generated from an excess, versus depletion, of a particular gene.

So…what did we find? First, what is the most commonly perturbed gene over all tissue types? The answer is EGR1, perturbed 1581 times, followed by SERPINA3, IFIT1, GDF15, and FOS. What is the most common gene which was never perturbed in a particular tissue? That distinction belongs to LUM (lumican), which was never perturbed in dendritic cells, despite being altered a total of 758 times over the other 19 tissues. Perhaps dendritic cells are adamant that they not be confused with other cell types that express lumican, which is largely an extracellular protein.

Of note, GAPDH, commonly used as a housekeeping control gene, was only perturbed 358 times. Actin-B was seen 610 times. Our gene lists primarily reflect perturbations, not abundance.

Below is one of the more obscure gene tables you’ll ever stumble across. It details the top genes that were never perturbed in particular cell types. The table is ordered by the count over all other tissues; thus, the cell types at the bottom of the table express a large array of transcripts/proteins.

CELL TYPE

GENE

COUNT OVER OTHER TISSUES

dendritic

lum

758

adipose

hopx

603

intestines

nav2

591

ovary

lam1

481

heart

jdp1

469

prostate

pck1

456

gland

hpp1

435

pancreas

cd38

422

muscle

slc6a14

405

colon

kcnk2

339

fibroblast

c1orf116

329

kidney

scca1

312

skin

sizn

230

brain

ugt2b15

205

lung

gpr37l1

183

breast

miat

175

stem

cyp2c9

164

liver

blcap

149

epithelial

ces1g

121

 

Another question: what are the genes that were uniquely perturbed in particular tissues? The champion is probably GM1818, a mouse gene that was perturbed 21 times in the brain, but never elsewhere. For human genes, we have FAM90A7P, which was perturbed 16 times in the brain, and never elsewhere. The brain, in fact, seems to have the largest number of uniquely perturbed genes by a large margin; the first case of a non-brain gene that was uniquely perturbed was the mouse gene AI132709 (liver), which was tweaked in 8 studies in our database…201 brain-unique genes are tweaked at least as frequently. The lncRNA Lnc-CHSY1-3 was uniquely expressed in lymphocytes, albeit with a mere 4 occurrences.

The special status of the brain is also seen in the heatmap below. We took our 40 perturbation lists and performed Fisher’s exact test on all combinations of lists, for a total of 760 P-values. 


If the image is too small, you could click on it to get a bigger view. The first row is labeled “LO_COL”, which means “genes that were least frequently perturbed in the colon.” Hopefully the other 39 labels are self-explanatory. The color key shows the –log(P-values). Combinations with very significant P-values tend to make sense…highly perturbed genes in ”breast” and “gland” overlap with extreme significance, as do non-perturbed genes in the lymphocyte/dendritic categories, and perturbed genes in the colon/intestine. There are, however, some very interesting overlaps that might not be so intuitively obvious. For example:

1) genes that are rarely perturbed in the brain are rarely perturbed in stem cells.

2) genes highly perturbed in glands are rarely perturbed in the brain.

3) genes highly tweaked in the breast are rarely tweaked in stem cells.

4) genes that are rarely perturbed in the brain are also rarely perturbed in muscle and lymphocytes.

5) looking at the “high_BR” (highly perturbed in the brain) group, the best matching “highly perturbed” cell type would be “stem”, with a –log(P-value) of about 4. This is a bit of a cheat, since stem cells and brain cells are not exclusive (i.e. some brain cells are stem cells). In truth, then, highly perturbed genes in brain cells do not overlap with the highly perturbed genes of “pure” cell types with any significance.

6) unlike the brain, the rarely perturbed genes in some tissues don’t overlap rarely perturbed genes in other tissues with great significance. For example, the rarely perturbed genes in the pancreas don’t overlap with rarely perturbed genes in other tissues with any amazing significance; the best match, in fact, would be to intestines, with a -log(P) of 7.

You can tinker with the data yourself at whatismygene.com. The table below gives you the dbase IDs that allow you to perform operations with our various apps.

DBASE ID

CELLS

132346123

most frequently perturbed in the brain

132346124

least frequently perturbed in brain

132346125

most frequently perturbed in the liver

132346126

least frequently perturbed in the liver

132346127

most frequently perturbed in skin

132346128

least frequently perturbed in skin

132346129

most frequently perturbed in muscle

132346130

least frequently perturbed in muscle

132346131

most frequently perturbed in lymphocytes

132346132

least frequently perturbed in lymphocytes

132346133

most frequently perturbed in stem cells

132346134

least frequently perturbed in stem cells

132346135

most frequently perturbed in the kidney

132346136

least frequently perturbed in the kidney

132346137

most frequently perturbed in the breast

132346138

least frequently perturbed in the breast

132346139

most frequently perturbed in the colon

132346140

least frequently perturbed in the colon

132346141

most frequently perturbed in the prostate

132346142

least frequently perturbed in the prostate

132346143

most frequently perturbed in the heart

132346144

least frequently perturbed in the heart

132346145

most frequently perturbed in the lung

132346146

least frequently perturbed in the lung

132346147

most frequently perturbed in the intestines

132346148

least frequently perturbed in the intestines

132346149

most frequently perturbed in glands

132346150

least frequently perturbed in glands

132346151

most frequently perturbed in the pancreas

132346152

least frequently perturbed in the pancreas

132346153

most frequently perturbed in dendritic cells

132346154

least frequently perturbed in dendritic cells

132346155

most frequently perturbed in ovaries

132346156

least frequently perturbed in ovaries

132346157

most frequently perturbed in adipose tissue

132346158

least frequently perturbed in adipose tissue

132346159

most frequently perturbed in fibroblasts

132346160

least frequently perturbed in fibroblasts

132346161

most frequently perturbed in epithelial cells

132346162

least frequently perturbed in epithelial cells

 

We’re not finished with our dissection of the perturbome. We’ll resume the discussion in a couple weeks.


whatismygene.com 

Thursday, July 15, 2021

On Cutoffs

Whenever an individual with some background in bioinformatics questions me on my approach to gene-enrichment, the most common question is probably, “what are your cutoffs?” Specifically, when entering lists of transcripts/protein/micro-RNAs (whatever…call them “genes”) into the database, what is the lowest level of significance required for a gene to make the cut? Or, perhaps, what is weakest fold-change? My answer: we don’t have strict cutoffs. Most typically, data is sorted according to some criteria (e.g. significance), and the top 200 up- and down-regulated portions are both entered into the database.

Some folks are not appreciative of this approach, preferring, for example, that only data with an FDR-adjusted P-value < .05 be entered. However, this would mean that many interesting datasets would be excluded. If FDR < .05 is strictly employed, all lists such as “tyrosine kinases”, or “common contaminants in mass spectrometry” would be eliminated. Journals sometimes offer ranked (1,2,3) gene lists without fold-change or significance measures.

Most importantly, even in studies where significance and/or fold-change can be measured, strict cutoffs can cause the loss of very interesting results. Below, I invoke real studies to convince skeptics that data that is “technically” insignificant can be very interesting. There are two forms of evidence. In the first group, a gene is targeted (e.g. by knockdown), yet fails to meet standard P-value cutoffs. Nevertheless, this gene is the single most strongly altered entity in the study when measured by mere fold-change. In the second group, a collection of insignificantly altered genes under a particular experimental condition matches up with extreme significance to another gene set in our database under similar experimental condition. For example, study A might treat cells with a drug vs. control, with genes being altered at mathematically insignificant levels. Study B, which is independent of A, uses a similar drug. We then find that despite A’s lack of significance, studies A and B overlap with extreme significance (as measured by Fisher’s exact test).

I’m not a mathematician, and won’t delve deeply into the apparent over-conservatism of standard adjustment methods (particularly Benjamini-Hochberg). Here’s one argument, however, that is fairly intuitive: Let’s imagine a study with 100 genes up-regulated at an adjusted significance of .75, with 9900 other genes being up-regulated at an adjusted significance of 1.0. The .75 figure, of course, is “insignificant.” Nevertheless, .75 also tells you that 25 of those 100 genes may have “really” been upregulated. You construct a gene list using those 100 genes. When applying Fisher’s exact test against another list of data, 25 “truly” upregulated genes, versus 1 or 2 that you might expect by randomly pulling 100 genes from the 10,000, can result in huge alterations in significance. Basically, a list of weakly altered genes can match up extremely significantly with other lists; mass effects in action.

Hopefully, the sum of examples below will be convincing. The examples are far from exhaustive...I merely collected them over the last few weeks upon off-handedly noticing (for the millionth time) that these insignificant "omics" results were actually very interesting.

Group 1

*In a GEO Dataset (GSE81399) from the study, Targeting of Mesenchymal Stromal Cells by Cre-Recombinase Transgenes Commonly Used to Target Osteoblast Lineage Cells, DMP1 should be overexpressed in “targeted” cells. It is, and in fact has the greatest fold change (about 10X) of 22,000 transcripts. However, following FDR adjustment, this alteration is insignificant (P = .67).

*In GSE83388, the gene KSRP is knocked down. After adjustment, the P-value (vs scrambled siRNA) is .82. Nevertheless, KSRP has the second-greatest fold change of 31,000 identified transcripts.

*In AKT isoforms modulate Th1-like Treg generation and function in human autoimmune disease, IFN-G+ tregs are separated from IFN-G- tregs. IFN-G itself fails to reach significance in IFN-G+ cells, though some other transcripts are indeed significantly altered. Nevertheless, IFN-G has the single-greatest fold-change of any of 31,742 transcripts.

*In Transcriptomic Analysis Unveils Correlations between Regulative Apoptotic Caspases and Genes of Cholesterol Homeostasis in Human Brain, CASP2 is knocked down. After adjustment, this knockdown is insignificant (P = 1.0). Nevertheless, when more than 30,000 genes in the transcriptomic set are ranked according to fold-change, CASP2 ranks #1.

*In Genome-Wide Analysis Identifies NURR1-Controlled Network of New Synapse Formation and Cell Cycle Arrest in Human Neural Stem Cells, NURR1 is overexpressed. After adjustment, this overexpression is not significant against controls. Nevertheless, NURR1 is the single most upregulated transcript (of more than 30,000) in the study as measured by fold-change.

*In GSE175853, TAZ is knocked-down. Relative to controls, the adjusted P-value is .492. Nevertheless, of more than 20,000 transcripts, it ranks #2 in terms of fold-change.

Group 2

*In Work, meaning, and gene regulation: Findings from a Japanese information technology firm (PMID 27434635, GSE79092), male workers were scored according to numerous psychological parameters (e.g. hedonia) and blood transcripts were examined. We chose to look at eudemonia. Despite the fact that no single transcript was significantly altered in comparison of high vs. low eudemonia subjects, sorting according to fold-change and then searching for datasets that strongly intersected this particular study generated some very significant and interesting results. For example, transcripts downregulated in high eudemonia subjects tended to be upregulated in responders to lithium treatment (log(P) = -77) and in sleep deprivation subjects (-47). Transcripts upregulated in high eudemonia subjects tended to be downregulated in mice upon “polytrauma” (-21). It is thus difficult to argue that these “insignificantly altered” transcripts are no different than randomly sorted transcripts.

*Comparing grade II vs grade I breast lobular carcinoma (GSE88770), not a single transcript showed an adjusted P-value less than 1. Nevertheless, a list of the most downregulated transcripts in this comparison best intersected with another breast cancer study (GSE49481). Specifically, these downregulated transcripts intersected with transcripts downregulated in invasive ductal carcinoma versus invasive lobular cancer at P=10-34.

*In GSE112943, the lupus lesional skin transcriptome is examined against healthy skin. Despite no single gene rising to the level of significance after adjustment, the database dataset that best matches upregulated transcripts in this study is another study of lupus lesional skin: GSE72535. The best match to downregulated transcripts in GSE112943 is found in another study of lesional skin: GSE136757.

*In GSE79721, a novel BET inhibitor is applied to breast cancer cells. Comparing these cells to controls, not a single transcript was altered at an adjusted P-value of <1.0. Nevertheless, the single best overlapping upregulated and downregulated datasets (with Fisher P-values of 10-75 and 10-54) both involved application of BET inhibitors to cells.

*In Knockdown of a novel lincRNA AATBC suppresses proliferation and induces apoptosis in bladder cancer, bladder cancer tissue is compared to adjacent tissue. After adjustment, no transcripts are significantly altered, possibly because only two cancer and two control samples were examined. Nevertheless, the database study that best overlaps transcripts upregulated in this study (at P=10-18) is yet another comparison of bladder cancer against adjacent tissue (GSE100926). On the downregulation side, two other cancer studies (esophageal and colorectal) outcompeted GSE100926 for significance.

*In GSE58591, a comparison of female vs male ES cells is made. Following FDR adjustment, only 4 transcripts are significantly altered. Nevertheless, Y-chromosome transcripts dominate the list of most strongly downregulated transcripts, as measured by fold-change alone (e.g. DDX3Y, USP9Y, EIF1AY, etc).

 


whatismygene.com 

Sunday, June 20, 2021

A New Feature at WhatIsMyGene.com

Most of the gene lists that compose our database are sorted according to some criteria. We may apply a significance cut-off to the data, and from there sort according to fold-change. Often, we divide log(fold-change) by the significance, combining these two measures in one step. In cases where significance measures are not available, or no genes are significantly altered, we’ll sort according to fold-change alone. You may wonder why we’d burden the database with studies wherein no genes at all were significantly altered…we’ll address that shortly in another post.

Regardless of the sorting method, in some cases there may be little difference between the top “most upregulated” gene and the 100th; a gradual decline. In other cases, there may be a steep dropoff from the first position to the 10th. An extreme case would be a knockdown experiment in which the targeted transcript is very significantly downregulated, but no other transcripts are altered to any great degree.

In cases where, say, only the top 25 genes in a list of 200 are “interestingly” altered, you might consider the remaining 175 genes to be more or less random garbage. This garbage could hide a significant result. When performing Fisher’s exact test against a selection of other studies, input composed of those 25 genes might render more significant results than input composed of the larger 200 gene list.

Or, perhaps you’re using our “relevant studies” tool. You want to find studies in which your gene of interest is very strongly (vs. moderately) altered. You’d thus like to eliminate the lower-ranked genes in our lists.

Our new feature simply allows you to whittle our database down to “top 100”, “top 50”, or “top 25” genes. Very roughly, most of the underlying lists are composed of about 200 genes, so the above options allow you to eliminate the 50%, 75%, and 88% lowest ranked genes in a list. The feature can be applied when you use the “Relevant studies”, “Coregulation”, “Fisher”, “Match Studies”, “Regulation”, and “Third Study” tools. You’ll see this “Restrict IDs” option at the the bottom left of a page. Using the tool has the additional benefit of speeding up the generation of output, as you’re tossing up to 88% of the database into the garbage.

Practically speaking, if you compare results with and without the “Restrict IDs” option, the most likely outcome is a lowered significance when restricting the database size. This is because a typical gene list in our database shows a gradual, not steep, change in significance (or fold-change…whatever). Thus we’d advise that you ignore this option when looking for broad trends and insights, and use the option when you seek to refine a result. The above does not apply to the “Relevant Studies” tool, as this tool simply searches for your gene in our database, and doesn’t generate any statistics. In the case of “Relevant Studies”, you may wish to begin with a restricted database size. In the case of "Relevant Studies", we've added "Top 10" and "Top 5" options, meaning you could restrict approximately 97.5% of the database.

In the case of the “Coregulation” tool, let’s say you restrict the database using “top 25.” In this case, your gene of interest must be found in the top 25 genes in a study, and its coregulated partners must also be found in the top 25.

At the beginning of this post, we say that most of our lists are sorted. In some cases, a journal simply provides a list of genes without any fold change data. Another sort of list would be represented by “518 human kinases”…we can’t say that one kinase is better than another, so we simply randomize these identifiers in our database. Also, early in our existence, we did not sort our lists; as time passes, we replace these early lists with improved lists that are sorted, but this is a slow process. When you use the "Restrict IDs" feature, in addition to tossing low-ranking genes, you're also tossing randomized (i.e. non-sorted) lists. Thus, if you wish to examine the maximum number of studies in our database, do not use this feature.

***

Given that most of our data is sorted, we thought it might be interesting to find the genes that most frequently occupy the #1 slot. The champion might be a bit unexpected: SPP1, appearing atop our lists in 30 different studies! Following that, we have TTR, EGR1, S100A8, FOS, LYZ, HBA1, and CCL5. If we adjust for rarity of the gene throughout the database, the list looks like this: EGR1, SPP1, TTR, ABL1, RAP1A, HBA1, and S100A8. Genes that are relatively rare over the entire database, yet nevertheless found themselves at the top position on multiple occasions include UBR4, FOXP3, ESRG, COX1, SNX3, EIF4A3, A1BG, and IGHD.

whatismygene.com 

Wednesday, May 19, 2021

Drug Resistance and Transcriptomics

Our database lists more than 100 studies in which drug resistant cells were compared against drug sensitive cells. Most commonly, a sensitive cell line is passaged in the presence of low drug levels until a resistant strain emerges, whereupon a transcriptomic comparison can be made. In other cases, tumor cells from resistant patients may be compared with cells from sensitive patients.

Given this plethora, we decided to gather all these studies and see if any particular genes emerged that were commonly upregulated or downregulated in the case of drug resistance. The result is a bit more complex than we had hoped. We looked at 95 studies, excluding those involving “radioresistance.” The gene that was most commonly altered in these studies was OAS1, appearing 21 times out of 190 opportunities (all 95 studies have up and down-regulated portions). That seems nice…a “big name” gene popping up at a frequency that, without crunching the numbers, appears to be significant. The problem is this: OAS1 appeared in both the resistance-upregulated and resistance-downregulated datasets (13 times up and 8 times down, to be specific). Therefore, we can’t make the blanket statement that OAS1 is upregulated in cases of drug resistance, nor can we surmise that suppressing the innate immune response (OAS1 is a big player there, after all) might overcome drug resistance. OAS1 is not unique in this respect.

That doesn’t mean that generation of lists of commonly up- and down-regulated genes involved in drug resistance would be entirely fruitless and couldn’t possibly spur insight. We’ve given these two lists the database IDs 129091122 and 129092122. In both cases, a gene had to occur at least 7 times (out of 95) to make the list, giving the lists a composition of 163 and 119 genes. 26 genes were found in both lists:

AREG

IFI27

C1orf24

IL1A

CA12

KYNU

CD24

LCN2

CEACAM6

OAS1

CXCR4

PEG10

DUSP6

SERPINB2

EMP1

SERPINE2

FAM129A

SOCS2

FSTL1

STC2

GPNMB

TSPAN8

HLA-DRA

UCHL1

HLA-DRB1

VCAN

Ignoring the fact that multiple genes are found in both lists, what broad categories of genes intersect with these lists? In the case of upregulation, the innate immune response does indeed seem relevant (e.g. genes upregulated early (vs late) in HCMV infection intersect with a P-value of 10-76). Erlotinib and neo-adjuvant therapy seem to do a fine job of upregulating common drug-resistance genes; if borne out in the lab/clinic, this would have obvious implications for cancer cocktail approaches. On the downregulation side, the metastatic (or not) nature of underlying tissues seems to be relevant. Specifically, transcripts downregulated in aggressively metastatic tissue overlap with transcripts that are downregulated in the case of drug resistance. For far deeper details, just plug either of the above dbase IDs into our “Fisher” app.

How about creating lists of, say, upregulated genes that weren’t found at all in the downregulated category? We tried that, but were met with discouragement. We found that RAB25 was the single best example of a transcript that was downregulated (in 8 studies), but never upregulated. Searching for validation of this characteristic of RAB25 in specific studies, the first study we stumbled upon was this: RAB25 confers resistance to chemotherapy by altering mitochondrial apoptosis signaling in ovarian cancer cells. There, it seems, RAB25 upregulation, not downregulation, correlates with resistance. Hmmmm.

Apparently, the same transcript may be upregulated in one resistance study, and downregulated in the next. My background in this field (a couple distant lectures and/or presentations) informed me that a handful of transporters are the main culprits in drug resistance, and I had hoped that this phenomenon would be obvious once multiple studies were compounded. This is not the case.

One might think that the above conundrum be resolved by examining particular drugs. That is, certain critical transcripts would always be upregulated in resistance to a particular drug. Cisplatin-resistance is the most commonly studied sort of resistance. Here, of 5 genes found in 4 out of 6 of the cisplatin-resistance studies we have on hand, 3 are both upregulated and downregulated, depending on the study: QPCT, SAA1, and MMP1 (TGFB2 is up in all four cases, and ANO1 is always down).

To attempt to clarify matters, we clustered the 190 datasets. Specifically, we performed Fisher’s exact test for each dataset against our entire database, generating millions of P-values. These P-values were the raw data for clustering (Cluster 3.0, k-means). 10 clusters were generated. The top genes found in each cluster are now found in our database with IDs 129137122, 129138122, 129139122, 129140122, 129141122, 129142122, 129143122, 129144122, 129145122, and 129146122.

Though the clusters did not nicely segregate according to drugs or cell types, as one might desire, some clarity was gained. Bearing in mind that both up- and down-regulated transcripts can be found in a single cluster, Cluster 0 transcripts tend to be upregulated on innate immune stimulation (e.g. via interferons). Cluster 3 transcripts tend to be upregulated in the case of metastasis and are enriched for cell-surface markers, while cluster 5 and 8 genes tend to be downregulated in metastatic cells. Cluster 6 genes have a strong tendency to be upregulated in cancer versus adjacent tissue and, rather bizarrely, downregulated on resveratrol treatment (P = 10-56). Other clusters are more nuanced. Thus, it would appear that investigators might wish to place special relevance on the status of cells with regard to innate immunity and metastasis when considering approaches that might mitigate drug resistance. Depending on the cluster, we see hints that particular drugs could, to some extent, reverse drug resistance: bromodomain inhibitors, noggin, losartan, gefitinib, etc. Other drugs, of course, could enhance drug resistance.

It is sometimes difficult to trust the output of clustering programs, so I performed an eyeball version of clustering in Excel. Give different colors to different significance levels (below, green indicates P<10-15), and then sort a column. Gather all columns where colors (indicating significance) percolated to the top. That’ll be cluster 1.Then move on to a column that doesn’t fall into cluster 1. Repeat. Believe it or not, this crude method matched up quite nicely with the software I used. To me, this sort of correspondence between the mathematical perfection of the clustering software and the childish simplicity of matching columns that have the same colors indicates that maybe we shouldn’t spend an excess amount of time/energy debating the merits of, say, “Euclidean distance” vs. “City block distance.” Below is a sliver of the result:


Again, for the fine details, just visit WhatIsMyGene, plug in database IDs (or your own datasets), and have fun.


Note 2/22/2022: We've added quite a few more studies involving resistance to our database. At this point, it's fairly obvious that genes upregulated in cells resistant to one sort of treatment may actually be downregulated on resistance to another treatment. This observation dampens hopes for across-the-board approaches to drug resistance. On the positive side, it may mean that resistance could be dealt with via drug cocktails; i.e. two drugs that trigger opposing resistance patterns could be combined in a treatment. At some point in the future, we'll re-cluster our resistance results. We'll be a bit more rigorous about finding an optimal number of clusters, look a bit deeper into commonalities in these clusters, and perhaps examine cases where 2 "resistance-complementary" drugs might be applied to particular maladies.

Note 9/17/2022: A quote from Comparative proteomic analysis identifies key metabolic regulators of gemcitabine resistance in pancreatic cancer: Surprisingly, a number of proteins that were downregulated in MIA-GR8 cells have been reported to promote drug resistance in other cancer types. It's nice to validate our view above, but it's also disappointing to see that many researchers may still be stuck in a one-dimensional view of drug resistance.


whatismygene.com 

Saturday, May 15, 2021

Did Covid-19 Emerge from the Wuhan Institute of Virology?

I’m going to address this topic with a minimum of drama. Go away if you’re a conspiracy buff. Stick around if you’re interested in a frank, somewhat introspective take on this question from a dude (albeit a low-impact dude) who has actually tinkered with viruses. For "safety", I'll spell out my #1 point right here: scientists have knee-jerk responses too. For even more safety, let me also spell out the following at the start: I still find it unlikely that the virus emerged from the Wuhan Institute of Virology.

There are two widely disseminated documents from credentialed authors providing arguments against and for the notion that Covid-19 was lab-generated. On the “against” side, we have a Nature article from March of 2020. On the “for” side, we have Nicholas Wade’s take.

I recall reading the Nature article last year, shaking my head at some of the refutations within, and then moving on to other topics. Wade’s article reminded me of my early skepticism. I’ll affirm two of Wade’s points:

1) The Nature article argues that the absence of a “previously used virus backbone”* within Covid-19 provides evidence that there was no lab-manipulation. Let me say: this is malarkey and, at best, an embarrassment for Nature. I’ve generated “backbone” free viruses myself (on dengue, to be specific). You insert the viral sequence into a plasmid, perform in vitro transcription, and infect cells with the resulting RNA. If you designed the plasmid correctly, there should be no evidence of “backbone.” Even if you erred, the virus may quickly shirk garbagy, non-optimal sequences upon multiple passaging (it can be frustrating to insert “loss of function” mutations into a virus, as the virus might dispense with them surprisingly quickly, if they don’t kill the virus from the very beginning).

In case anyone wishes to nitpick: yes, the 30kb length of coronaviruses makes ordinary plasmid insertion tricky, if not impossible. But there are plenty of methods to generate these long viruses without evidence of a backbone.

It’s hard to believe that the esteemed authors of the Nature article weren’t aware of these viral basics. Why did they choose to offer this lame argument?

2) The argument is made that the spike protein’s interaction with the ACE2 receptor is not optimal; therefore, Covid-19 could not be the product of manipulation.

Again, this is absurd. You have to assume that any and all lab experiments involving Covid-19 would involve insertion of the theoretically optimal (for ACE2 binding) spike protein sequence. Here’s an example of an experiment that I would consider interesting: perform some sort of guided evolution to generate a myriad of spike protein sequences, and test them ALL for both ACE2 affinity and infectivity**. Take the “winners” of this process, insert them into the virus, and write a paper. That’s just one of a near infinite number of experiments you could perform.

Let’s imagine that Dr. Evil is indeed behind the Covid-19 pandemic. He, like any competent virologist, would not automatically assume that the virus that best binds ACE2 has the highest potential to wipe out the human race. It wouldn’t surprise me at all to find that such a virus would be severely handicapped, refusing to let go of ACE2 at any step, and unable to perform its various pleiotropic functions.

Again, it’s odd that virologists would even attempt to pass this argument off in a Nature article.

There’s further lameness in the Nature article. For example: some of the mutations in Covid-19 haven’t been mentioned in the literature as yet. The idea, I guess, is that any lab-generated mutations would already have been described. I won’t even bother refuting that.

I have to question at least one of Wade’s other arguments, however. This regards the appearance of a furin cleavage site within the virus. This is supposed to be some sort of smoking gun for lab experimentation. The site is only 4 amino acids long. It’s not easy to estimate the probability that nature would come up with this mutation. Bear in mind that coronaviruses are the absolute champions of a process called “RNA recombination.” Without going into detail, the furin site doesn’t have to emerge via a step-wise series of mutations…it could enter in one fell swoop. Again, if there’s any “garbage” RNA left over from recombination, it could be eliminated quickly via evolution, including further recombination. If a paper attempts to address the furin cleavage site appearance from a probabilistic perspective, be skeptical about the underlying assumptions about what viruses do and don’t do.

On the other hand, everybody in the virology world inserts furin cleavage sites in their viruses and “replicons.” It’s something we do.

So, where do I stand? The most dramatic thing I can say without feeling guilty is this: we’re far from eliminating the possibility of a lab-generated Covid-19. Nothing I’ve seen convinces me that the virus couldn’t have emerged from the lab in Wuhan. Certainly not the Nature commentary.

I’ll take Wade at his word when he says that the Wuhan lab is China’s #1 coronavirus research facility. Rather odd, no? The counter-argument, I guess, might be that Wuhan is an optimal location to study coronaviruses, because that region of China is coronavirus heaven. I don't know.

To be clear, there’s a huge difference between a lab accident and intentional release. I don’t see any reason to assume the latter. How has China emerged from this mess? With an economy that’s not any stronger than anyone else’s, and the clunkiest vaccines on the market. Infections have been minimized in China, but the emergence of variants threatens that. India surprised everyone with a minimum of infections and deaths…last year.

Returning to the question of why Nature published its lame refutation, let me offer a bit of introspection. I don’t want a lab accident to be the cause of Covid-19 and I feel compelled to argue against the possibility. Just as a big-time developer tires of apparently nit-picky regulations covering endangered insect species, virologists don’t want further restrictions on their activities. We feel like we know what we’re doing. I suspect that the authors of the Nature article feel the same.

Finally, if you point a gun at my head and inquire as to the most probable source of Covid-19, I'd have to lean strongly on the side of natural origin. If you've read the above and have concluded I'd think otherwise, sorry to disappoint. There are plenty of arguments to support the natural origin of Covid-19; most of them, unfortunately, are not very accessible to layfolk. Here's the one that I find most difficult to refute: the 97% similarity between Covid-19 and its closest relative, RatG13, means that Covid-19 diverged from RatG13 no later than the early 1980's, and probably earlier. Thus, Dr. Evil (or Dr. Carelessness) would need to introduce about 1,000 mutations into RatG13 over the years. Whether by site-directed mutagenesis, lab passaging, or directed evolution, that's a figure that nearly unimaginable to virologists. Note also that these sequence differences are spread all over the viral genome; there's no sign that, for example, the spike protein was singled out for special treatment.

Given the above, if you're dead-set on blaming the Wuhan Institute, the only remotely possible scenario that I see would be the following: WIV scientists gathered the Covid-19 virus, or something very closely related, and brought it into the lab, whereupon it escaped with few or no mutations. Given that folks have not identified any virus with higher Covid-19 similarity than RatG13, one might surmise that such a virus may have been collected outside of China. There are indeed studies wherein WIV scientists gathered viruses outside of China (e.g. Africa). Now, if Covid-19 has a natural origin in China, what's more likely: it spread in the chaotic environment of a wet-market, or it spread in the controlled environment of a virology institute? In the case of import from outside of China, one could accuse the WIV of carelessly handling a virus to which the local population may have little immunity. All very speculative, with no evidence at all at this point.

*I note that some metavirology folks use the term "backbone" to refer to the conserved portions of a viral sequence. However, that's not the case in the Nature paper, which points to a paper on Coronavirus construction methods, not broad sequence comparisons, following the term "backbone".

**In fact, a bit of Googling shows that the folks at Wuhan are familiar with Selex, a method that lets evolution, as opposed to "rational design", determine an experimental outcome. Check out, for example, A SELEX-Screened Aptamer of Human Hepatitis B Virus RNA Encapsidation Signal Suppresses Viral Replication. To be clear, this particular paper optimizes RNA, not a protein.


whatismygene.com 


Mixing WIMG and AI

We've been tinkering with incorporating AI (Gemini) with our Fisher app. The idea is not so tricky...you feed the AI an identity/ruleboo...