Wednesday, May 19, 2021

Drug Resistance and Transcriptomics

Our database lists more than 100 studies in which drug resistant cells were compared against drug sensitive cells. Most commonly, a sensitive cell line is passaged in the presence of low drug levels until a resistant strain emerges, whereupon a transcriptomic comparison can be made. In other cases, tumor cells from resistant patients may be compared with cells from sensitive patients.

Given this plethora, we decided to gather all these studies and see if any particular genes emerged that were commonly upregulated or downregulated in the case of drug resistance. The result is a bit more complex than we had hoped. We looked at 95 studies, excluding those involving “radioresistance.” The gene that was most commonly altered in these studies was OAS1, appearing 21 times out of 190 opportunities (all 95 studies have up and down-regulated portions). That seems nice…a “big name” gene popping up at a frequency that, without crunching the numbers, appears to be significant. The problem is this: OAS1 appeared in both the resistance-upregulated and resistance-downregulated datasets (13 times up and 8 times down, to be specific). Therefore, we can’t make the blanket statement that OAS1 is upregulated in cases of drug resistance, nor can we surmise that suppressing the innate immune response (OAS1 is a big player there, after all) might overcome drug resistance. OAS1 is not unique in this respect.

That doesn’t mean that generation of lists of commonly up- and down-regulated genes involved in drug resistance would be entirely fruitless and couldn’t possibly spur insight. We’ve given these two lists the database IDs 129091122 and 129092122. In both cases, a gene had to occur at least 7 times (out of 95) to make the list, giving the lists a composition of 163 and 119 genes. 26 genes were found in both lists:

AREG

IFI27

C1orf24

IL1A

CA12

KYNU

CD24

LCN2

CEACAM6

OAS1

CXCR4

PEG10

DUSP6

SERPINB2

EMP1

SERPINE2

FAM129A

SOCS2

FSTL1

STC2

GPNMB

TSPAN8

HLA-DRA

UCHL1

HLA-DRB1

VCAN

Ignoring the fact that multiple genes are found in both lists, what broad categories of genes intersect with these lists? In the case of upregulation, the innate immune response does indeed seem relevant (e.g. genes upregulated early (vs late) in HCMV infection intersect with a P-value of 10-76). Erlotinib and neo-adjuvant therapy seem to do a fine job of upregulating common drug-resistance genes; if borne out in the lab/clinic, this would have obvious implications for cancer cocktail approaches. On the downregulation side, the metastatic (or not) nature of underlying tissues seems to be relevant. Specifically, transcripts downregulated in aggressively metastatic tissue overlap with transcripts that are downregulated in the case of drug resistance. For far deeper details, just plug either of the above dbase IDs into our “Fisher” app.

How about creating lists of, say, upregulated genes that weren’t found at all in the downregulated category? We tried that, but were met with discouragement. We found that RAB25 was the single best example of a transcript that was downregulated (in 8 studies), but never upregulated. Searching for validation of this characteristic of RAB25 in specific studies, the first study we stumbled upon was this: RAB25 confers resistance to chemotherapy by altering mitochondrial apoptosis signaling in ovarian cancer cells. There, it seems, RAB25 upregulation, not downregulation, correlates with resistance. Hmmmm.

Apparently, the same transcript may be upregulated in one resistance study, and downregulated in the next. My background in this field (a couple distant lectures and/or presentations) informed me that a handful of transporters are the main culprits in drug resistance, and I had hoped that this phenomenon would be obvious once multiple studies were compounded. This is not the case.

One might think that the above conundrum be resolved by examining particular drugs. That is, certain critical transcripts would always be upregulated in resistance to a particular drug. Cisplatin-resistance is the most commonly studied sort of resistance. Here, of 5 genes found in 4 out of 6 of the cisplatin-resistance studies we have on hand, 3 are both upregulated and downregulated, depending on the study: QPCT, SAA1, and MMP1 (TGFB2 is up in all four cases, and ANO1 is always down).

To attempt to clarify matters, we clustered the 190 datasets. Specifically, we performed Fisher’s exact test for each dataset against our entire database, generating millions of P-values. These P-values were the raw data for clustering (Cluster 3.0, k-means). 10 clusters were generated. The top genes found in each cluster are now found in our database with IDs 129137122, 129138122, 129139122, 129140122, 129141122, 129142122, 129143122, 129144122, 129145122, and 129146122.

Though the clusters did not nicely segregate according to drugs or cell types, as one might desire, some clarity was gained. Bearing in mind that both up- and down-regulated transcripts can be found in a single cluster, Cluster 0 transcripts tend to be upregulated on innate immune stimulation (e.g. via interferons). Cluster 3 transcripts tend to be upregulated in the case of metastasis and are enriched for cell-surface markers, while cluster 5 and 8 genes tend to be downregulated in metastatic cells. Cluster 6 genes have a strong tendency to be upregulated in cancer versus adjacent tissue and, rather bizarrely, downregulated on resveratrol treatment (P = 10-56). Other clusters are more nuanced. Thus, it would appear that investigators might wish to place special relevance on the status of cells with regard to innate immunity and metastasis when considering approaches that might mitigate drug resistance. Depending on the cluster, we see hints that particular drugs could, to some extent, reverse drug resistance: bromodomain inhibitors, noggin, losartan, gefitinib, etc. Other drugs, of course, could enhance drug resistance.

It is sometimes difficult to trust the output of clustering programs, so I performed an eyeball version of clustering in Excel. Give different colors to different significance levels (below, green indicates P<10-15), and then sort a column. Gather all columns where colors (indicating significance) percolated to the top. That’ll be cluster 1.Then move on to a column that doesn’t fall into cluster 1. Repeat. Believe it or not, this crude method matched up quite nicely with the software I used. To me, this sort of correspondence between the mathematical perfection of the clustering software and the childish simplicity of matching columns that have the same colors indicates that maybe we shouldn’t spend an excess amount of time/energy debating the merits of, say, “Euclidean distance” vs. “City block distance.” Below is a sliver of the result:


Again, for the fine details, just visit WhatIsMyGene, plug in database IDs (or your own datasets), and have fun.


Note 2/22/2022: We've added quite a few more studies involving resistance to our database. At this point, it's fairly obvious that genes upregulated in cells resistant to one sort of treatment may actually be downregulated on resistance to another treatment. This observation dampens hopes for across-the-board approaches to drug resistance. On the positive side, it may mean that resistance could be dealt with via drug cocktails; i.e. two drugs that trigger opposing resistance patterns could be combined in a treatment. At some point in the future, we'll re-cluster our resistance results. We'll be a bit more rigorous about finding an optimal number of clusters, look a bit deeper into commonalities in these clusters, and perhaps examine cases where 2 "resistance-complementary" drugs might be applied to particular maladies.

Note 9/17/2022: A quote from Comparative proteomic analysis identifies key metabolic regulators of gemcitabine resistance in pancreatic cancerSurprisingly, a number of proteins that were downregulated in MIA-GR8 cells have been reported to promote drug resistance in other cancer types. It's nice to validate our view above, but it's also disappointing to see that many researchers may still be stuck in a one-dimensional view of drug resistance.


whatismygene.com 

Saturday, May 15, 2021

Did Covid-19 Emerge from the Wuhan Institute of Virology?

I’m going to address this topic with a minimum of drama. Go away if you’re a conspiracy buff. Stick around if you’re interested in a frank, somewhat introspective take on this question from a dude (albeit a low-impact dude) who has actually tinkered with viruses. For "safety", I'll spell out my #1 point right here: scientists have knee-jerk responses too. For even more safety, let me also spell out the following at the start: I still find it unlikely that the virus emerged from the Wuhan Institute of Virology.

There are two widely disseminated documents from credentialed authors providing arguments against and for the notion that Covid-19 was lab-generated. On the “against” side, we have a Nature article from March of 2020. On the “for” side, we have Nicholas Wade’s take.

I recall reading the Nature article last year, shaking my head at some of the refutations within, and then moving on to other topics. Wade’s article reminded me of my early skepticism. I’ll affirm two of Wade’s points:

1) The Nature article argues that the absence of a “previously used virus backbone”* within Covid-19 provides evidence that there was no lab-manipulation. Let me say: this is malarkey and, at best, an embarrassment for Nature. I’ve generated “backbone” free viruses myself (on dengue, to be specific). You insert the viral sequence into a plasmid, perform in vitro transcription, and infect cells with the resulting RNA. If you designed the plasmid correctly, there should be no evidence of “backbone.” Even if you erred, the virus may quickly shirk garbagy, non-optimal sequences upon multiple passaging (it can be frustrating to insert “loss of function” mutations into a virus, as the virus might dispense with them surprisingly quickly, if they don’t kill the virus from the very beginning).

In case anyone wishes to nitpick: yes, the 30kb length of coronaviruses makes ordinary plasmid insertion tricky, if not impossible. But there are plenty of methods to generate these long viruses without evidence of a backbone.

It’s hard to believe that the esteemed authors of the Nature article weren’t aware of these viral basics. Why did they choose to offer this lame argument?

2) The argument is made that the spike protein’s interaction with the ACE2 receptor is not optimal; therefore, Covid-19 could not be the product of manipulation.

Again, this is absurd. You have to assume that any and all lab experiments involving Covid-19 would involve insertion of the theoretically optimal (for ACE2 binding) spike protein sequence. Here’s an example of an experiment that I would consider interesting: perform some sort of guided evolution to generate a myriad of spike protein sequences, and test them ALL for both ACE2 affinity and infectivity**. Take the “winners” of this process, insert them into the virus, and write a paper. That’s just one of a near infinite number of experiments you could perform.

Let’s imagine that Dr. Evil is indeed behind the Covid-19 pandemic. He, like any competent virologist, would not automatically assume that the virus that best binds ACE2 has the highest potential to wipe out the human race. It wouldn’t surprise me at all to find that such a virus would be severely handicapped, refusing to let go of ACE2 at any step, and unable to perform its various pleiotropic functions.

Again, it’s odd that virologists would even attempt to pass this argument off in a Nature article.

There’s further lameness in the Nature article. For example: some of the mutations in Covid-19 haven’t been mentioned in the literature as yet. The idea, I guess, is that any lab-generated mutations would already have been described. I won’t even bother refuting that.

I have to question at least one of Wade’s other arguments, however. This regards the appearance of a furin cleavage site within the virus. This is supposed to be some sort of smoking gun for lab experimentation. The site is only 4 amino acids long. It’s not easy to estimate the probability that nature would come up with this mutation. Bear in mind that coronaviruses are the absolute champions of a process called “RNA recombination.” Without going into detail, the furin site doesn’t have to emerge via a step-wise series of mutations…it could enter in one fell swoop. Again, if there’s any “garbage” RNA left over from recombination, it could be eliminated quickly via evolution, including further recombination. If a paper attempts to address the furin cleavage site appearance from a probabilistic perspective, be skeptical about the underlying assumptions about what viruses do and don’t do.

On the other hand, everybody in the virology world inserts furin cleavage sites in their viruses and “replicons.” It’s something we do.

So, where do I stand? The most dramatic thing I can say without feeling guilty is this: we’re far from eliminating the possibility of a lab-generated Covid-19. Nothing I’ve seen convinces me that the virus couldn’t have emerged from the lab in Wuhan. Certainly not the Nature commentary.

I’ll take Wade at his word when he says that the Wuhan lab is China’s #1 coronavirus research facility. Rather odd, no? The counter-argument, I guess, might be that Wuhan is an optimal location to study coronaviruses, because that region of China is coronavirus heaven. I don't know.

To be clear, there’s a huge difference between a lab accident and intentional release. I don’t see any reason to assume the latter. How has China emerged from this mess? With an economy that’s not any stronger than anyone else’s, and the clunkiest vaccines on the market. Infections have been minimized in China, but the emergence of variants threatens that. India surprised everyone with a minimum of infections and deaths…last year.

Returning to the question of why Nature published its lame refutation, let me offer a bit of introspection. I don’t want a lab accident to be the cause of Covid-19 and I feel compelled to argue against the possibility. Just as a big-time developer tires of apparently nit-picky regulations covering endangered insect species, virologists don’t want further restrictions on their activities. We feel like we know what we’re doing. I suspect that the authors of the Nature article feel the same.

Finally, if you point a gun at my head and inquire as to the most probable source of Covid-19, I'd have to lean strongly on the side of natural origin. If you've read the above and have concluded I'd think otherwise, sorry to disappoint. There are plenty of arguments to support the natural origin of Covid-19; most of them, unfortunately, are not very accessible to layfolk. Here's the one that I find most difficult to refute: the 97% similarity between Covid-19 and its closest relative, RatG13, means that Covid-19 diverged from RatG13 no later than the early 1980's, and probably earlier. Thus, Dr. Evil (or Dr. Carelessness) would need to introduce about 1,000 mutations into RatG13 over the years. Whether by site-directed mutagenesis, lab passaging, or directed evolution, that's a figure that nearly unimaginable to virologists. Note also that these sequence differences are spread all over the viral genome; there's no sign that, for example, the spike protein was singled out for special treatment.

Given the above, if you're dead-set on blaming the Wuhan Institute, the only remotely possible scenario that I see would be the following: WIV scientists gathered the Covid-19 virus, or something very closely related, and brought it into the lab, whereupon it escaped with few or no mutations. Given that folks have not identified any virus with higher Covid-19 similarity than RatG13, one might surmise that such a virus may have been collected outside of China. There are indeed studies wherein WIV scientists gathered viruses outside of China (e.g. Africa). Now, if Covid-19 has a natural origin in China, what's more likely: it spread in the chaotic environment of a wet-market, or it spread in the controlled environment of a virology institute? In the case of import from outside of China, one could accuse the WIV of carelessly handling a virus to which the local population may have little immunity. All very speculative, with no evidence at all at this point.

*I note that some metavirology folks use the term "backbone" to refer to the conserved portions of a viral sequence. However, that's not the case in the Nature paper, which points to a paper on Coronavirus construction methods, not broad sequence comparisons, following the term "backbone".

**In fact, a bit of Googling shows that the folks at Wuhan are familiar with Selex, a method that lets evolution, as opposed to "rational design", determine an experimental outcome. Check out, for example, A SELEX-Screened Aptamer of Human Hepatitis B Virus RNA Encapsidation Signal Suppresses Viral Replication. To be clear, this particular paper optimizes RNA, not a protein.


whatismygene.com 


Saturday, April 17, 2021

A Handful of Neoantigen-Vaccine Biotechs

A couple disclaimers before proceeding:

1) I own a tad of stock in two of the below-mentioned companies (Gritstone Oncology and Genocea Biosciences).

2) My historical stock-picking record is probably not so different than monkeys throwing darts at a list of candidates. Possibly worse.

Despite the above, I’m excited about a narrow field of research, that of neoantigen cancer vaccines. For intimate details on the subject, you can check out a wonderful review here. There may be some bias in that “wonderful” appraisal. In brief, researchers have long sought druggable targets that are unique to cancers, not healthy tissue. Historically, such targets have been elusive. However, we now realize the following:

1) Cancers often have a lot of mutated DNA.

2) A decent % of these mutations are represented in the form of mutated proteins.

3) A decent % of mutated proteins get chopped up into little bits, and these bits (specifically, 8-11 amino-acid peptides) get “displayed” on the outside of cells. These are “neoantigens.”

4) A certain % of these displayed peptides are recognized by the immune system as foreign.

5) Of these, some have the potential to trigger an anti-cancer immune response.

6) Clinical introduction of these peptides has provoked profound responses in some cases.

Of course, the above steps simplify the subject and make it sound easy. If you have any background in molecular biology, you’ll recognize that it’s a hell of a lot of work to get to step #6. Specifically, you’ll do exome sequencing, RNA-seq, possibly mass spectrometry and/or ribosomal profiling, and a lot of bioinformatics to arrive at a list of peptide candidates. Inject the peptides into the patient, and they may fail to stimulate the real immune system despite your high-tech tetramer assays. All of this work must be done fast, as the cancer grows and spreads in the patient’s body.

Genocea (GNCA) claims that your hard-earned peptide candidates may actually inhibit an immune response. This is a unique claim, prompting GNCA to create a name for such neoantigens: “inhibigens.” You can read GNCA’s recent, high-impact paper on the subject here. GNCA goes further, saying that some common assumptions in the field may be misguided; specifically, GNCA minimizes notions that peptide-HLA-binding affinity is strongly relevant to an immune-triggering effect (“immunogenicity”) and that “exotic” neoantigens, generated by events like gene fusions, aberrant splicing, translation of supposedly “non-coding” RNA (etc.), would be especially potent candidates.

GNCA can make these claims because it takes a step nobody else seems interested in: assaying every possible neoantigen candidate (potentially hundreds) from a patient for likely immunogenicity, rather than making the aforementioned assumptions to whittle down the candidate list to, say, 5 neoantigens. You can go to GNCA’s website to check out various presentations and the clinical status of its offerings (in brief: there’s nothing in phase III and no overtly amazing results as yet).

What I like here is that if GNCA is correct, the competition may be significantly handicapped. Meanwhile, GNCA accumulates real-world data on the parameters that really provoke the immune system, with a minimum of bioinformatic whittling. The stock may go to zero, get bought out, or skyrocket…I doubt it’ll tread water for the next 5 years.

Gritstone (GRTS) offers a more conventional approach to neoantigen generation, but they seem to be the most prominent and advanced “pure play” in the field. Until January, the company quietly advanced its neoantigen programs for cancer. That changed with the announcement of a Gates-Foundation sponsored program to use the underlying technology to discover and target Covid-related antigens (why not...these foreign peptides also get displayed on the surface of cells and provoke immune reactions), causing the stock to quadruple overnight. Another recent announcement involves an HIV-related collaboration with biotech biggy Gilead. The stock has come down to earth since January.

Gritstone’s potential edge derives largely from its method of using AI and bioinformatics to deconvolute big datasets to arrive at optimal neoantigen candidates. See GRTS’s own high-impact paper here. In addition to purely “personal” treatments, the company is also testing “off the shelf” peptides that are found in the most commonly mutated proteins in cancer. This sounds like a no-brainer, but cancers often weed out the mutated peptides that provoke strong reactions. Both the personalized and off-the-shelf products are in clinical trials, with intriguing early results. Again, for results and presentations, check out the website.

If GNCA and/or GRTS are lacking in any of the attributes that you seek in a biotech, you could also check out Agenus (AGEN), Achilles (ACHL), Personalis (PSNL), Vaccibody (VACC; Oslo stock exchange), or Instil (TIL). We note that one of the early neoantigen companies, Neon Therapeutics, was bought out in January of 2020 by Biontech (BNTX), which is now famous for its Covid-19 mRNA vaccine; the big guys are watching these entities. Moderna, in collaboration with Merck, is also in the field. A couple academic and private companies also have a stake in the future of neoantigen vaccines. Check out table 2 in a very recent, comprehensive review to get a sense of where neoantigen clinical trials stand at the moment.

In theory, proper neoantigen candidates should not provoke any reaction to healthy tissue. They’re just short peptides (or the RNA that codes for them, as in mRNA vaccines) that are already found in cancers, not entirely novel molecules that need to be carefully studied for toxicity in a number of tissues. I expect that the FDA will quickly adapt to the fact that these personalized peptides are generally benign once the positive data starts flowing at a high rate. Nevertheless, the market may be viewing neoantigen vaccine companies as typical biotechs, with narrow-target drugs that take 8 or more years, and billion dollar expenditures, to receive approval; there’s an opportunity here.

For future reference, GNCA and GRTS are now priced at $2.21 and $8.64.

***

Users of the WhatIsMyGene database might be interested to know that, whenever possible, we make note that a particular dataset was contributed by a biotech company (as opposed to most studies in the database, which derive from academic institutions). Thus, you can open the “relevant studies” app, enter “Syros” (for Syros Pharmaceuticals) and see a list of studies relevant to their drug candidates. You can then test these datasets against the entire database with the “Fisher” app and get a sense of what the drug candidate actually does. Of course, many biotechs keep data solely in-house, seeing secrecy as more important than the exposure that could result from a high-impact publication. Other times, biotech-relevant data may be published by an entity that appears purely academic. If you see a case where we’ve missed the fact that a biotech is behind a particular dataset, let us know!


whatismygene.com 


Wednesday, March 10, 2021

Cytokine Storms, Cancer and Metastasis, and Interferons

Here, we just wish to point out that a number of “canonical” lists lurk within our database and you may find them useful.

First of all, we have compilations of studies relating to “cytokine storms.” There’s a lot of debate about exactly what constitutes a cytokine storm. We see a number of “survivors vs. non-survivors” studies where folks die of viral infections without dramatic upregulation of cytokines in the blood. At the anecdotal level, here in Thailand, where hundreds of folks routinely die per year from dengue, I’ve encountered doctors who claim they’ve never seen anyone actually die of a “cytokine storm” per se. So here’s how we define “cytokine storm”: it’s when you die from an acute viral infection in rapid fashion. Thus we toss chronic infections like HBV and HIV from the dataset. We also toss Ebola, as the viral infection may be so potent that death is more likely than survivorship.

The “cytokine storm” lists are assembled from a mere 7 blood-based studies: GSE95233, GSE43777, GSE17924, GSE97287, GSE101702, GSE111368, and GSE111368. The upregulation list takes the database ID 118771101, and the downregulation list 118772101. We found that the upregulation list strongly overlaps with a GO “secretory granule” set, so we also constructed an upregulation list with secretory granule components excluded: 116037101.

We have plenty to say about cytokine storms, but we won’t start blathering at this point.

Next, we have some cancer-related lists. Based on more than 50 studies, we constructed broad lists of transcripts/proteins that are up- or downregulated in cancer tissue vs. adjacent tissue: 118765101 and 118766101. We also have lists relating to transcripts that are up- or downregulated in metastatic vs. primary tissue: 118767101 and 118768101.

You may intuit that cancer is so heterogeneous across tissues and cell-types that such lists would not be particularly informative. You'd be wrong. At the level of transcript and protein abundance, the same entities are altered again and again. It’s not surprising at all for us to come across a new cancer study, enter the upregulated (or downregulated) entities into our database, and find that the single best mimic of this cancer study comes from a canonical cancer list, with P-values that may exceed 10^-100. Again, we have plenty to say about cancer…but not now.

Finally, we have “canonically altered under interferon treatment” lists based on 20 studies (folks love to hit cells with interferons, with largely similar effects every time). We’ve excluded type II interferons (interferon gamma) from the list, as these have very different effects than type I/III interferons: 118769101 and 118770101.

In the future, we’ll update our cancer lists with new studies; we see perhaps 2 new studies per month. We could also build lists for specific cancer types, if the volume of data warrants it. Another task is to double-check every metastasis-related study, as comparisons can be made between primary tissues from which metastasis has and hasn’t arisen, as well as primary cancer tissue vs cancer tissue that has actually lodged in distant locales. We don’t expect that the cytokine lists will be updated soon, as the underlying studies are few and far between. We’ll probably never update the interferon lists, as it’s pretty clear what happens when you blast cells with these antiviral compounds.

whatismygene.com 


Wednesday, February 17, 2021

Underrated Genes

Nature magazine recently published a study of...biological studies. A number of questions were asked, one of them being, “which genes are most represented in the literature?” Not surprisingly, TP53 is the champion, with 9,232 publications. It’s a good read.

A question not addressed is, “What are the most under-represented genes in the literature?” Of course, it’s trivial to find genes that have no mentions at all. What we can do, however, is use our own database and ask, “Which genes have the largest disparity between inclusions in our database and inclusions in the literature?” The exercise is simple on its face, but there are a number of technicalities that make it a bit tricky. If we were writing an academic paper, we’d have to do 100X the work we’re putting into this post. Basically, though, our procedure works like this:  Download a list of genes ranked according to literature mentions. Convert these gene IDs into the format used in the database. Generate a frequency table of all genes in our database. Compare the frequencies in our database against the frequencies in the literature.

The list of genes according to literature mentions is found here: ftp://ftp.ncbi.nih.gov/gene/GeneRIF/generifs_basic.gz

With the understanding that there are a number of ways in which results can be skewed, here’s a list of the most under-rated players in the genomic universe:

RTP4

VSIG2

CLIC6

FAM198B

HIST1H2BD

MT1L

MOXD1

CENPK

ANKRD37

CMBL

PLBD1

TUBA1C

ARHGAP11A

TMEM154

HIST1H2BI

NMES1

TMEM140

PKIA

ADGRL2

KBTBD11

NT5DC2

C15orf15

RSL24D1

RPL27A

FAM49A

PGM5

RGL1

CLMN

EVI2A

TFEC

RPL18A

RPL21

SRM

CALML4

OLFML2A

RPS8

ENDOD1

KDELR3

RPS11

GNG4

TMEM56

SH3BGRL2

CIART

ENPP5

GBP6

RSRP1

COX6A2

GPRIN3

GPRC5C

TMEM71

NRIP3

MFAP3L

CPNE2

ABLIM3

SMIM14

HIST1H2BM

SLC46A3

EVI2B

PCP4L1

TRNP1

GBP4

SLC16A14

RBP7

SLFN13

FAM84A

RAPGEF5

TM6SF1

NSG2

VAT1L

EPPK1

RPL27

DNAJA4

PGAM2

TTC39C

TRANK1

GBP7

N4BP2L2

MEGF6

CDH19

FIBIN

TINAGL1

CCDC3

LONRF2

DDX60L

MXRA7

GPR137B

CENPV

GNG12

CCDC85A

GRAMD3

FAM105A

STRBP

ZNF608

KIAA1551

LRRC2

UAP1L1

MEGF9

EPB41L4A

PLEKHA4

METTL7B


RTP4 is the champion, with few mentions in the literature but more than 700 appearances in our database. Googling RTP4, it seems that there’s no dearth of studies on this gene, but we’re sticking with the above NIH list of literature mentions. Next on the list is VSIG2. A Google search does seem to indicate that nobody cares about this sad gene. It’s hard to even get a clue as to its function.* Nevertheless, it appears 699 times in the database; perturb a cell and there’s a decent chance you’ll alter VSIG2 expression.

We ran a Fisher analysis of an extended, 500-ID list of undervalued genes against our entire database. As might be expected, there’s no massive enrichment for any particular group. There does seem to be a tendency for genes with short transcripts and genes that are depleted in P-bodies to be represented on the list (unadjusted log(P-values) of -7.5 and -5.8). Eyeballing the list, a number of ribosomal proteins can be seen. Perhaps folks view the ribosome as a big unified glob, and don’t care to tinker with its individual components.

The opposite task, that of generating a list of “overrated” genes, is even trickier, and we won’t bother with it here. In the end, genes like TP53 would dominate the list and, given TP53’s role in cancer, labeling it “overrated” or “overstudied” would hardly be fair.

 

*Let’s say you want to know about VSIG2’s function. You can use our tools. First, you enter VSIG2 into our Coregulation app. You’ll get a list of coregulated genes. Take that list and enter it into the Fisher app. To spare you this [minimum] trouble, the swarm of genes with which VSIG2 is coexpressed looks to be hugely involved in the cell cycle, altered by a large array of common drugs (e.g. glucosamine), and also relevant to viral infections. Using the coregulation tool alone, individual genes that are strongly coexpressed with VSIG2 include TRIB3, CHAC1, ASNS, and many more. You can also note that CA9, FAM111B, NREP, and more have a fairly strong tendency to be expressed in the opposite direction to VSIG2 (i.e. when VSIG2 is up, CA9 tends to be down).


whatismygene.com 


Tuesday, February 9, 2021

Alzheimer's According to Various Brain Tissues

Below is a table of all individual studies that contributed to our "canonically altered in the Alzheimer's brain" lists. The numbers show log(P-values) for the intersections between the canonical upregulated and downregulated lists against individual studies. If this were an academic paper, a referee might (justifiably) complain about our method; technically, one shouldn't intersect a compendium of studies with studies from which the compendium is derived and go on to derive P-values. A better approach would be to remove all contributions of a particular study from the canonical lists, and then perform the Fisher test. This exercise has to be performed for all studies; a lot of work. Practically speaking, however, we entered such a large volume of studies into these canonical lists that the P-values won't be exaggerated to any great extent. 

I've broken down the studies according to brain regions. Hopefully, my inadequacy in basic neuroanatomy won't be too obvious (note that I bundled everything with the terms "cortex" and "cortical" into a single group). The goal is to ascertain whether Alzheimer's, as defined by the "canonical" lists, is particularly potent in particular regions. One could also ask if Alzheimer's trends reverse in particular tissues of Alzheimer's patients.


region

study

up

down

all

downregulated (MS) in Alzheimer's brain (vs controls)(fc/p: s: Large-scale proteomic analysis of Alzheimer’s disease brain)

0.91

-3.98

downregulated in Alzheimer's brain (all regions incl)(fc/p:  GSE5281: anatomically and functionally distinct regions of the normal aged)

0.21

-25.4

downregulated in Alzheimer's brain (all regions incl)(fc/p: GSE118553: asymptomatic and symptomatic alzheimer brains)

0.21

-18.17

downregulated in Alzheimer's disease in three brain tissues (GSE131617)

-2.64

-10.51

upregulated (MS) in Alzheimer's brain (vs controls)(fc/p: s: Large-scale proteomic analysis of Alzheimer’s disease brain)

-1.45

-0.75

upregulated in Alzheimer's brain (all regions incl)(fc/p:  GSE5281: anatomically and functionally distinct regions of the normal aged)

-16.78

0.45

upregulated in Alzheimer's brain (all regions incl)(fc/p: GSE118553: asymptomatic and symptomatic alzheimer brains)

-20.18

0.45

upregulated in Alzheimer's disease in three brain tissues (GSE131617)

-2.85

0.48

astrocytes

downregulated in alzheimer's astrocytes (with c9orf72 mutations) vs. healthy astrocytes (raw data fc/p: GSE142730)

-0.51

-6.73

downregulated in astrocytes of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

0.52

-3.65

downregulated in astrocytes of alzheimer's patients at braak stage III or greater (GSE29652: astrocyte transcriptome in the aging brain)

0.22

-20.52

upregulated in alzheimer's astrocytes (with c9orf72 mutations) vs. healthy astrocytes (raw data fc/p: GSE142730)

-1.11

0.49

upregulated in astrocytes of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

-11.29

0.59

upregulated in astrocytes of alzheimer's patients at braak stage III or greater (GSE29652: astrocyte transcriptome in the aging brain)

-2.86

0.48

cerebellum

downregulated in cerebellum of "asymptomatic alzheimers" vs normal (GSE118553)

-1.13

-4.9

upregulated in cerebellum of "asymptomatic alzheimers" vs normal (GSE118553)

-7.57

0.43

choroid plexus

downregulated in choroid plexus of alzheimers patients (GSE61196)

0.19

-1.91

upregulated in choroid plexus of alzheimers patients (GSE61196)

-1.21

-0.43

cortex

downregulated (MS) in frontal cortex of Alzheimer's patients (s4: fc/p: proteomic analysis of the frontal cortex in Alzheimer’s)

0.32

-4.58

downregulated in alzheimer's cortex (s3: fc/p: link between amyloidosis and neuroinflammation)

0.22

-12.37

downregulated in cortex of Alzheimer's patients (fc/p: GSE15222: brain transcript expression in Alzheimer disease)

0.26

-43.23

downregulated in dorsolateral prefrontal cortex of alzheimer's patients (raw data w/ttest: GSE53697: ELAV-like protein binding to coding and non-coding)

-0.74

0.22

downregulated in entorhinal cortex of alzheimer's patients (fc/p: GSE48350: normal brain aging are sexually dimorphic)

0.22

-1.14

downregulated in entorhinal cortex of alzheimer's patients (GSE26972:  loss of hnRNP-A/B in Alzheimer's disease)

0.12

-53.23

downregulated in frontal cortex of Alzheimer patients (GSE36980: Altered expression of diabetes-related genes in Alzheimer's disease brains)

0.2

-48.11

downregulated in frontal cortex synaptoneurosome of Alzheimer's patients (fc/p: GSE12685: synaptoneurosomes identifies neuroplasticity genes overexpressed)

-0.45

0.5

downregulated in neocortex of alzheimer's patients (GSE37263: gene expression in the neocortex of Alzheimer's )

0.24

-96.02

downregulated in neocortex of alzheimer's patients (GSE37264: alternative splicing in Alzheimer's disease)

0.14

-66.08

downregulated in parietal cortex of Alzheimer's patients (GSE16759: miRNA and mRNA expression in Alzheimer's disease)

-1.21

-18.78

downregulated in prefrontal cortex of alzheimer's patients (fc/p: GSE33000: human prefrontal cortex underlies two neurodegenerative)

0.19

-29.5

downregulated in temporal and frontal cortex of Alzheimer's patients (GSE139384: Pathomechanism of Kii ALS/PDC)

-3.62

0.69

downregulated in temporal cortex of alzheimers patients (vs. vascular dementia patients)(GSE122063)

0.17

-32.84

upregulated (MS) in frontal cortex of Alzheimer's patients (s4: fc/p: proteomic analysis of the frontal cortex in Alzheimer’s)

-2.19

0.77

upregulated in alzheimers (frontal pole and occipital cortex)(GSE84422: regional vulnerability to Alzheimer's disease)

-26.42

0.41

upregulated in alzheimer's cortex (s3: fc/p: link between amyloidosis and neuroinflammation)

-4.97

0.49

upregulated in cortex of Alzheimer's patients (fc/p: GSE15222: brain transcript expression in Alzheimer disease)

-13.67

0.59

upregulated in entorhinal cortex of alzheimer's patients (fc/p: GSE48350: normal brain aging are sexually dimorphic)

-13.4

-0.7

upregulated in entorhinal cortex of alzheimer's patients (GSE26972:  loss of hnRNP-A/B in Alzheimer's disease)

-16.26

0.19

upregulated in frontal cortex of Alzheimer patients (GSE36980: Altered expression of diabetes-related genes in Alzheimer's disease brains)

-25.66

0.28

upregulated in frontal cortex synaptoneurosome of Alzheimer's patients (fc/p: GSE12685: synaptoneurosomes identifies neuroplasticity genes overexpressed)

0.31

-4.64

upregulated in neocortex of alzheimer's patients (GSE37263: gene expression in the neocortex of Alzheimer's )

-30.18

0.44

upregulated in neocortex of alzheimer's patients (GSE37264: alternative splicing in Alzheimer's disease)

-15.5

-0.45

upregulated in parietal cortex of Alzheimer's patients (GSE16759: miRNA and mRNA expression in Alzheimer's disease)

-9.24

-1.83

upregulated in prefrontal cortex of alzheimer's patients (fc/p: GSE33000: human prefrontal cortex underlies two neurodegenerative)

-19.04

0.42

upregulated in temporal and frontal cortex of Alzheimer's patients (GSE139384: Pathomechanism of Kii ALS/PDC)

0.78

-6.02

upregulated in temporal cortex of alzheimers patients (vs. vascular dementia patients)(GSE122063)

-2.48

0.3

endothelial

downregulated in prefrontal cortical endothelial cells of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

0.09

-4.14

upregulated in prefrontal cortical endothelial cells of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

-3.67

1.25

gray matter

downregulated in incipient Alzheimers gray matter (Microarray analyses of laser-captured hippocampus)

-3.24

-0.8

downregulated in severe Alzheimers gray matter (GSE28146: Microarray analyses of laser-captured hippocampus)

0.16

-5.83

upregulated in incipient Alzheimers gray matter (Microarray analyses of laser-captured hippocampus)

0.17

0.39

upregulated in severe Alzheimers gray matter (GSE28146: Microarray analyses of laser-captured hippocampus)

-0.57

0.38

hippocampus

downregulated in hippocampus of alzheimer's patients (fc/p: GSE29378: cell type changes in Alzheimer's disease)

0.21

-4.02

downregulated in hippocampus of alzheimer's patients (raw data fc/p: GSE67333: Alzheimer's Disease Reveal Neurovascular Defects)

0.41

-4.44

downregulated in hippocampus of alzheimer's patients (raw data w/ttest (p<.01): GSE113524: autism/intellectual disability somatic mutations in Alzheimer's brains)

0.06

0.13

downregulated in hippocampus of severe Alzheimer's patients (Incipient Alzheimer's disease: microarray correlation analyses)

0.19

-34.13

upregulated in hippocampus of alzheimer's patients (fc/p: GSE29378: cell type changes in Alzheimer's disease)

-37.91

-0.74

upregulated in hippocampus of alzheimer's patients (raw data fc/p: GSE67333: Alzheimer's Disease Reveal Neurovascular Defects)

-0.43

0.59

upregulated in hippocampus of alzheimer's patients (raw data w/ttest (p<.01): GSE113524: autism/intellectual disability somatic mutations in Alzheimer's brains)

0.21

0.46

upregulated in hippocampus of severe Alzheimer's patients (Incipient Alzheimer's disease: microarray correlation analyses)

-4.23

0.41

microglia

downregulated in prefrontal cortical microglia of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

0.46

0.64

upregulated in prefrontal cortical microglia of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

-0.46

0.63

neuron

downregulated in excitatory neurons of alzheimer's patients (S2: fc/p: Single-cell transcriptomic analysis of Alzheimer?s disease)

-0.48

-21.72

downregulated in prefrontal cortical excitatory neurons of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

-0.43

-8.22

downregulated in prefrontal cortical inhibitory neurons of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

0.04

-0.78

upregulated in excitatory neurons of alzheimer's patients (S2: fc/p: Single-cell transcriptomic analysis of Alzheimer?s disease)

-0.48

-1.47

upregulated in prefrontal cortical excitatory neurons of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

0.45

1.27

upregulated in prefrontal cortical inhibitory neurons of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

0.04

0.09

nscs

downregulated in neural progenitor cells from sporadic alzheimer's patients (GSE117586: iPSC Models of Alzheimer's Disease)

-0.51

-12.37

upregulated in neural progenitor cells from sporadic alzheimer's patients (GSE117586: iPSC Models of Alzheimer's Disease)

0.22

-5.74

olfactory

downregulated in olfactory of advanced alzheimers patients (vs. control and initial disease)(GSE93885: olfactory bulb transcriptome during Alzheimer®s disease evolution)

-0.55

-8.59

upregulated in olfactory of advanced alzheimers patients (vs. control and initial disease)(GSE93885: olfactory bulb transcriptome during Alzheimer®s disease evolution)

-5.57

-0.43

oligodendrocyte

downregulated in prefrontal cortical oligodendrocytes of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

0.39

-0.75

upregulated in prefrontal cortical oligodendrocytes of alzheimer's patients (single nucleus sequencing) (s3: endothelial cells and neuroprotective glia in Alzheimer’s)

-1.38

0.54

posterior cingulate

downregulated in posterior cingulate of early onset alzheimer's patients (fc/p: GSE39420: sporadic and monogenic early-onset Alzheimer's disease)

0.2

-40.22

upregulated in posterior cingulate of early onset alzheimer's patients (fc/p: GSE39420: sporadic and monogenic early-onset Alzheimer's disease)

-20.14

0.45

temporal gyrus

downregulated in middle temporal gyrus of alzheimers patients (GSE132903: Alzheimer's Disease Middle Temporal Gyrus: Importance)

0.2

-123.42

downregulated in temporal gyrus of alzheimer patients (GSE109887: Alzheimer's disease-associated (hydroxy)methylomic)

0.2

-122.84

upregulated in middle temporal gyrus of alzheimers patients (GSE132903: Alzheimer's Disease Middle Temporal Gyrus: Importance)

-38.49

0.49

upregulated in temporal gyrus of alzheimer patients (GSE109887: Alzheimer's disease-associated (hydroxy)methylomic)

-43.17

0.4

temporal lobe

downregulated in lateral temporal lobe of Alzheimer's patients (vs healthy elderly)(raw data w/ttest (p<.001): GSE104704: landscape of normal aging in Alzheimer's disease)

0.22

-20.52

upregulated in lateral temporal lobe of Alzheimer's patients (vs healthy elderly)(raw data w/ttest (p<.001): GSE104704: landscape of normal aging in Alzheimer's disease)

-3.87

0.49



Initially, I included the 5 groups, both up and down-regulated (10 more columns), from the previously mentioned Alzheimer's clustering study. It generated a confusing, difficult-to-post mess, so I simplified. In general, these 10 groups intersected the individual studies as might be expected, with groups B1 and B2 often reversing the trends seen above. 

To answer the initial question: the up/down-regulation patterns in various tissues strongly tended to conform to up/down-regulation in our canonical lists. In other words, these canonical fingerprints don't strongly depend on tissues. Particularly strong intersections were seen between the broad canonical lists and two temporal gyrus studies. The posterior cingulate and general cortex also showed nice P-values.

Red text signifies cases where the direction (+ or -) of the canonical lists versus individual studies are not the same (e.g. intersection of the canonically upregulated list with transcripts downregulated in a particular study generates an interesting P-value). There are only four of these cases. Most interesting, in my opinion, is the case in which transcripts upregulated in Alzheimer's neural stem cells significantly overlap canonically downregulated Alzheimer's transcripts. Is it possible that the levels of particular transcripts (and their accompanying proteins) in NSCs generates a reverse signature in surrounding tissues?

A number of tissues do not strongly overlap with the canonical lists: oligodendrocytes, the choroid plexus, and microglia in particular. One should not naively conclude that such tissues are irrelevant to Alzheimer's...such tissues might be especially relevant.

In general, we can't claim that this exercise was particularly insightful. The main point is to show that the Alzheimer's signature is evident in multiple regions of the Alzheimer's brain.

whatismygene.com 


Mixing WIMG and AI

We've been tinkering with incorporating AI (Gemini) with our Fisher app. The idea is not so tricky...you feed the AI an identity/ruleboo...