aqDeveloper environment for foundation models

Replication of Single-Nucleus Transcriptomic Findings in Neurodegeneration

Aquin is starting a cross-study replication of single-nucleus transcriptomics. One disease, every public dataset for it, the same pipeline throughout. Target: December 2026.

Cross-Study Replication of Single-Nucleus Transcriptomic Findings in Neurodegeneration

Bar charts comparing original Mathys et al. analysis to Murphy and Skene reanalysis: differentially expressed genes fell from 23,923 to 26; median absolute log2 fold change rose from 0.16 to 2.34

Genes called differentially expressed

FDR < 0.05 · log-scaled bars · 549× fewer genes

Size of those gene expression changes

median absolute log₂ fold change · 15× larger effects

Mathys et al. (2019), the first Alzheimer's snRNA-seq study, reprocessed with donor-level statistics. Source: Murphy & Skene, eLife 2023;12:90214. Far fewer genes survive correct analysis, but those that do show much larger effects.

Single-nucleus RNA sequencing of postmortem brain has become the standard method for identifying cell-type-specific changes in neurodegenerative disease, yet its findings are rarely tested across independent studies. A published reanalysis of the first Alzheimer's snRNA-seq dataset found that applying donor-level statistics and standard quality control reduced the number of differentially expressed genes 549-fold, raising the question of how many reported effects across the field would hold up under the same scrutiny. We will select a neurodegenerative disease that has several public snRNA-seq datasets but no existing integration, reprocess all of them through a single uniform pipeline using donor-level analysis with technical confounders controlled, and determine which published cell-type findings replicate. We will then identify which replicated genes encode proteins known to be secreted or detectable in cerebrospinal fluid or plasma, producing a shortlist of candidate biomarkers grounded in findings that hold across independent datasets rather than a single study. The purpose is to give that disease's research community a clear account of which of its single-nucleus findings are reliable, and a well-supported starting point for biomarker selection. The project will produce an open harmonization pipeline, an integrated dataset, a replication table, and a candidate biomarker shortlist, with completion targeted for December 2026.

Why findings don't travel

Single-nucleus RNA sequencing of postmortem brain is a primary way the field finds cell-type-specific changes in neurodegenerative disease. Those findings are almost never tested across independent studies. Different labs, different pipelines, different quality control. Claims move into reviews and follow-on work without a clear map of what survives someone else's data.

Murphy and Skene made that concrete in eLife in 2023, reanalyzing Mathys et al., the first Alzheimer's snRNA-seq dataset. Applying donor-level statistics and standard quality control cut the number of differentially expressed genes from 23,923 to 26, a 549-fold drop. The genes that survived had effects roughly 15 times larger than the originals, with a median absolute log2 fold change of 2.34 against 0.16. Their point was not that the original study was careless, since its methods were standard at the time, but that the analysis choices matter enormously. If one careful pass does that to a landmark study, the same question applies across the literature.

The approach

We are selecting a neurodegenerative disease with several public snRNA-seq datasets that have not been analyzed together. All of them get reprocessed through a single pipeline: donor-level statistics, technical confounders controlled, the same quality control end to end. Against that shared baseline we ask which published cell-type findings replicate. From the ones that do, we pull genes encoding proteins known to be secreted or detectable in CSF or plasma. That is the biomarker shortlist, grounded in effects that hold across datasets rather than a single paper.

What we're working on

An open harmonization pipeline, an integrated multi-study dataset, a replication table of which cell-type findings hold, and a shortlist of CSF and plasma candidate biomarkers. The aim is to give that disease's research community a shared baseline they can check their own results against, and a better-supported starting point for what to measure next.

This is part of Aquin's broader neuro work: building the infrastructure that makes findings comparable across studies.

Not sure if Aquin is right for you?