1 转录组 RNA-seq
1.1 Source 1
1.1.1 DESeq2 Differential Expression Analysis
跳转: DESeq2 Differential Expression Analysis
Differential expression analysis of RNA-seq data was performed using the DESeg2 R package (v1.36.0). Genes with absolute fold change (IFC= > 1.8) and Benjamini-Hochberg adjusted p-value (padj) <0.01 were regarded asdifferentially expressed genes. Three pairwise comparisons were conducted:
- B vs. A: AD group (astrocyte + tau PFF) vs. control group (astrocyte)
- F vs. A: PF-562271 intervention group (astrocyte + tau PFF + PF271) vs. control group
- F vs. B: PF-562271 intervention group Vs. AD group
References:
- Karthivashan et al., 2026. 5xFAD AD model; Methods 2.6 describes mRNA RNA-seq and DESeq2-based analysis. PMC13240024
1.1.2 Functional analysis of DEGs
跳转:Functional analysis of DEGs
To explore the functions of DEGs across different groups, Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses were performed using the ClusterProfiler package version 4.4.4. An adjusted p-value (padj) < 0.05 was considered statistically significant.
1.2 Source 2
1.2.1 mRNA RNAseq and data analysis
跳转:mRNA RNAseq and data analysis
Paired-ended sequencing was performed on Illumina’s NovaSeq 6000 sequencing system (LC Sciences). Sequence adapters and low-quality reads were removed using Cutadapt. Quality control checks on raw sequence data were performed with FastQC. Sequencing reads were then mapped to the mouse reference genome (release-96) using the HISAT program. After the final transcriptome was generated, the R package DESeq2 was used to estimate the expression levels of all transcripts (Supplementary-Data File 1). Differentially expressed genes (DEGs) from different comparisons were identified using the following criteria: fold change (FC) ≥ 1 for upregulated and ≤ −1 for downregulated DEGs with adjusted p < 0.01. Principal component analysis (PCA) was performed to determine the variability and repeatability of samples as described earlier.37 Volcano plots were generated by the R package EnhancedVolcano (1.20.0) to visualize the overall distribution of DEGs. Gene Ontology (GO) analysis of DEGs was performed using the R Package clusterProfiler (version 3.18.1), and the significance level for GO terms was set at p < 0.05. Gene set enrichment analysis (GSEA, version 4.2.3) was performed using the RNA-seq data (Supplementary-Data File 1) against the gene sets described in the Molecular Signatures Database (version 7.5.1, C2 and C5 categories). The R package vissE (version 1.10.0) was used to exploit this gene set–gene set overlap relationship to enable the interpretation of clusters of gene sets and further summarize results. To identify modules of highly correlated genes, WGCNA was performed on RNA-seq data by the WGCNA package in R (version 1.72-5). To identify modules with different expression patterns, a soft threshold power was assigned to create co-expression networ
1.3 Source 3

1.4 Source 4
1.5 Source 5
2 单细胞 scRNA-seq / snRNA-seq
2.1 Source 6
DOI: 10.1016/j.jhep.2025.11.008
2.1.1 Single-Cell Sequencing
Single cell suspensions collected from CASH samples were stained with antibodies against CD45 for FACS sorting. scRNA-seq libraries were prepared by NovelBio Bio-Pharm Technology Co., Ltd (Shanghai, China). using Chromium Next GEM Single Cell 5’ V2 kit (10X Genomics, Pleasanton, CA) according to the manufacturer’s protocol. Count matrices were qualified and analyzed using Seurat, a R package (version 5.10). Other details are supplied in the supplementary materials and methods.
2.1.2 Single-Cell Data Processing
Count matrices were qualified and analyzed using Seurat, a R package (version 5.10). Low quality cells were defined as those with less than 500 detected genes or over 6000 genes ormore than 20% mitochondrial UMI counts, were excluded. All samples were normalized by LogNormalize method, and scaled to achieve a mean of zero and a standard derivation of one.The top 2000 variable features were chosen for principal component analysis (PCA). Anchor-based CCA integration was utilized to remove batch effect and integrate multiple samples. The top 30 principals were used to construct a shared neighborhood graph and clusters with a givenset of resolution presented in the following sections. The reduction was achieved by usinguniform manifold approximation and projection (UMAP).
2.1.3 Cell Type Annotation
In Figure 1C, the major clusters were achieved by performing the first-round clustering with aresolution of 0.1. The second-round clustering on myeloid with a resolution of 0.8, T cell witha resolution of 0.7, and NKs with a resolution of 0.4. Differential expressed genes were identified by FindAllMarkers function with thresholds on log-scale fold change of >0.25 and p value of <0.05 (Wilcoxon rank-sum test, and the p values have been corrected by falsediscovery rate (FDR)).
2.2 Source 7
2.2.1 Single-Cell Sequencing via 10x Genomics
scRNA - seq library preparation was performed according to the 10x Genomics Chromium Single Cell 3’ Reagents KitV3, user guide except that sample volumes containing 25,000 cells are loaded onto the 10x Genomics flow cell to capture ~10,000 total cells. The 10x Genomics workflow was then followed according to the manufacturer protocol and libraries are pooled at equimolar concentrations for sequencing on an llumina NovaSeq 6000, targeting ~50,000 reads per cell FASTQ files are aligned to the human GRCh38 transcriptome using the Cell Ranger (version 3.0.2) count command, with the expected cells set to 10,000 and no secondary analysis performed.
2.2.2 scRNA - seq data visualization and differential gene analysis
跳转:Single-Cell Data Processing
UMI count tables were read into Seurat (version 3) for preprocessing and clustering analysis. Initial quality control(QC) was performed by log - normalizing and scaling (default settings) each dataset followed by principal componentanalysis (PCA) performed using all genes in the dataset. Seurat’s “ElbowPlot” function was used to select principal components (PCs) to be used for clustering along with a resolution parameter of 0.5 and clusters identified as beingdoublets, gene - poor, or dividing are removed from the dataset prior to downstream analysis. Secondary QC cutoffs were then applied to retain only cells with < 20%-25% ribosomal genes, 12.5% mitochondrial genes,> 500 genes but less than double the median gene count, and 500 UMI but less than double the median UM count. At this point one hPS19 and one hPS - 5X sample were discarded due to abnormal 10x cell capture and transcript amplification. Thefinal samples that yielded clean and high viability isolates and passed QC are as follows: four hWT mice (pooled twoper 10x run into two 10x runs); three h5xFAD mice (two pooled into one 10x run, one in a 10x run independently;two hPS19 mice pooled into one 10x run; one hPS - 5X mouse in one 10x run. Cells passing QC for each sample were then merged using Seurat’s “merge” function and datasets were processed using Seurat’s integrated analysisworkflow. Briefly, samples from individual mice were integrated using the “FindntegationAnchors” and”IntegrateData” commands using dimensions 1:25. Datasets were then scaled, and sources of technical variation areregressed out (number of genes, percent ribosomal genes, and percent mitochondrial genes) and PCA wasperformed using Seurat’s “RunPCA” command. A shared nearest neighbor (SNN) plot was generated using Seurat’s”FindNeighbors” function using PCs 1:40 as input, clustering was performed using the “FindClusters” function and a resolution parameter of 0.3, and dimension reduction was performed using the “RunuMAP” function with the same PCs used for generating the SNN plot. Differentially expressed genes (DEGs) were determined between clusters usingthe “FindAllMarkers” function, which employs a Wilcoxon rank sum test with and FDR cutoff of 0.01, an LFC cutoff of 0.25, and the requirement that the gene be expressed in at least 10% of the cluster and clusters are labeledaccording to manual curation of the differential gene lists.
2.3 Source 8
文章跳转:Single-cell transcriptomic analysis of Alzheimer’s disease | Nature
2.3.1 Quality control for cell inclusion
跳转:Quality control for cell inclusion
The initial dataset contained 80,660 cells with a median value of 1,496 counts. As initial reference, the entire dataset was projected to the 2D space using t-distributed stochastic neighbor embedding (t-SNE) on the top 10 principal components. The tSNE coordinates were used to visualize potential biases in apparent cell similarity due to differential cellquality. For each cell, the following quality measures were quantified: (i) the number of genes for which at least one read was mapped, which is indicative of library complexity, (ii)the total number of counts, (iii) the percentage of counts mapping to the top 50 genes, and(iv) the percentage of reads mapped to mitochondrial genes, which may be used toapproximate the relative amount of endogenous RNA, and it is commonly used as a measureof cell quality. Cells with a high ratio of Mitochondrial (MT) relative to endogenous RNAs had low starting amounts of RNA, which might indicate that source cells were dead or stressed,resulting RNA degradation. Outlier cells in these quality metrics were found to clustertogether in the tSNE 2D space. Based on these observations and subsequent scatter plotanalyses, cells with less than 200 detected genes, and cells with abnormally high ratio ofcounts mapping to MT genes, relative to the total number of detected genes, were removed.Specifically, given a highly skewed empirical distribution of the MT ratio values (i.e., having an elbow shape clearly separating high and low scores) outlier cells were classified in twogroups using the k-means clustering algorithm (k=2) on the MT ratio, and subsequently removed. Only counts associated with protein-coding genes were considered;mitochondrially encoded genes, and genes detected in less than 2 cells were excluded. After applying these filtering steps, the dataset included 17,926 genes profiled in 75,060 nuclei.
2.3.2 Cell clustering
All 75,060 cells were combined into a single dataset. Normalization and clustering were done with the scanpy package. Briefly, counts for all nuclei were scaled by the total library size multiplied by 10,000, and transformed to log space. 3,188 highly variable genes wereidentified based on dispersion and mean, the technical influence of the total number of counts was regressed out, and the values rescaled. These preprocessing steps were performed by sequentially using the functions normalize_per_cell, filter_genes_dispersion,loglp.regress_out, and scale in scanpy. PCA was performed over the variable genes, and tSNE was run over the top 10 PCs using the MulticoreTSNE package by Ulanov (2017). The top 50 PCs were used to build a k-nearest-neighbors cell-cell graph with k=30 neighbors,subsequently spectral decomposition over the graph was performed with 15 components,and the Louvain graph clustering algorithm was applied to identify cell clusters. These analyses were performed using the functions pca, neighbors, and louvain in scanpy. Wec onfirmed that a number of PCs greater that 30 captures 100% of the variance of the data.The initial pre-clustering analysis resulted in 20 pre-clusters with a median number of 2,990 cells, ranging from 413 to 15,900 cells, after excluding 2 pre-clusters of 360 and 791 cellsthat reflected low-quality cells (these showed mixed cell-type markers, extreme complexitywith many more genes expressed than other cells, either too many or too few reads, and inone case cells isolated almost exclusively from one individual). For each pre-cluster,differentially expressed genes were detected using the variance adjusted t-test asimplemented in the function rank_genes_groups in scanpy. The top 500 ranking genes were extracted for each cluster, and used to test for overlap with markers as previously reported.The same clustering protocol was used for both pre- and sub-clustering analyses. During sub-clustering, additional potentially spurious clusters representing low-quality or doubletcells were detected on the basis of extreme separation from the rest of sub-clusters from thesame cell-type. From these, those having distinctly high number of total counts and mixedexpression of markers from different cell-types were tagged as potential doublets and notconsidered for downstream analyses, resulting in a total of 70634 cells
2.3.3 Cell type annotation and sub-clustering
跳转:Cell type annotation and sub-clustering
For each pre-cluster, we assigned a cell-type label using statistical enrichment for sets ofmarker genes1821, and manual evaluation of gene expression for small sets of known markergenes. Enrichment was statistically assessed using the hypergeometric distribution (Fisher’sexact test) and FDR correction over all gene sets and pre-clusters. Broad cell-type clusters were defined by grouping together all pre-clusters corresponding to the same cell type. Sub-clustering analysis was performed independently over each broad cell-type cluster.
2.3.4 Marker Identification
For sub-clusters a set of markers (specifically overexpressed) genes was defined by differential expression analysis of the cells grouped in each sub-cluster against the remaining cells within the corresponding broad cell-type cluster. This analysis was applied to all celltypes independently. Significantly over-expressed genes were defined based on the Wilcoxon-rank-sum test with a FDR corrected p-value <=0.01 and a logFC of 0.5. Only genes detected in at least 25% of the cell within the given subcluster were considered. Gene Ontology (GO) enrichment analyses were performed using Metascape and the p.value ranked gene lists as input. For all cases we used the set of 17,926 protein-coding genesincluded in the QC data as background.
2.3.5 Gene differential expression analysis
跳转:Gene differential expression analysis
snRNAseq-based differential expression analysis was assessed using two tests. First, a cell-level analysis was performed using the Wilcoxon-rank-sum test and FDR multiple testingcorrection. Second, a Poisson mixed model accounting for the individual of origin for nucleiand for unwanted sources of variability was performed using the R packages Ime4 and RUV-seq, respectively. Consistency of DEGs detected using the cell-level analysis model with thoseobtained with the Poisson mixed model was assessed by comparing DEGs directionality andrank for in the two models. Consistency in directionality for all cell-types was measured by counting what fraction of the top 1000 DEGs (ranked by FDR scores) detected in cell-levelanalysis show consistent direction in the mixed model. High consistency was found, with amedian fraction of 0.99 (Extended Data Table S2). Global consistency between the twomodels was assessed statistically using a resampling test. We tested whether the differentialp-value and z-score ranks corresponding to genes detected as up-regulated or down-regulated in the cell-level analysis was significantly higher or lower than expected by chancewhen computed using the mixed model. Expected scores were estimated by randomly sampling same-sized gene sets (n=1000, replicates). Both significance rank and direction deviate significantly from expectation, with directionality consistent for up- or down-regulated genes. Results of consistency tests are included in (Extended Data Table S2).Results from both cell-level and mixed-model differential tests are reported in Extended DataTable S2. For analyses involving DEG counts, only genes that are significantly supported byboth models using the criteria FDR-corrected P < 0.01 in two sided Wilcoxon-rank-sum test,absolute Log-fold > 0.25, and FDR-corrected P < 0.05 in Poisson mixed model are considered.Per and End populations were excluded from differential analyses due to their small cell counts.
Bulk RNA-seq differential analyses was performed by fitting a linear model using the Rpackage limma, while accounting for the covariates age, RNA integrity number, post-mortem interval, and plate batches. The pathological definition of AD groups as provided by ROSMAP was used. This classification defines AD or Non-AD groups based exclusively on theoverall pathological burden, without considering clinical diagnosis. We used the labels low-AD-pathology groups for consistency.
Consistency of gene expression perturbations in the different cell-types observed in snRNAdata with those detected in tissue-level bulk was assessed using a resampling approach. To test whether the genes identified as DEGs in single-cell data are also detected as high-ranking in the differential analysis in bulk, a z-score statistic was computed to quantify thedeviation of the observed differential (p-value) rank scores obtained in the bulk analysis forthe genes detected as DEGs in single-cells, relative to those observed in 1000 randomlychosen gene sets (Extended Data Table S2). This analysis was performed for each cell-typeindependently.
2.4 Source 9
文章跳转:Tau Accumulation Induces Microglial State Alterations in Alzheimer’s Disease Model Mice | eNeuro
2.4.1 Single-nucleus RNA-seq sequence data analysis
跳转:Single-nucleus RNA-seq sequence data analysis
Sequence data were analyzed using Seurat (https://satijalab.org/seurat/0) version 5.0.2(Hao et al, 2024) implemented in R version 4.3.2. First, filtered gene expression matrices were converted into Seurat objects using the CreateSeuratObject function. We selected cellswith >400 genes expressed and a mitochondrial gene ratio <1%. Possible doublet populations were identified by DoubletFinder (McGinnis et al, 2019), with an estimated doublet ratio of 20%, and were removed for the following analyses. The average numbers ofreads and genes per cell were 11,218 and 3,103, respectively. Data from each sequence were normalized using the SCTransform function (Hafemeister and Satija, 2019) before being integrated with the FindIntegrationAnchors and IntegrateData functions (Stuart et al., 2019)The principal components were calculated from the variable genes of the integrated Seuratobject. Clusters were identified using FindNeighbors and FindClusters, with a resolution of0.05. We generated two-dimensional projections using the Uniform Manifold Approximationand Projection (UMAP) algorithm based on the top 30 principal components. Cluster annotation was conducted by referring to the expression patterns of existing cell-typemarkers. Additional clustering analyses were then performed on Hexbt microglia, with aresolution of 0.2. Microglia clusters expressing other cell-type markers, such as St18 (foroligodendrocytes) and Rbfox3 (for neurons), were judged as doublets and eliminated fromthe subsequent analyses. Marker genes for each subcluster of microglia were searched using the FindAllMarkers function. Differential expression analysis was performed according to aprevious study (Zhan et al, 2020) using the FindMarkers function with default Wilcoxonrank-sum test and log fold-change threshold set at 0.5. Enrichment analyses were performed in Metascape (Zhou et al, 2019) using the identified marker genes. The single-cell RNA-seq and spatial transcriptomic datasets generated in the present study have been deposited inthe Gene Expression Omnibus (GEO) repository under accession number GSE270389.
3 常用方法参考文献
- DESeq2:Love et al., 2014, Genome Biology. Bioconductor DESeq2
- clusterProfiler:Yu et al., 2012, OMICS. Bioconductor clusterProfiler
- GSEA:Subramanian et al., 2005, PNAS. GSEA
- Seurat:Stuart et al., 2019, Cell. Seurat
- UMAP:McInnes et al. UMAP paper
References:
评论