The UCSF Genomics CoLab provides bioinformatic support for data generated through our laboratory as well as selected externally generated datasets. Our goal is to move each project from raw instrument output to well-documented, interpretable results that address the investigator's biological question.
Analysis is scoped with the investigator before work begins. The exact workflow and deliverables depend on the assay, experimental design, data quality, available metadata, and level of biological interpretation requested.
Jump to analysis services by assay
Bulk RNA-sequencing
Small RNA sequencing
Whole-genome and whole-exome sequencing
Bulk ATAC-seq, CUT&RUN and CUT&Tag
Single-cell and single-nucleus RNA sequencing
Immune-receptor and Feature Barcode libraries
Single-nucleus ATAC and RNA-plus-ATAC Multiome
Visium HD spatial transcriptomics
Xenium spatial transcriptomics
Analysis-only and externally generated data
Levels of analysis support
Primary processing
Primary processing converts raw sequencing or imaging output into analysis-ready files. Depending on the assay, this may include read quality assessment, alignment, feature quantification, platform-specific processing, generation of count matrices, peak calling, spatial transcript assignment, and technical quality-control metrics.
Standard project analysis
Standard analysis evaluates the structure and quality of a dataset and performs the comparisons defined in the experimental design. It may include normalization, exploratory analysis, sample or cell clustering, differential testing, annotation, standard visualizations, and a summary discussion with the investigator.
Collaborative or custom analysis
Projects requiring extensive integration, advanced modeling, iterative biological interpretation, publication-oriented figures, or development of a specialized workflow are scoped separately. These projects require active scientific input from the investigator and may be best handled as a larger collaboration.
Bulk RNA sequencing
Primary processing and quality control
- Quality assessment of raw sequencing reads.
- Alignment to an appropriate reference genome and transcript annotation.
- Gene-level quantification and generation of a sample-by-gene count matrix.
- Assessment of sequencing depth, mapping, library complexity, strandedness, gene detection, and other assay-appropriate quality metrics.
- Identification of samples that may be affected by low quality, contamination, unexpected composition, or other technical concerns.
Standard analysis
- Count normalization and exploratory assessment of the dataset.
- Principal component analysis, sample correlation, and hierarchical clustering.
- Review of whether biological replicates group as expected and whether technical
covariates or batch effects are evident.
- Differential-expression testing for defined contrasts, including multifactorial
designs when supported by the sample size and experimental structure.
- MA plots, volcano plots, expression heat maps, and gene-level visualizations.
- Functional or pathway-enrichment analysis when included in the project scope.
- Genome-browser-compatible signal files when appropriate.
Typical deliverables
Deliverables may include read-quality summaries, alignment and quantification metrics, raw and normalized count tables, differential-expression results, standard figures, analysis-ready data objects, and a methods summary.
Small RNA and microRNA sequencing
Small-RNA data require processing that differs from conventional mRNA sequencing. The analysis is planned around the RNA classes of interest, library chemistry, species, and available annotations.
Primary processing and quality control
- Raw-read quality assessment and removal of library adapters.
- Assessment of read-length and insert-size distributions.
- Alignment and annotation against appropriate small-RNA reference collections.
- Generation of miRNA or other small-RNA count matrices, as supported by the organism and project design.
- Review of library composition and the fraction of reads assigned to the intended RNA classes.
Standard analysis
- Normalization, sample correlation, principal component analysis, and clustering.
- Differential-abundance testing for defined biological comparisons.
- Heat maps, MA plots, volcano plots, and visualizations of selected small RNAs.
- Integration with matched bulk mRNA results or target/pathway information when included in a custom analysis scope.
Typical deliverables
Deliverables may include trimmed reads, processing and composition metrics, annotated count tables, normalized abundance tables, differential results, figures, and a methods summary.
Whole-genome and whole-exome sequencing
DNA-sequencing analysis depends strongly on the organism, study design, desired variant classes, target coverage, and whether the project concerns germline, somatic, or another type of variation. These requirements must be defined before sequencing and analysis are finalized.
Primary processing and quality control
- Quality assessment of raw sequencing reads.
- Alignment to the selected reference genome.
- Generation of analysis-ready alignment files.
- Assessment of mapping, duplication, insert size, sequencing depth, and coverage uniformity.
- For exome projects, assessment of target capture and on-target coverage.
- Coverage summaries for the genome, exome, or other agreed regions of interest.
Project-dependent analysis
- Calling of single-nucleotide variants and small insertions or deletions.
- Variant-level quality filtering and annotation.
- Comparison with matched controls when required by the study design.
- Copy-number, structural-variant, low-frequency, or other specialized analyses when
the data and project scope support them.
Variant interpretation and specialized analyses are not assumed to be part of every WGS or WES project. The CoLab does not provide clinical diagnostic interpretation; research variant analysis must be defined and reviewed as part of the project scope.
Typical deliverables
Deliverables may include sequencing-QC reports, alignment files, coverage summaries, and, when included, filtered and annotated research variant files and summary figures.
Bulk ATAC-seq, CUT&RUN and CUT&Tag
ATAC-seq measures chromatin accessibility across a population of cells or nuclei. Analysis evaluates both the technical quality of the library and differences in accessible regulatory regions among biological conditions. CUT&RUN and CUT&Tag map genomic regions associated with a selected histone modification, transcription factor, or other chromatin-bound protein. Analysis must account for the target, expected peak shape, antibody performance, controls, and any
spike-in strategy used in the experiment.
Primary processing and quality control
- Raw-read quality assessment and alignment to the reference genome.
- Review of library complexity, duplicate reads, mitochondrial reads, fragment-size distribution, transcription-start-site enrichment, and signal-to-background.
- Peak calling and creation of a consensus set of accessible regions when supported by the project design.
- Generation of normalized signal tracks for genome-browser visualization.
Standard analysis
- Sample correlation, principal component analysis, and clustering using accessible regions.
- Identification of regions with differential accessibility between conditions.
- Annotation of peak regions to nearby genes and genomic features.
- Heat maps and aggregate signal plots across selected or differential regions.
- Motif-enrichment analysis when included in the project scope.
- Integration with matched gene-expression data as a collaborative analysis option.
Typical deliverables
Deliverables may include quality-control metrics, alignment files, peak or region files, normalized signal tracks, enrichment matrices, differential results, annotations, figures, and a methods summary.
Single-cell and single-nucleus RNA sequencing
This section covers 10x Genomics 3' and 5' Gene Expression, On-Chip Multiplexing (OCM), Flex fixed-sample Gene Expression, and compatible single-nucleus workflows. Primary processing and analysis are adapted to the chemistry and sample-barcoding strategy used in the experiment.
Primary processing and quality control
- Cellranger processing of sequencing reads and assignment of reads to cell or nucleus barcodes.
- Generation of gene-by-cell or gene-by-nucleus count matrices.
- For OCM and Flex, demultiplexing and review of sample-barcode assignment.
- Review of recovered cells or nuclei, reads per cell, genes detected, library saturation, sequencing depth, and other platform metrics.
- Assessment of low-quality barcodes, ambient RNA, multiplets, and sample-specific quality differences, as included in the agreed workflow.
Standard analysis
- Filtering, normalization, dimensionality reduction, unsupervised clustering and batch correction.
- Identification of cluster markers and visualization of selected genes.
- Cell-type or cell-state annotation using investigator knowledge, established markers, and suitable reference information.
- Comparison of cell composition and expression patterns among samples or groups.
- Differential-expression testing within relevant cell populations using a method appropriate for the number of biological replicates and experimental design.
- Assessment and correction of technical or batch effects when justified.
Collaborative analysis options
- Integration across experiments, batches, treatments, or sample cohorts.
- Reference mapping or transfer of annotations from a suitable external dataset.
- Cell-state scoring, pathway analysis, trajectory or lineage analyses, and other project-specific methods.
- Preparation of publication-oriented figures and iterative biological interpretation.
Typical deliverables
Deliverables may include cellranger reports, filtered count matrices, analysis-ready data objects, cell metadata, dimensionality-reduction coordinates, cluster and annotation assignments, marker and differential-expression tables, standard figures, and a methods summary.
Immune-receptor and Feature Barcode libraries
Compatible 10x projects may include T-cell or B-cell receptor libraries, cell-surface protein measurements, CRISPR Guide Capture, or other supported Feature Barcode libraries. These data are analyzed alongside the associated gene-expression data.
Primary processing and quality control
- TCR or BCR sequence reconstruction and clonotype assignment.
- Summary of clonotype frequency, diversity, sharing, and expansion.
- Processing of sequencing protein or feature barcodes reads and assignment of reads to cell or nucleus barcodes.
- Generation of TCR/BCR/Protein/FB-by-cell/nuclei count matrices.
Standard analysis
- Linking immune-receptor information to gene-expression clusters and cell annotations.
- Processing and normalization of compatible antibody-derived tag or other feature
count matrices.
- Visualization of protein or feature measurements alongside RNA expression.
- Assignment and evaluation of captured CRISPR guides when supported by the
experimental design.
Typical deliverables
Exact deliverables depend on the library type and the quality and completeness of the paired gene-expression dataset.
Single-nucleus ATAC and RNA-plus-ATAC Multiome
Single-nucleus ATAC resolves chromatin accessibility by nucleus. Multiome measures gene expression and chromatin accessibility from the same nucleus, allowing the two modalities to be analyzed jointly.
Primary processing and quality control
- Cellranger processing of alignment, barcode assignment, and generation of feature matrices.
- Assessment of recovered nuclei, fragment counts, transcription-start-site enrichment, nucleosomal patterning, signal-to-background, and other assay metrics.
- Generation of accessible-region matrices and peak sets.
- For Multiome, generation and review of both RNA and chromatin-accessibility matrices for the same nuclei.
Standard and collaborative analysis
- Filtering, dimensionality reduction, clustering, and marker-region analysis.
- Cell-type or cell-state annotation using accessibility, gene activity, RNA
expression, and suitable reference information.
- Differential-accessibility analysis between biological conditions or cell types.
- Motif and transcription-factor activity analysis when included in scope.
- Joint RNA and accessibility visualization and integrated clustering for Multiome.
- Project-specific analysis of relationships between regulatory regions and gene
expression when the data and study design support it.
Typical deliverables
Deliverables may include processing reports, peak and feature matrices, analysis-ready data objects, nucleus metadata, clusters and annotations, differential-accessibility results, motif summaries, integrated Multiome results, figures, and a methods summary.
Visium HD spatial transcriptomics
Visium HD analysis combines sequencing-based expression measurements with tissue
images and spatial coordinates. The appropriate resolution and interpretation depend
on tissue morphology, image quality, chemistry, sequencing depth, and the chosen
binning or segmentation strategy.
Primary processing and quality control
- Spaceranger processing of sequencing data and spatial images.
- Alignment and generation of spatial gene-expression matrices.
- Review of tissue coverage, reads and genes detected, sequencing saturation, image alignment, and other platform quality metrics.
- Generation of expression outputs at an agreed bin size or using an agreed image-guided segmentation approach.
Standard and collaborative analysis
- Visualization of gene expression in tissue coordinates.
- Dimensionality reduction and clustering of spatial expression profiles.
- Annotation of tissue regions, spatial domains, or putative cell types using morphology, marker genes, and suitable references.
- Identification and visualization of spatially patterned genes.
- Comparison of regions, samples, or experimental groups when biological replication supports the comparison.
- Integration with matched histology, single-cell or single-nucleus reference data, or other molecular measurements as a custom analysis option.
Typical deliverables
Deliverables may include platform processing reports, spatial expression matrices, processed images, coordinates or segmentation outputs, analysis-ready data objects, spatial clusters and annotations, gene-expression maps, comparison tables, figures, and a methods summary.
Xenium spatial transcriptomics
Xenium produces targeted, spatially resolved transcript measurements together with tissue images and cell-segmentation results. Analysis focuses on transcript-detection quality, segmentation, cell-level expression, spatial organization, and the biological questions supported by the selected gene panel.
Primary processing and quality control
- Review of platform processing and run-level quality metrics.
- Assessment of transcript counts, negative-control measurements, tissue coverage, image quality, and cell-segmentation results.
Standard and collaborative analysis
- Visualization of transcripts and gene expression in tissue coordinates.
- Cell clustering and annotation using the gene panel, morphology, and suitable
reference information.
- Comparison of cell types, cell states, or regions among samples and conditions.
- Spatial-neighborhood, cell-proximity, tissue-domain, or region-of-interest analyses when supported by the panel and project design.
- Integration with Visium, single-cell, single-nucleus, histologic, or other data as a custom analysis option.
Typical deliverables
Deliverables may include platform processing reports, transcript coordinates, cell-segmentation outputs, cell-by-gene matrices, analysis-ready data objects, cell annotations, spatial maps, neighborhood or region summaries, figures, and a methods summary.
Analysis-only and externally generated data
The CoLab may accept data generated outside our laboratory after reviewing its format, completeness, quality, provenance, and compatibility with the requested analysis. Analysis-only projects should include:
- Raw or processed data in an agreed format.
- Complete sample metadata and a sample-name key.
- Experimental groups, biological replicates, covariates, and batch information.
- Organism, reference genome, annotation, and relevant panel or chemistry details.
- A description of prior processing and any available quality-control reports.
- The biological questions, planned comparisons, and desired deliverables.
- Any applicable data-access, security, transfer, or deadline requirements.
Incomplete metadata or a study design in which biology is confounded with processing batch can substantially limit the conclusions that can be drawn.
Data delivery and project review
The files delivered depend on the assay and agreed analysis scope. Projects may include raw data, processed matrices, alignment or peak files, analysis-ready data objects, statistical result tables, annotations, visualizations, and a methods summary. Raw, intermediate, and long-term-storage responsibilities should be agreed upon at project initiation.
When analysis is included, we normally schedule a review with the investigator to discuss data quality, major findings, limitations, and appropriate next steps. More extensive interpretation or additional analyses can then be scoped separately.
Start an analysis project
To request bioinformatic support, provide the assay or data type, number of samples, organism, experimental design, available data formats, requested comparisons, desired outputs, and any grant, manuscript, or project deadline.