Admera Health 126 Corporate Boulevard, South Plainfield, NJ 07080
+1-908-222-0533 • custom-services@admerahealth.comwww.admerahealth.com
CLIA ID: 31D2038676

PacBio IsoSeq Long-Read RNA Sequencing Analysis

Project IDdemo-project
Analysis DateJanuary – February 2026
ReferenceGRCh38 / GENCODE v39
Samples4

1. Project Overview

This project performs full-length transcript isoform analysis using PacBio IsoSeq long-read RNA sequencing. Five human samples were sequenced to identify, classify, and quantify transcript isoforms at single-molecule resolution. The pipeline processes full-length non-chimeric (FLNC) reads through clustering, genome alignment, isoform collapsing, classification, ORF prediction, differential expression, and functional enrichment analysis.

BioSample IDSample IDSample Name
BioSample_2Sample2Sample2
BioSample_3Sample3Sample3
BioSample_4Sample4Sample4
BioSample_5Sample5Sample5

2. Analysis Pipeline

1
Cluster

isoseq cluster2
Groups FLNC reads into transcript clusters; generates consensus sequences.

2
Map

pbmm2 align
Aligns clustered reads to GRCh38 using the ISOSEQ preset.

3
Collapse

isoseq collapse
Merges redundant isoforms; generates per-sample FLNC counts.

4
Classify

pigeon classify/filter
Annotates isoforms against GENCODE v39; removes artifacts.

5
ORF Prediction

TransDecoder
Identifies coding potential; flags NMD candidates.

6
DE Analysis

edgeR
Pairwise differential expression with TMM normalization.

7
Enrichment

Enrichr API
GO, KEGG, WikiPathway, Reactome enrichment and visualization.

3. Overall Isoform Statistics

Total Transcripts
139,730
(Before Filter)
Transcripts
86,219
(After Filter)
Unique Genes
12,787
(After Filter)
Total FL Reads
1,861,862
(After Filter)
Coding Transcripts
94.2%
(After Filter)
Mean ORF Length
467 aa
(After Filter)
MetricBefore FilterAfter Filter
Total Transcripts 139,730 86,219
Unique Genes 18,675 12,787
Total FL Reads 2,268,370 1,861,862
Coding Transcripts 118,393 81,194
Non-coding Transcripts 21,337 5,025
Coding Percentage (%) 84.7% 94.2%
Mean ORF Length (aa) 450.6 467.1
Median ORF Length (aa) 366.0 383.0

4. Structural Categories

Distribution After Filter (86,219 transcripts)

Before vs After Filter by Category

Structural Category Before Filter% After Filter%
FSM full-splice_match 45,16432.3% 36,74642.6%
ISM incomplete-splice_match 46,97633.6% 25,78029.9%
NIC novel_in_catalog 23,69817.0% 15,50518.0%
NNC novel_not_in_catalog 17,23812.3% 7,7729.0%
Fusion 4200.3% 2410.3%
Other genic / antisense / intergenic 6,2344.5% 1750.2%

5. Per-Sample Statistics (After Filter)

BioSampleName TranscriptsGenes FL ReadsCodingNon-coding FSMISMNIC NNCFusion
BioSample_2Sample2 52,990 11,169 437,110 50,208 2,782 25,749 14,717 8,594 3,734 103
BioSample_3Sample3 55,264 10,930 458,711 52,449 2,815 26,069 15,397 9,247 4,341 140
BioSample_4Sample4 59,336 11,423 558,138 56,186 3,150 28,429 15,759 10,1464,790 138
BioSample_5Sample5 44,652 10,299 407,097 42,246 2,406 22,276 13,285 6,110 2,849 67

FL Reads per Sample (After Filter)

Structural Category Distribution per Sample (After Filter)

6. ORF Prediction & Coding Potential

ORF prediction was performed using TransDecoder. A transcript is classified as coding if its longest ORF is ≥300 bp (100 aa) OR if ORF coverage is ≥30%. The analysis also reports nonsense-mediated decay (NMD) candidates, indel counts from alignments, and per-isoform TPM values normalized for transcript length and sequencing depth.

BioSampleName With ORF (Unfiltered)With ORF (Filtered) Mean ORF – Unfiltered (aa)Mean ORF – Filtered (aa)
BioSample_2Sample2 74,101 50,208 461.8479.4
BioSample_3Sample3 74,918 52,449 473.9489.7
BioSample_4Sample4 79,850 56,186 467.3480.5
BioSample_5Sample5 54,810 42,246 428.6432.6

7. Differential Expression Analysis

Pairwise DE was performed with edgeR exactTest using fixed dispersion (BCV = 0.4), TMM normalization, minimum count ≥5, |log2FC| ≥1, and FDR <0.05 (Benjamini-Hochberg). All 10 pairwise comparisons among 5 samples are reported. Volcano and MA plots are provided per comparison.

Comparison TestedUp-regulated Down-regulatedNo Change % Up% Down
Sample2 vs Sample3 16,88183 20 16,778 0.5% 0.1%
Sample2 vs Sample4 18,571242 118 18,211 1.3% 0.6%
Sample2 vs Sample5 16,363428 533 15,402 2.6% 3.3%
Sample3 vs Sample4 17,7450 0 17,745 0.0% 0.0%
Sample3 vs Sample5 16,21481 164 15,969 0.5% 1.0%
Sample4 vs Sample5 17,35451 102 17,201 0.3% 0.6%
Volcano Plot – Sample2 vs Sample3
Volcano Plot – Sample2 vs Sample3
MA Plot – Sample2 vs Sample3
MA Plot – Sample2 vs Sample3

8. Functional Enrichment Analysis

Enrichment analysis was performed via the Enrichr API on up-regulated, down-regulated, and all DE gene lists from each pairwise comparison. Four databases were queried and significant terms (adjusted p <0.05) were visualized using dotplots, barplots, gene-concept networks, and enrichment maps.

GO Biological Process 2023 GO Molecular Function 2023 GO Cellular Component 2023 KEGG 2021 Human WikiPathway 2023 Human Reactome 2022
Plot TypeDescription
Dotplot Top 20 enriched terms; dot size = gene count; dot color = −log10(adj. p-value)
Barplot Top 15 terms; bar length = −log10(adj. p-value); dashed line at p = 0.05
Gene-Concept NetworkBipartite network connecting pathway terms (blue) to member genes (orange)
Enrichment Map Network of enriched terms; edges by Jaccard similarity ≥ 0.2; reveals functional clusters
Example top enriched GO Biological Process terms for Sample2 vs Sample3 (up-regulated): Myofibril Assembly, Cardiac Muscle Tissue Morphogenesis, Heart Contraction, Actomyosin Structure Organization. KEGG pathway visualization (pathview) colors DE genes by log2 fold change on official pathway maps.

Showing: GO Biological Process 2023 – Up-regulated genes (Sample2 vs Sample3)

Dotplot
Dotplot
Barplot
Barplot
Gene-Concept Network
Gene-Concept Network
Enrichment Map
Enrichment Map

KEGG Pathway Visualization Example: KEGG Pathway: Hypertrophic Cardiomyopathy (hsa05410) – Sample2 vs Sample4 (all DE genes)

Hypertrophic Cardiomyopathy (hsa05410)
Hypertrophic Cardiomyopathy (hsa05410)

9. Key Output Files

FileDescription
04.pigeon/*_isoform_classification_complete.txt Full isoform classification table – all 139,730 isoforms (54 columns)
04.pigeon/*_isoform_classification.filtered_complete.txt High-confidence filtered isoforms – 86,219 isoforms
04.pigeon/classification_summary.txt Overall and per-sample statistics summary
05.differential_expression/DE_summary_statistics.txt Summary of all 10 pairwise DE comparisons
05.differential_expression/heatmap_top_DE_isoforms.pdf Heatmap of top differentially expressed isoforms across all samples
05.differential_expression/<comp>/DE_<comp>.txt Full DE results per comparison (isoform, log2FC, p-value, FDR, regulation)
05.differential_expression/<comp>/volcano_<comp>.pdf Enhanced volcano plot with top 20 DE gene labels
05.differential_expression/<comp>/MA_plot_<comp>.pdf MA plot (mean expression vs fold change)
05.differential_expression/<comp>/GO_*/enrichment_*.txt GO enrichment results tables per comparison and direction
05.differential_expression/<comp>/KEGG_*/ KEGG enrichment results and pathway diagrams

10. Software & Methods

IsoSeq Pipeline
  • isoseq cluster2 – transcript clustering
  • pbmm2 align – genome alignment (ISOSEQ preset)
  • isoseq collapse – isoform collapsing
  • pigeon v1.0.0 – classification, filtering, reporting
  • samtools / pbindex – BAM utilities
ORF Prediction
  • TransDecoder – LongOrfs + Predict (≥100 aa, --single_best_only)
  • Coding criteria: ORF ≥ 300 bp OR coverage ≥ 30%
Differential Expression
  • R 4.4.0 – statistical computing
  • edgeR – exactTest, TMM normalization
  • ggplot2 + ggrepel – volcano and MA plots
Enrichment & Visualization
  • Enrichr API – GO, KEGG, WikiPathway, Reactome
  • igraph (R) – network & enrichment map
  • pathview (R) – KEGG pathway diagrams
  • KEGGREST (R) – KEGG database access

Reference Genome & Annotation

ResourceDetails
Reference GenomeGRCh38 – human_GRCh38_no_alt_analysis_set.fasta
Gene AnnotationGENCODE v39 – gencode.v39.annotation.sorted.gtf
Additional ResourcespolyA.list.txt, refTSS BED, Intropolis TSV