This project performs full-length transcript isoform analysis using PacBio IsoSeq long-read RNA sequencing. Five human samples were sequenced to identify, classify, and quantify transcript isoforms at single-molecule resolution. The pipeline processes full-length non-chimeric (FLNC) reads through clustering, genome alignment, isoform collapsing, classification, ORF prediction, differential expression, and functional enrichment analysis.
| BioSample ID | Sample ID | Sample Name |
|---|---|---|
| BioSample_2 | Sample2 | Sample2 |
| BioSample_3 | Sample3 | Sample3 |
| BioSample_4 | Sample4 | Sample4 |
| BioSample_5 | Sample5 | Sample5 |
isoseq cluster2
Groups FLNC reads into transcript clusters; generates consensus sequences.
pbmm2 align
Aligns clustered reads to GRCh38 using the ISOSEQ preset.
isoseq collapse
Merges redundant isoforms; generates per-sample FLNC counts.
pigeon classify/filter
Annotates isoforms against GENCODE v39; removes artifacts.
TransDecoder
Identifies coding potential; flags NMD candidates.
edgeR
Pairwise differential expression with TMM normalization.
Enrichr API
GO, KEGG, WikiPathway, Reactome enrichment and visualization.
| Metric | Before Filter | After Filter |
|---|---|---|
| Total Transcripts | 139,730 | 86,219 |
| Unique Genes | 18,675 | 12,787 |
| Total FL Reads | 2,268,370 | 1,861,862 |
| Coding Transcripts | 118,393 | 81,194 |
| Non-coding Transcripts | 21,337 | 5,025 |
| Coding Percentage (%) | 84.7% | 94.2% |
| Mean ORF Length (aa) | 450.6 | 467.1 |
| Median ORF Length (aa) | 366.0 | 383.0 |
| Structural Category | Before Filter | % | After Filter | % |
|---|---|---|---|---|
| FSM full-splice_match | 45,164 | 32.3% | 36,746 | 42.6% |
| ISM incomplete-splice_match | 46,976 | 33.6% | 25,780 | 29.9% |
| NIC novel_in_catalog | 23,698 | 17.0% | 15,505 | 18.0% |
| NNC novel_not_in_catalog | 17,238 | 12.3% | 7,772 | 9.0% |
| Fusion | 420 | 0.3% | 241 | 0.3% |
| Other genic / antisense / intergenic | 6,234 | 4.5% | 175 | 0.2% |
| BioSample | Name | Transcripts | Genes | FL Reads | Coding | Non-coding | FSM | ISM | NIC | NNC | Fusion |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BioSample_2 | Sample2 | 52,990 | 11,169 | 437,110 | 50,208 | 2,782 | 25,749 | 14,717 | 8,594 | 3,734 | 103 |
| BioSample_3 | Sample3 | 55,264 | 10,930 | 458,711 | 52,449 | 2,815 | 26,069 | 15,397 | 9,247 | 4,341 | 140 |
| BioSample_4 | Sample4 | 59,336 | 11,423 | 558,138 | 56,186 | 3,150 | 28,429 | 15,759 | 10,146 | 4,790 | 138 |
| BioSample_5 | Sample5 | 44,652 | 10,299 | 407,097 | 42,246 | 2,406 | 22,276 | 13,285 | 6,110 | 2,849 | 67 |
ORF prediction was performed using TransDecoder. A transcript is classified as coding if its longest ORF is ≥300 bp (100 aa) OR if ORF coverage is ≥30%. The analysis also reports nonsense-mediated decay (NMD) candidates, indel counts from alignments, and per-isoform TPM values normalized for transcript length and sequencing depth.
| BioSample | Name | With ORF (Unfiltered) | With ORF (Filtered) | Mean ORF – Unfiltered (aa) | Mean ORF – Filtered (aa) |
|---|---|---|---|---|---|
| BioSample_2 | Sample2 | 74,101 | 50,208 | 461.8 | 479.4 |
| BioSample_3 | Sample3 | 74,918 | 52,449 | 473.9 | 489.7 |
| BioSample_4 | Sample4 | 79,850 | 56,186 | 467.3 | 480.5 |
| BioSample_5 | Sample5 | 54,810 | 42,246 | 428.6 | 432.6 |
Pairwise DE was performed with edgeR exactTest using fixed dispersion (BCV = 0.4), TMM normalization, minimum count ≥5, |log2FC| ≥1, and FDR <0.05 (Benjamini-Hochberg). All 10 pairwise comparisons among 5 samples are reported. Volcano and MA plots are provided per comparison.
| Comparison | Tested | Up-regulated | Down-regulated | No Change | % Up | % Down |
|---|---|---|---|---|---|---|
| Sample2 vs Sample3 | 16,881 | 83 | 20 | 16,778 | 0.5% | 0.1% |
| Sample2 vs Sample4 | 18,571 | 242 | 118 | 18,211 | 1.3% | 0.6% |
| Sample2 vs Sample5 | 16,363 | 428 | 533 | 15,402 | 2.6% | 3.3% |
| Sample3 vs Sample4 | 17,745 | 0 | 0 | 17,745 | 0.0% | 0.0% |
| Sample3 vs Sample5 | 16,214 | 81 | 164 | 15,969 | 0.5% | 1.0% |
| Sample4 vs Sample5 | 17,354 | 51 | 102 | 17,201 | 0.3% | 0.6% |
Enrichment analysis was performed via the Enrichr API on up-regulated, down-regulated, and all DE gene lists from each pairwise comparison. Four databases were queried and significant terms (adjusted p <0.05) were visualized using dotplots, barplots, gene-concept networks, and enrichment maps.
| Plot Type | Description |
|---|---|
| Dotplot | Top 20 enriched terms; dot size = gene count; dot color = −log10(adj. p-value) |
| Barplot | Top 15 terms; bar length = −log10(adj. p-value); dashed line at p = 0.05 |
| Gene-Concept Network | Bipartite network connecting pathway terms (blue) to member genes (orange) |
| Enrichment Map | Network of enriched terms; edges by Jaccard similarity ≥ 0.2; reveals functional clusters |
Showing: GO Biological Process 2023 – Up-regulated genes (Sample2 vs Sample3)
KEGG Pathway Visualization Example: KEGG Pathway: Hypertrophic Cardiomyopathy (hsa05410) – Sample2 vs Sample4 (all DE genes)
| File | Description |
|---|---|
04.pigeon/*_isoform_classification_complete.txt | Full isoform classification table – all 139,730 isoforms (54 columns) |
04.pigeon/*_isoform_classification.filtered_complete.txt | High-confidence filtered isoforms – 86,219 isoforms |
04.pigeon/classification_summary.txt | Overall and per-sample statistics summary |
05.differential_expression/DE_summary_statistics.txt | Summary of all 10 pairwise DE comparisons |
05.differential_expression/heatmap_top_DE_isoforms.pdf | Heatmap of top differentially expressed isoforms across all samples |
05.differential_expression/<comp>/DE_<comp>.txt | Full DE results per comparison (isoform, log2FC, p-value, FDR, regulation) |
05.differential_expression/<comp>/volcano_<comp>.pdf | Enhanced volcano plot with top 20 DE gene labels |
05.differential_expression/<comp>/MA_plot_<comp>.pdf | MA plot (mean expression vs fold change) |
05.differential_expression/<comp>/GO_*/enrichment_*.txt | GO enrichment results tables per comparison and direction |
05.differential_expression/<comp>/KEGG_*/ | KEGG enrichment results and pathway diagrams |
| Resource | Details |
|---|---|
| Reference Genome | GRCh38 – human_GRCh38_no_alt_analysis_set.fasta |
| Gene Annotation | GENCODE v39 – gencode.v39.annotation.sorted.gtf |
| Additional Resources | polyA.list.txt, refTSS BED, Intropolis TSV |