From short-read reconstruction to full-length molecules: what actually changed
The baseline that long-read transcriptomics displaces is computational assembly of short fragments. Short-read sequencing generates tens of billions of highly accurate reads, but reconstructing full-length RNA structures from short fragments requires assembly that frequently fails to resolve splice variants, overlapping transcripts, alternative transcript ends, and structures in repetitive or low-complexity regions [1]. The anchor review states this limitation as a structural property of the read length rather than a fixable pipeline problem, and illustrates it in two settings where full-length information matters most: alternatively spliced human genes and compact viral genomes with extensive transcriptional overlap [1]. The foundational work behind this framing established that alternative splicing and alternative transcription start and end site usage generate isoforms from the same gene, and that disruption of these processes is associated with hereditary disease and cancer [2]. That earlier work also identified fragmentation bias as an inherent property of short-read RNA-seq, which is the specific technical reason isoform reconstruction is difficult rather than merely underpowered [2].
The anchor review's contribution is to organize the field around what full-length reads actually resolve: transcription start-site choice, splicing, transcription end-site selection, and polyadenylation coordinated within the same RNA molecule [1]. This is a conceptual shift from detecting expression to defining transcript architecture. It also reframes the analytical problem, because boundary fidelity and transcript-model confidence become the quantities that determine biological and clinical impact, not read counts alone [1].
Platform chemistry sets the error and boundary ceiling before any analysis begins
PacBio and ONT are not interchangeable inputs. PacBio HiFi reads are generated by circular consensus sequencing and often exceed 99.9% accuracy at read lengths commonly in the 10–25 kb range, while ONT reads have historically shown median read accuracy of 87–92% for legacy direct RNA kits, improving to approximately 95% or above with the SQK-RNA004 chemistry and current basecalling models [1]. The anchor review is explicit that residual errors persist in both chemistries, with PacBio errors concentrated in homopolymers and low-complexity regions and cDNA protocols susceptible to reverse-transcription and PCR artifacts, while ONT retains context-dependent deletions [1]. This matters because splice-site definition and RNA modification detection are directly sensitive to those error modes, and reference-guided alignment only partly compensates in well-annotated genomes [1].
The precursor benchmarking evidence agrees on the direction of this trade-off. ISAtools was developed specifically for PacBio CCS data, and its authors state that benchmarking efforts from LRGASP consistently showed PacBio CCS data achieve superior isoform reconstruction accuracy compared with ONT data, largely owing to lower error rates and more reliable splice junction detection [3]. That is a convergent conclusion reached from a different starting point, which strengthens the platform ordering for splice-structure reconstruction. The anchor review extends it by adding the boundary dimension: direct RNA sequencing preserves native RNA features and typically supports robust 3′ end recovery, but incomplete 5′ coverage can reduce transcription start-site fidelity, whereas cDNA workflows improve full-length recovery at the cost of internal priming and template switching [1]. The practical reading is that platform choice should follow the analytical question, with PacBio HiFi/Kinnex usually best suited for high-confidence splice-structure reconstruction and ONT data requiring stronger attention to basecalling model, chemistry, alignment strategy, and validation [1].
Depth, not read length, is the binding constraint for quantification
The anchor review draws a sharp asymmetry: read length and accuracy primarily determine isoform reconstruction quality, whereas sequencing depth principally governs quantification precision, so reliable quantification of rare isoforms requires substantially higher depth than isoform discovery alone [1]. This is the single most consequential statement for clinical readers, because it separates two questions that are often conflated. A transcript model can be structurally well supported by a handful of full-length reads while its expression level remains statistically unusable. The review notes that neither PacBio nor ONT currently achieves the per-sample read depth of short-read RNA-seq at comparable cost, which limits statistical power for transcript-level differential expression, especially for low-abundance transcripts [1].
The limitation evidence from LongBench sharpens this into a protocol-level warning. Across bulk, single-cell, and single-nucleus samples from eight lung cancer cell lines profiled with seven sequencing protocols across three platforms, all protocols showed specific length and quantitation biases [6]. PacBio Kinnex's improved throughput enabled differential analysis with highly accurate reads, but transcript lengths were constrained at both ends of the spectrum, particularly with underrepresentation of transcripts below 1 kb in the standard bulk protocol; ONT cDNA was less costly but yielded shorter and less accurate reads with reduced sensitivity for gene, transcript, splice junction detection, and differential expression; ONT direct RNA had the lowest bias of the bulk methods but requires relatively large RNA inputs and has lower accuracy [6]. Critically, LongBench found that the three long-read protocols agreed on the major isoform for only 67–69% of multi-isoform genes, while short-read results were notably more divergent [6]. That agreement figure is the honest ceiling on cross-platform isoform concordance and should temper any claim that a single-platform isoform ratio is a settled measurement.
Clinical interpretation gains are real but narrower than the technology's promise
The strongest clinical case in the anchor review comes from recent long-read RNA-seq studies in rare disorders. PacBio Kinnex/HiFi RNA-seq resolved clinically relevant intron retention, multiple exon skipping, leaky splicing, variant phasing, and isoform switching, while targeted strategies such as STRIPE/TEQUILA increased depth across disease-gene panels and supported haplotype-resolved interpretation of variants of uncertain significance and molecular diagnosis in previously undiagnosed individuals [1]. The scale of these experiments is worth stating precisely: the Kinnex/HiFi study generated approximately 13.2 million full-length non-chimeric reads per blood library and approximately 11.7 million per fibroblast sample, whereas STRIPE/TEQUILA achieved median on-target rates of 64.2% and 87.1%, corresponding to 71.8- and 39.5-fold enrichment over untargeted whole-transcriptome long-read sequencing in the two tested panels [1]. Those enrichment factors are the mechanism by which targeted long-read RNA-seq reaches the depth that untargeted whole-transcriptome work cannot.
The validation evidence sets the boundary. In a well-characterized rare disease cohort of 1,462 families, short-read genome sequencing identified diagnostic structural variants in 5.4% of families, and 98% of those diagnostic structural variants were located in genomic regions readily accessible to short-read genome sequencing, with only three clinically interpretable variants localized to highly repetitive segmental duplications or simple repeat sequences [5]. The authors conclude that the current relative incremental improvement in diagnostic yield from long-read genome sequencing is modest when a comprehensive short-read analysis is performed [5]. This is a different endpoint from transcript-level isoform interpretation, so it does not refute the anchor review's clinical claims, but it does constrain how much of the undiagnosed rare-disease burden long-read approaches are likely to absorb through variant discovery alone. The anchor review's own evidence boundary applies here: it is a methodological review and provides no new clinical diagnostic efficacy data, and its platform recommendations will shift as chemistry and software versions change [1].
Viral, single-cell, and spatial applications define where the method still stops
Compact viral genomes function as a stress test because transcript-unit resolution, not expression quantification, determines biological interpretation. The anchor review documents that across multiple herpesviruses, long-read sequencing uncovered previously unrecognized mRNAs and noncoding RNAs, alternative transcription start and end site isoforms, splice variants, polygenic RNAs, and extensive networks of transcriptional overlap, with similar complexity observed in poxviruses, baculoviruses, African swine fever virus, and SARS-CoV-2 [1]. In hepatitis B virus, targeted long-read sequencing of patient liver biopsies resolved cDNA-derived and integrant-derived transcripts and identified previously uncharacterized spliced, truncated, and chimeric viral RNAs [1]. Competing evidence from bacterial and viral work shows the same platform trade-offs operating at smaller scale: in Escherichia coli, direct RNA sequencing, direct cDNA, and PCR-cDNA protocols were compared, and the study concluded that nanopore RNA-seq is suitable for quantitative measurement and correlates well with Illumina short-read data, while noting that direct RNA sequencing requires more than 10 µg of starting RNA to yield enough mRNA after rRNA depletion and lacks straightforward barcoding [4]. In Senecavirus A, PCR-cDNA produced much higher viral yield and lower raw error rates than direct RNA sequencing, but direct RNA sequencing showed a strong linear relationship between input viral copies and output reads (r² = 0.99) whereas PCR-cDNA did not (r² = 0.54), indicating that direct RNA sequencing was quantitative while PCR-cDNA was not [7]. Those two results pull in opposite directions on accuracy versus quantitation, which is exactly the trade-off the anchor review describes as platform-specific rather than resolvable by a single best protocol [1].
Single-cell and spatial long-read work is conceptually feasible but depth-limited. The anchor review states that the main barrier is no longer conceptual feasibility but joint optimization of throughput, barcode fidelity, sequencing accuracy, and isoform-level statistical analysis, and warns that because reads are distributed across many cells, rare-isoform quantification is often limited by depth even when individual molecules are structurally informative [1]. The foundational single-cell methodology work identifies the same constraint from the library-preparation side, noting that droplet-based platforms can capture 10,000 cells per library but suffer high dropout rates especially for low-expression genes, and that analysis tools are often tailored to a specific experiment type and may fail on customized barcodes or non-standard flanking sequences [2]. The anchor review also notes that modification calling from direct RNA sequencing remains strongly dependent on modification type, basecalling model, signal model, and sequencing depth, and should be interpreted with caution [1]. The open question is not whether these applications work in principle but what depth threshold makes a rare-isoform call in a single cell defensible, and the supplied evidence does not yet supply that number.
About These Sources
This research page is built on 7 peer-reviewed studies — published from 2020 to 2026, 5 from 2024 or later — selected as the most relevant from 8 studies that passed quality screening, drawn from 61 papers retrieved from a database of over 500 million.
Sources used in this answer
Long-read transcriptomics - opportunities and challenges
The anchor review by Tombácz and colleagues synthesizes platform-specific trade-offs across PacBio HiFi/Kinnex cDNA, ONT cDNA, and direct RNA sequencing, and argues that read length and accuracy govern isoform reconstruction while sequencing depth governs quantification precision, making platform choice and depth calibration preconditions for defensible transcript models in rare-disease, viral, and single-cell applications [1].
Computational methodology for analysis of cell type specific RNA isoform expression with long read sequencing
This foundational methodology work establishes that alternative splicing and alternative transcription start and end site usage generate isoforms from the same gene, that disruption of these processes is associated with hereditary disease and cancer, and that fragmentation bias is inherent to short-read RNA-seq, while also documenting that droplet-based single-cell platforms suffer high dropout rates and that analysis tools often fail on customized barcodes [2].
Efficient reconstruction of full-length RNA isoforms using ISAtools and large-scale PacBio circular consensus sequencing data
ISAtools was developed as a scalable reference-free framework for full-length isoform reconstruction from PacBio CCS data, and its authors report that LRGASP benchmarking consistently showed PacBio CCS data achieve superior isoform reconstruction accuracy compared with ONT data, with ISAtools completing reconstruction and quantification of approximately 1 billion simulated reads in about 4 hours using about 8 GB of memory while IsoQuant, Bambu, and Mandalorion failed at a 300 GB memory limit [3].
Nanopore sequencing of RNA and cDNA molecules in Escherichia coli.
This competing evidence compares ONT direct RNA, direct cDNA, and PCR-cDNA protocols in Escherichia coli and concludes that nanopore RNA-seq is suitable for quantitative measurement and correlates well with Illumina short-read data, while noting that direct RNA sequencing requires more than 10 µg of starting RNA to yield enough mRNA after rRNA depletion and lacks straightforward barcoding [4].
Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing
This validation study of 1,462 rare disease families found diagnostic structural variants by short-read genome sequencing in 5.4% of families, with 98% of those variants in regions readily accessible to short-read sequencing and only three clinically interpretable variants in highly repetitive regions, concluding that the current incremental diagnostic yield from long-read genome sequencing is modest when comprehensive short-read analysis is performed [5].
Benchmarking long-read RNA-sequencing technologies with LongBench: a cross-platform reference dataset profiling cancer cell lines with bulk and single-cell …
LongBench profiled eight lung cancer cell lines across bulk, single-cell, and single-nucleus modalities using seven sequencing protocols across three platforms and found that all protocols show specific length and quantitation biases, that the three long-read protocols agreed on the major isoform for only 67–69% of multi-isoform genes, and that PacBio Kinnex underrepresents transcripts below 1 kb in the standard bulk protocol [6].
Characterization of Emerging Swine Viral Diseases through Oxford Nanopore Sequencing Using Senecavirus A as a Model.
Using Senecavirus A as a model, this competing study found that PCR-cDNA sequencing produced much higher viral yield and lower raw error rates than direct RNA sequencing, but direct RNA sequencing showed a strong linear relationship between input viral copies and output reads (r² = 0.99) whereas PCR-cDNA did not (r² = 0.54), indicating that direct RNA sequencing was quantitative while PCR-cDNA was not [7].
