Why penaeid shrimp genomes stayed fragmented for two decades
The difficulty is not sequencing cost but genome architecture. Yuan et al. surveyed four penaeid species and found bimodal k-mer depth distributions in all of them, which they traced to abundant homo-duplicated repeats rather than whole-genome duplication, since Ks and allele-frequency analyses and a single Hox cluster gave no support for a duplication event [8]. They estimated heterozygosity at 2.43% in L. vannamei, 1.95% in Fenneropenaeus chinensis, 4.95% in Penaeus monodon, and 4.49% in Marsupenaeus japonicus [8]. High repeat content plus high heterozygosity produces exactly the collapsed or fragmented contigs that plagued early assemblies.
The black tiger shrimp case shows the same constraint from a different angle. Uengwetwanit et al. noted that available P. monodon assemblies lacked contiguity and completeness because of the highly repetitive genome and the technical difficulty of extracting high-quality, high-molecular-weight DNA [6]. Their own PacBio-only draft reached a contig N50 of just 79 kb before Chicago and Hi-C scaffolding raised the scaffold N50 to 44.86 Mb [6]. In other words, long reads alone did not solve penaeid genomes; scaffolding and careful DNA handling did much of the work.
Comparative work in other crustaceans reinforces that contiguity is achievable when repeat content is lower or the assembly strategy is tuned. Baldwin-Brown et al. produced a 120 Mb clam shrimp assembly with a contig N50 of 18 Mb from PacBio and Illumina data, describing it as the most contiguous crustacean assembly of its time [3]. That genome is roughly sixteen times smaller than L. vannamei's, so the comparison is about method, not scale: long reads that span repeats can resolve them when the repeat landscape is tractable.
What the new assembly actually changed, and what it did not
The anchor paper's central move is strategic rather than algorithmic. Ma et al. generated 170.44 Gb of PacBio HiFi data (reads N50 15,986 bp), 143.25 Gb of ONT ultra-long data (reads N50 100 kb), and 336.69 Gb of Hi-C clean data, then tested three contig strategies before scaffolding [1]. The HiFi-only assembly gave 1.22 Gb with a 41,821 bp contig N50; a HiFi-plus-ONT hybrid gave 0.93 Gb with a 67,583 bp N50; and an ONT-only assembly polished with HiFi gave 1.99 Gb with a 454,920 bp N50 [1]. They carried the ONT-based route forward, then applied three rounds of gap filling and Hi-C scaffolding to reach 1.95 Gb, a 1.97 Mb contig N50, a 46.08 Mb scaffold N50, and 97.25% anchoring across 44 pseudochromosomes [1].
That is a real jump over the prior frontier. The 2019 L. vannamei assembly had a 57.65 kb contig N50 and 605.56 kb scaffold N50 [1]; the 2025 assembly reached 201.33 kb contig N50 and 42.33 Mb scaffold N50 with a BUSCO completeness of 84.6% [1]. The new contig N50 is roughly ten times the 2025 value, and the scaffold N50 is slightly higher, though the BUSCO figure is lower, which the authors do not reconcile [1].
The annotation adds 26,830 protein-coding genes with 23,952 (89.27%) functionally annotated, an average gene length of 15,514 bp, and about 6.15 exons per gene [1]. Repetitive sequences total 1.195 Gb, or 71.95% of the assembly, with interspersed repeats alone at 61.28% and DNA transposons the largest class at 43.54% [1]. For breeding-oriented readers, the practical gain is that 97.25% of the genome now sits on chromosomes, which makes linkage, selection-scan, and synteny analyses far more tractable than on a contig-level draft.
Platform choice is not neutral: what the HiFi-versus-ONT comparison shows
The internal comparison in the anchor paper is unusually informative because the same sample was pushed through three pipelines. The ONT-only route produced both the largest (1.99 Gb) and by far the most contiguous (454,920 bp N50) contig set, while the HiFi-only route produced the smallest (1.22 Gb) and least contiguous (41,821 bp N50) [1]. The authors note the ONT assembly was substantially larger and more contiguous than a recently published 1.7 Gb P. vannamei genome [1]. This is consistent with broader evaluations: Rayamajhi et al. assembled the bald notothen across short-read, hybrid, and long-read-only phases and concluded that long-read-only contig assembly was the best choice, with phase I and phase II assemblies of lower quality [7].
But platform comparisons are context-dependent. Udaondo et al. compared PacBio and ONT for transcriptomic analysis in P. monodon and found that read-length bias and throughput differences strongly influenced results, concluding that platform choice should follow the research question [9]. The anchor paper's own numbers support a nuanced reading: HiFi reads mapped back at 97.14% with 67.89% of the genome at 5x or greater coverage, while ONT reads mapped at 100.00% with 98.58% at 5x or greater coverage [1]. The ONT advantage here is partly coverage depth, not only read length.
The gap-filling sequence also matters for interpretation. Ma et al. used HiFi-based LR_GapCloser, then quarTeT's gapfiller with contigs from both the HiFi and hybrid assemblies, then ONT-based LR_GapCloser, followed by Racon polishing [1]. Because the final contig N50 of 1.97 Mb is described as a substantial increase from the 454.92 kb achieved after three rounds of gap-filling, the largest single gain appears to come from scaffolding and gap closure rather than from the initial contig assembly [1].
The completeness ceiling: what 70% BUSCO means for downstream claims
The most important boundary is stated plainly in the paper. BUSCO assessment against eukaryota_odb10 (255 genes) returned 181 complete BUSCOs (70.98%), comprising 123 single-copy (48.24%) and 58 duplicated (22.75%), with 4 fragmented (1.57%) and 70 missing (27.45%) [1]. The proteome mode was similar: 178 complete (69.80%), 126 single-copy (49.41%), 52 duplicated (20.39%), 13 fragmented (5.10%), and 64 missing (25.10%) [1]. Roughly a quarter of conserved single-copy orthologs are absent from the assembly.
That figure sits below the 84.6% BUSCO completeness reported for the 2025 L. vannamei assembly [1], which is a genuine tension the paper does not resolve. Possible explanations supported by the supplied material include differences in lineage dataset, assembly size, and repeat masking, but the sources do not let us adjudicate among them. What can be said is that a 70% BUSCO score is not evidence of a poor assembly in a repeat-rich genome; it is evidence that a substantial fraction of conserved genes fall in regions the assembly does not capture.
Tool choice also affects how completeness is measured. Huang and Li showed that compleasm, which uses miniprot instead of BUSCO's two-round MetaEuk approach, reported 99.6% completeness for T2T-CHM13 versus BUSCO's 95.7%, and that 557 of 562 compleasm-specific complete genes were supported by NCBI annotation [5]. They also caution that compleasm has limited sensitivity to distant homologs and recommend combining it with BUSCO for divergent assemblies [5]. For a crustacean genome with 71.95% repeats [1], that caveat is directly relevant: the reported 70.98% may understate true completeness, but no compleasm run is reported in the anchor paper.
The practical consequence for aquaculture genetics is that the new assembly is a strong scaffold for chromosome-level analyses but not yet a finished reference. Whole-genome resequencing work in L. vannamei has already identified 37 million SNPs at an average density of 22.5 SNPs/Kb across 180 shrimp from six breeds, with selection signatures enriched for macromolecule metabolism, proteolysis, structural molecule activity, and stimulus response [2]. Those analyses benefit from better anchoring, but any selection scan or GWAS that depends on gene models in the missing 27.45% will inherit the assembly's blind spots.
What this means for breeding, and what remains open
The clearest downstream value is structural. With 97.25% of sequence on 44 pseudochromosomes and a 46.08 Mb scaffold N50 [1], the assembly supports the kind of synteny and chromosome-level comparison that Uengwetwanit et al. used to detect one-to-one and one-to-two relationships between P. monodon and L. vannamei pseudochromosomes, while explicitly flagging that some apparent one-to-two patterns could reflect scaffolding errors in either assembly [6]. Better anchoring in L. vannamei reduces one source of that ambiguity.
The annotation also enables gene-family and trait work that earlier drafts could not support cleanly. Liu et al. used a chromosome-level Cherax quadricarinatus assembly to identify expansion of immune-related gene families and positive selection on KDM3A, KDM5A, and HMOX2, then validated KDM5A's role in hypoxia tolerance with RNAi [4]. That study also illustrates a caution: their 3.95 Gb assembly was 2.07 Gb smaller than the estimated 6.02 Gb genome size, yet had higher BUSCO completeness and a higher Merqury QV than the larger published assembly [4]. Assembly size and completeness can move in opposite directions, which is exactly the pattern seen between the new L. vannamei assembly and its 2025 predecessor [1].
Several questions remain genuinely open. First, whether the 70.98% BUSCO reflects true missing sequence, lineage-dataset mismatch, or measurement tool sensitivity is unresolved in the supplied evidence [1][5]. Second, the discrepancy between the new assembly's completeness and the 2025 assembly's 84.6% is not explained [1]. Third, the assembly derives from a single 62 g male from one commercial hatchery [1], so it captures one haplotype's worth of structural variation in a species with multiple breeding lines adapted to different environments [1][2]. Fourth, the 71.95% repeat content [1] means that repeat-mediated structural variation, a plausible contributor to trait variation in penaeids [8], remains largely inaccessible. The genome is a better map, not a finished one.
About These Sources
This research page is built on 9 peer-reviewed studies — published from 2017 to 2026, 2 from 2024 or later, 4 in Q1 journals, collectively cited 543 times — selected as the most relevant from 13 studies that passed quality screening, drawn from 177 papers retrieved from a database of over 500 million.
Sources used in this answer
An improved chromosome-scale genome assembly of whiteleg shrimp, Litopenaeus vannamei
Ma et al. present a 1.95 Gb chromosome-scale L. vannamei assembly integrating PacBio HiFi, ONT ultra-long, and Hi-C data, with a 1.97 Mb contig N50, 97.25% anchoring to 44 pseudochromosomes, 26,830 protein-coding genes, 71.95% repetitive sequence, and BUSCO completeness of 70.98% (genome) and 69.80% (proteome).
Selection Signatures of Pacific White Shrimp Litopenaeus vannamei Revealed by Whole-Genome Resequencing Analysis
Wang et al. resequenced 180 L. vannamei from two selective breeds and four companies, identifying 37 million SNPs at 22.5 SNPs/Kb and selection signatures enriched for macromolecule metabolism, proteolysis, structural molecule activity, and stimulus response.
A New Standard for Crustacean Genomes: The Highly Contiguous, Annotated Genome Assembly of the Clam Shrimp Eulimnadia texana Reveals HOX Gene Order and Identifies the Sex Chromosome
Baldwin-Brown et al. assembled the 120 Mb clam shrimp Eulimnadia texana genome with a contig N50 of 18 Mb from PacBio and Illumina data, describing it as the most contiguous crustacean assembly of its time and using it to locate the sex chromosome.
Genome assembly of redclaw crayfish (Cherax quadricarinatus) provides insights into its immune adaptation and hypoxia tolerance
Liu et al. built a chromosome-level Cherax quadricarinatus assembly (3.95 Gb, contig N50 739.45 kb, scaffold N50 33.93 Mb, 100 pseudomolecules) with higher BUSCO completeness and Merqury QV than a larger published assembly, and validated KDM5A's role in hypoxia tolerance via RNAi.
compleasm: a faster and more accurate reimplementation of BUSCO
Huang and Li show compleasm is 14 times faster than BUSCO and reports 99.6% completeness for T2T-CHM13 versus BUSCO's 95.7%, while cautioning that compleasm has limited sensitivity to distant homologs and should be combined with BUSCO for divergent assemblies.
A chromosome‐level assembly of the black tiger shrimp ( <i>Penaeus monodon</i> ) genome facilitates the identification of growth‐associated genes
Uengwetwanit et al. produced the first chromosome-level P. monodon assembly (44 pseudochromosomes, scaffold N50 44.86 Mb) from PacBio plus Chicago and Hi-C data, and used synteny with L. vannamei to identify growth-associated genes while flagging possible scaffolding errors in either assembly.
Evaluating Illumina-, Nanopore-, and PacBio-based genome assembly strategies with the bald notothen,<i>Trematomus borchgrevinki</i>
Rayamajhi et al. compared short-read, hybrid, and long-read-only assemblies of Trematomus borchgrevinki and concluded that long-read-only contig assembly is the current best choice, with phase I and phase II assemblies of lower quality and hidden mate-pair scaffolding errors degrading hybrid assemblies.
Genome Sequencing and Assembly Strategies and a Comparative Analysis of the Genomic Characteristics in Penaeid Shrimp Species
Yuan et al. surveyed four penaeid shrimp species and attributed their poor assembly to high heterozygosity or abundant homo-duplicated repeats rather than whole-genome duplication, estimating heterozygosity at 2.43% in L. vannamei, 1.95% in F. chinensis, 4.95% in P. monodon, and 4.49% in M. japonicus.
Comparative Analysis of PacBio and Oxford Nanopore Sequencing Technologies for Transcriptomic Landscape Identification of Penaeus monodon
Udaondo et al. compared PacBio and ONT for transcriptomic analysis across three P. monodon tissues and found that read-length bias and throughput differences strongly influence results, concluding that platform choice should be guided by the research question.
