An engineered phage is only as reliable as the prophage coordinates it is built from. Before any CRISPR edit or in vitro reboot, the dormant phage sequence hidden in a bacterial chromosome must be found, bounded, and ranked for engineering value. Creative BioMart Microbe delivers In Silico Prophage Discovery for Phage Engineering — a computational service that converts raw bacterial genome data into annotated, editing-ready prophage loci.
We bridge bacterial genome sequencing to actionable phage engineering targets. Predicted prophage regions are delivered with precise attL/attR boundaries, functional module maps, and rebooting-feasibility scores, eliminating the trial-and-error of selecting targets for an editing campaign. The output feeds directly into our downstream phage genome editing and large-scale production pipelines.
Prophages are dormant genetic reservoirs, not discarded sequence. Recognizing their engineering potential is the first step toward deliberate phage design rather than opportunistic screening.

Figure 1. In silico prophage discovery platform overview — converting raw bacterial genome data into annotated, editing-ready prophage loci through computational prediction, boundary localization, functional annotation, and engineering-suitability assessment.
Our discovery pipeline moves from initial customer consultation through data intake, computational analysis, and engineering handoff in six stages, with each stage producing a defined, handoff-ready artifact and clear client touchpoints.

Multi-Algorithm Prophage Detection
We run complementary prediction tools — including geNomad, PIDE, PHASTER, and VIBRANT — in parallel and combine their calls into a consensus set. Cross-tool agreement minimizes single-algorithm bias and false positives. Deliverable: a ranked candidate prophage list.

Boundary Refinement & att Site Localization
For each candidate we refine attL/attR coordinates and validate integrase binding sites, producing boundaries tight enough for primer design and rebooting experiments. Coordinates are reported with precision reflecting assembly quality and algorithm confidence.

Integrity Assessment & Functional Annotation
Each prophage is graded for completeness (Intact / Questionable / Incomplete) and annotated for structural proteins, lysis modules, regulatory elements, and auxiliary metabolic genes (AMGs). Output is delivered in annotation-ready formats.

Engineering Suitability Scoring & Blueprint
Candidates are prioritized by rebooting feasibility, host-range potential, and therapeutic relevance, with custom scoring for phage therapy, biocontrol, or synthetic biology. The result is a blueprint naming which loci to excise, modify, or mine.
| Deliverable | Format | Downstream Use |
|---|---|---|
| Annotated Prophage Genomes | FASTA + GFF3 / GBK | Direct import into editing pipelines |
| Engineering Target Report | Ranked candidate list + att sequences | Integrase-mediated excision design |
| Functional Module Map | Modular decomposition + cargo-insertion sites | Domain swapping, chimeric phage construction |
Predicted prophage sequences are transferred directly to CRISPR/Cas9 phage genome editing design, and att site coordinates enable precise homologous-recombination template construction. We offer a combined path — in silico discovery → in vitro rebooting → genome editing → scale-up production — managed as a single project from bacterial genome data to engineered phage delivery.
| Required Information | Optional Information | Not Accepted |
|---|---|---|
|
|
|
| Strategy | Platform | Assembly & Notes |
|---|---|---|
| ONT long-read | Oxford Nanopore long-read sequencing | Single-run finished genome; particularly effective for resolving att sites and complete contigs |
| Illumina short-read | Illumina paired-end sequencing (standard to high coverage) | High-accuracy base calls; suited to well-characterized strains |
| Hybrid | ONT + Illumina | Combined continuity and accuracy; recommended for novel or AT-rich hosts |
Genomic DNA and assemblies are accepted as digital files; for physical samples, ship on dry ice with glycerol backups and contact our team via contact us before sending BSL-2 materials.

Therapeutic Phage Revival
Reactivation of cryptic prophages from clinical isolates for phage-based therapeutic development.

Biocontrol Agent Development
Prophage mining from environmental strains for agricultural and food-safety biocontrol agents.

Probiotic Strain Safety
Prophage inventory and risk assessment during industrial bacterial strain qualification.

Synthetic Biology Chassis
Prophage-derived regulatory elements and structural modules for engineered phage construction.
geNomad is a classification and annotation framework that combines an alignment-free neural network (IGLOO) with a gene-based marker classifier. For integrated provirus detection, it employs a conditional random field model that scores each gene using marker specificity values from surrounding genes, then merges closely located viral islands and extends boundaries up to nearby tRNAs or integrases. Benchmarked against the TIGER database, this approach achieved high precision and sensitivity at the gene level while maintaining low contamination. In virus classification benchmarks it achieved a Matthews correlation coefficient of 95.3% (and 77.8% for plasmids), substantially outperforming other tools. Its marker database of 227,897 protein profiles provides functional gene annotation and taxonomic assignment of viral genomes. Scaled across 2.7 trillion base pairs of sequencing data, it enabled the discovery of millions of viruses and plasmids.

Figure 2. geNomad uses marker information to demarcate provirus boundaries. (Camargo et al. 2024)
VIBRANT is a neural-network viral identification and functional annotation platform that uses protein annotation signatures from KEGG, Pfam, and VOG databases together with a unique v-score metric to maximize identification of lytic viral genomes and integrated proviruses. Against 29,926 viral fragments it correctly identified 98.43% (versus VirSorter 40.03%–96.53%, VirFinder 76.23%–89.03%, MARVEL 93.79%), with specificity of 99.90% against genomic and 98.90% against plasmid fragments (precision 99.87%, accuracy 99.15%, F1 0.991). It also estimates genome quality across complete circular, high-quality draft, medium-quality draft, and low-quality draft categories, and distinguishes lytic from lysogenic viruses. When extracting integrated provirus regions from host scaffolds, it matched PHASTER (17 each) and recovered two putative proviruses the other programs missed.

Figure 3. Application of quality metrics to 2,466 NCBI RefSeq Caudovirales viruses with decreasing genome completeness from 100 to 10%, respective of total sequence length. (Kieft, et al. 2020)
A: We accept ONT long-read, Illumina short-read, and hybrid assemblies, delivered as FASTA or FASTQ. Unlike services locked to a single platform, our multi-platform intake supports both well-characterized strains and novel or AT-rich hosts. Assembled genomes are preferred, but we can work from raw reads with provided metadata.
A: We run geNomad, PIDE, PHASTER, and VIBRANT in parallel and take a cross-tool consensus. Candidates flagged by multiple independent algorithms are prioritized, which sharply reduces the false-positive rate that any single predictor would produce on its own.
A: You receive annotated prophage genomes as FASTA with precise boundary coordinates, plus GFF3 or GBK annotations for direct import into editing pipelines. The engineering target report adds a ranked candidate list, att site sequences for integrase-mediated excision, and a functional module map.
A: Yes. Our toolset includes protein-language and annotation-based models (PIDE, VIBRANT) that are not dependent on nucleotide homology alone, so they remain effective in AT-rich genomes where homology-based tools lose sensitivity. Hybrid assemblies further improve boundary resolution.
A: Predicted prophage sequences feed directly into editing design, and attL/attR coordinates let us build precise homologous-recombination templates. We can manage the full chain — discovery, in vitro rebooting, genome editing, and scale-up — as one project. See our phage genome editing service.
A: Yes. We build a prophage inventory and evaluate spontaneous-induction risk during industrial strain qualification, and we flag virulence-associated prophages as targets for attenuation or removal. This supports safer probiotic strain development.
A: No. Even incomplete prophages are valuable sources of modular genetic elements — tail fibers, endolysins, and depolymerases — that can be harvested for domain swapping and chimeric phage construction. Our functional module map identifies these non-essential regions and cargo-insertion sites.
References:
Enter your email here to subscribe