PlantBGC predicts compact candidate biosynthetic gene cluster loci from plant genomic or CDS sequences.

Upload your genome, receive annotated BGC candidate regions by email.

Submit a Job View Example
About PlantBGC
PlantBGC is a Transformer-based framework for plant BGC discovery. It converts plant sequences into Pfam-domain representations, predicts candidate BGC loci using a three-stage training strategy (microbial supervision → label-free plant domain adaptation → optional weak supervision), and returns downloadable results by email.
Submit a Job

Job Submission

Click to browse or drag & drop
Accepted formats: .fna  ·  .fa  ·  .fasta  ·  .gbk  ·  .gbff
Allow use for training
Contributed files may be used to improve future PlantBGC models
Uploading…
Output

Results will be sent to the provided email address with a download link. Returned files may include:

*.bgc.tsv
Candidate BGC prediction table with locus coordinates and scores
*.pfam.tsv
Per-protein Pfam domain annotation and BGC-likeness scores
*.full.gbk
Full annotated GenBank file with all predicted features
*.bgc.gbk
GenBank file filtered to BGC candidate regions only
LOG.txt
Full run log for reproducibility and debugging

No online visualization is provided. Results are returned as downloadable files.

Example Usage

PlantBGC can also be run locally via the command line:

# Full prediction pipeline: annotate + detect BGC candidate loci
plantbgc Predict --output prediction_output/ --score 0.5 --min-proteins 3 input.fna
# Protein FASTA input (skip gene calling)
plantbgc Predict --protein --output prediction_output/ proteins.fasta
Documentation

Full documentation including installation instructions, data format specifications, and parameter descriptions are available in the PlantBGC GitHub repository.

Contact

For questions or issues, please open an issue on GitHub or contact the development team at blalu@ncsu.edu or yzhao66@ncsu.edu.