Gist 4127915bd94a1fc0518bb0cceea3b612
✓ Published0🌍 Public
NNpfn
Last edited Feb 8, 2019
Created on Feb 8, 2019
This example shows a Perl script that filters complete CEGs (Core Eukaryotic Genes) from a CEGMA output FASTA file, extracting only sequences corresponding to KOG identifiers not listed as missing in the completeness report. It reads three input files—the CEGMA FASTA, a completeness cutoff table, and the completeness report—then writes a new FASTA file. The script uses Perl’s built-in file handling and regular expression matching to parse KOG IDs, track a flag for sequence inclusion, and output the selected sequences.
AI-generated description