Analyze and fix .gtf file for Drop-seq pipeline to remove warnings and skipping for genes with issues
This R script demonstrates a bioinformatics workflow for cleaning a modified GTF genome annotation file by identifying and filtering problematic gene entries before running a Drop-seq pipeline. It shows how duplicated gene names and multiple gene IDs cause warnings and skips during processing. The code uses the `refGenome` package to parse the GTF and extract gene positions, then leverages `dplyr` and `tidyr` for data manipulation to categorize issues. It cross-references gene expression data from two DGE files and queries the Ensembl database via `biomaRt` to retrieve current gene IDs. The visualization ultimately classifies problematic genes by duplication, chromosomal disagreement, or strand disagreement, filtering them out to create a clean annotation set.
AI-generated description