Online Resource

RNA-Seq Data readiness for machine learning applications

This lesson provides a practical guide to sourcing and Preprocessing a bulk RNA-Seq dataset for use in a machine learning classification task. The lessons explains the characteristics of a dataset required for this type of analysis, how to search for and download a dataset from each of the main public functional genomics repositories, and then provides guidelines on how to pre-process a dataset to make it machine learning ready, with detailed examples. The lesson finally explains some of the additional data filtering and transformation steps that will improve the performance of machine learning algorithms using RNA-Seq count data.

The lesson is written in the context of a supervised machine learning classification modelling task, where the goal is to construct a model that is able to differentiate two different disease states (e.g. disease vs. healthy control) based on the gene expression profile.

Keywords: Dataset, FASTQ, Metadata, Short reads, bioinformatics, classification, cross-validation, data processing, feature engineering, genomics, machine learning, model evaluation, regression, reproducibility, scikit-learn, statistics, supervised learning

Target audience: Postgraduate student, Postgraduate researcher, Researcher, Clinician, Engineer, Trainer, Industry, Public sector, Research Software Engineer, Bioinformatician, Open Research Specialist, Research skills educator

Resource type: Online Resource

Authors: Edward Parkinson, Munazah Andrabi, Katarzyna Kamieniecka


Activity log