Machine learning with the TCGA-HNSC dataset: Improving usability by addressing inconsistency, sparsity, and high-dimensionality

20Citations
Citations of this article
96Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Background: In the era of precision oncology and publicly available datasets, the amount of information available for each patient case has dramatically increased. From clinical variables and PET-CT radiomics measures to DNA-variant and RNA expression profiles, such a wide variety of data presents a multitude of challenges. Large clinical datasets are subject to sparsely and/or inconsistently populated fields. Corresponding sequencing profiles can suffer from the problem of high-dimensionality, where making useful inferences can be difficult without correspondingly large numbers of instances. In this paper we report a novel deployment of machine learning techniques to handle data sparsity and high dimensionality, while evaluating potential biomarkers in the form of unsupervised transformations of RNA data. We apply preprocessing, MICE imputation, and sparse principal component analysis (SPCA) to improve the usability of more than 500 patient cases from the TCGA-HNSC dataset for enhancing future oncological decision support for Head and Neck Squamous Cell Carcinoma (HNSCC). Results: Imputation was shown to improve prognostic ability of sparse clinical treatment variables. SPCA transformation of RNA expression variables reduced runtime for RNA-based models, though changes to classifier performance were not significant. Gene ontology enrichment analysis of gene sets associated with individual sparse principal components (SPCs) are also reported, showing that both high- and low-importance SPCs were associated with cell death pathways, though the high-importance gene sets were found to be associated with a wider variety of cancer-related biological processes. Conclusions: MICE imputation allowed us to impute missing values for clinically informative features, improving their overall importance for predicting two-year recurrence-free survival by incorporating variance from other clinical variables. Dimensionality reduction of RNA expression profiles via SPCA reduced both computation cost and model training/evaluation time without affecting classifier performance, allowing researchers to obtain experimental results much more quickly. SPCA simultaneously provided a convenient avenue for consideration of biological context via gene ontology enrichment analysis.

References Powered by Scopus

Gene ontology: Tool for the unification of biology

32326Citations
N/AReaders
Get full text

A review of feature selection techniques in bioinformatics

4125Citations
N/AReaders
Get full text

Bias in random forest variable importance measures: Illustrations, sources and a solution

2559Citations
N/AReaders
Get full text

Cited by Powered by Scopus

The application of deep learning in cancer prognosis prediction

267Citations
N/AReaders
Get full text

Computational Oncology in the Multi-Omics Era: State of the Art

66Citations
N/AReaders
Get full text

An up-to-date overview of computational polypharmacology in modern drug discovery

56Citations
N/AReaders
Get full text

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Cite

CITATION STYLE

APA

Rendleman, M. C., Buatti, J. M., Braun, T. A., Smith, B. J., Nwakama, C., Beichel, R. R., … Casavant, T. L. (2019). Machine learning with the TCGA-HNSC dataset: Improving usability by addressing inconsistency, sparsity, and high-dimensionality. BMC Bioinformatics, 20(1). https://doi.org/10.1186/s12859-019-2929-8

Readers' Seniority

Tooltip

PhD / Post grad / Masters / Doc 25

58%

Researcher 14

33%

Professor / Associate Prof. 3

7%

Lecturer / Post doc 1

2%

Readers' Discipline

Tooltip

Computer Science 12

28%

Medicine and Dentistry 12

28%

Biochemistry, Genetics and Molecular Bi... 10

23%

Agricultural and Biological Sciences 9

21%

Article Metrics

Tooltip
Mentions
Blog Mentions: 1

Save time finding and organizing research with Mendeley

Sign up for free