Author: Millán Arias, Pablo; Alipour, Fatemeh; Hill, Kathleen A.; Kari, Lila
Title: DeLUCS: Deep Learning for (Unsupervised) Clustering of DNA Sequences Cord-id: ae4o2kbm Document date: 2021_8_23
ID: ae4o2kbm
Snippet: We present a novel Deep Learning method for the (Unsupervised) Clustering of DNA Sequences (DeLUCS) that does not require sequence alignment, sequence homology, or (taxonomic) identifiers. DeLUCS uses Chaos Game Representations (CGRs) of primary DNA sequences, and generates “mimic†sequence CGRs to self-learn data patterns (genomic signatures) through the optimization of multiple neural networks. A majority voting scheme is then used to determine the final cluster assignment for each sequenc
Document: We present a novel Deep Learning method for the (Unsupervised) Clustering of DNA Sequences (DeLUCS) that does not require sequence alignment, sequence homology, or (taxonomic) identifiers. DeLUCS uses Chaos Game Representations (CGRs) of primary DNA sequences, and generates “mimic†sequence CGRs to self-learn data patterns (genomic signatures) through the optimization of multiple neural networks. A majority voting scheme is then used to determine the final cluster assignment for each sequence. The clusters learned by DeLUCS match true taxonomic groups for large and diverse datasets, with accuracies ranging from 77% to 100%: 2,500 complete vertebrate mitochondrial genomes, at taxonomic levels from sub-phylum to genera; 3,200 randomly selected 400 kbp-long bacterial genome segments, into clusters corresponding to bacterial families; three viral genome and gene datasets, averaging 1,300 sequences each, into clusters corresponding to virus subtypes. DeLUCS significantly outperforms two classic clustering methods (K-means++ and Gaussian Mixture Models) for unlabelled data, by as much as 47%. DeLUCS is highly effective, it is able to cluster datasets of unlabelled primary DNA sequences totalling over 1 billion bp of data, and it bypasses common limitations to classification resulting from the lack of sequence homology, variation in sequence length, and the absence or instability of sequence annotations and taxonomic identifiers. Thus, DeLUCS offers fast and accurate DNA sequence clustering for previously intractable datasets.
Search related documents:
Co phrase search for related documents- accurate classification and loss function: 1, 2
- adam optimizer and loss function: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10
- adam optimizer and loss function calculate: 1
Co phrase search for related documents, hyperlinks ordered by date