Category
page 1Datasets in machine learning
Common Voice
voice dataset by Mozilla
training, validation, and test data sets
three datasets used in machine learning
Iris flower data set
1936 dataset of flowers
BookCorpus
BookCorpus (also sometimes referred to as the Toronto Book Corpus) is a dataset consisting of the text of around 7,000 self-published books scraped from the indie ebook distribution website Smashwords. It was the main corpus used to train the initial GPT model by OpenAI, and has been used as training data for other early large language models including Google's BERT. The dataset consists of around 985 million words, and the books that comprise it span a range of genres, including romance, science fiction, and fantasy.