Datasets

Training datasets with quantum-safe provenance

lishaoyong/latex-formulas-80M HF Unverified

For more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository. IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set

Size_categories:10M<n<100MFormat:parquetModality:imageModality:textLibrary:datasetsLibrary:dask
MahmoodLab/hest HF Unverified

Model Card for HEST-1k What is HEST-1k? A collection of 1,276 spatial transcriptomic profiles, each linked and aligned to a Whole Slide Image (with pixel size < 1.15 µm/px) and metadata. HEST-1k was assembled from 180 public and internal cohorts encompassing: 26 organs 2 species (Homo Sapiens and Mus Musculus) 398 cancer samples from 25 cancer types. HEST-1k processing enabled the identification of >1.5 million expression/morphology pairs and >76 million nuclei… See the full description on the dataset page: https://huggingface.co/datasets/MahmoodLab/hest.

Task_categories:image-ClassificationTask_categories:feature-ExtractionTask_categories:image-SegmentationLanguage:enSize_categories:100B<n<1TSpatial-Transcriptomics
uoft-cs/cifar100 HF Unverified

Dataset Card for CIFAR-100 Dataset Summary The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses. There are two labels per image - fine label (actual class) and coarse label (superclass). Supported Tasks and Leaderboards image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar100.

Task_categories:image-ClassificationAnnotations_creators:crowdsourcedLanguage_creators:foundMultilinguality:monolingualSource_datasets:extended|other-80-Million-Tiny-ImagesLanguage:en
saffatgazi/fish-vista HF Unverified

Dataset Card for Fish-Visual Trait Analysis (Fish-Vista) Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images. See Example Code to Use the Segmentation Dataset Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset. Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/saffatgazi/fish-vista.

Task_categories:image-ClassificationTask_categories:image-SegmentationLanguage:enSize_categories:10K<n<100KFormat:csvModality:image
Skylion007/openwebtext HF Unverified

Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset files: 13.51 GB Size of the… See the full description on the dataset page: https://huggingface.co/datasets/Skylion007/openwebtext.

Task_categories:text-GenerationTask_categories:fill-MaskTask_ids:language-ModelingTask_ids:masked-Language-ModelingAnnotations_creators:no-AnnotationLanguage_creators:found
DWKPartners/trading-data HF Unverified

DartLab Data Structured company data from DART & EDGAR disclosure filings DART 전자공시 + EDGAR 공시 데이터 — 한국 2,700사 / 미국 970사 What is this? Pre-collected Parquet files from DartLab — a Python library that turns DART (Korea) and EDGAR (US) disclosure filings into one structured company map. 한국 DART 전자공시 시스템과 미국 SEC EDGAR에서 수집한 기업 공시 데이터입니다. This dataset is the data layer behind DartLab. When you run dartlab.Company("005930"), the library automatically downloads the… See the full description on the dataset page: https://huggingface.co/datasets/DWKPartners/trading-data.

Task_categories:table-Question-AnsweringTask_categories:text-ClassificationLanguage:koLanguage:enSize_categories:1K<n<10KFinance
X779/Danbooruwildcards HF Unverified

This is a set of wildcards for danbooru tags. Artist:Prompts for random artist styles, covering approximately 0.6M different artists.Please select the appropriate version of the collection, ranging from 128 to 5000, based on the model's capabilities.The full version is not recommended for use as it includes too many artists with only one image on danbooru or other websites. Almost no model can generate a style that corresponds to these artists . Characters:"Characters" is a set of wildcards… See the full description on the dataset page: https://huggingface.co/datasets/X779/Danbooruwildcards.

Task_categories:text-GenerationLanguage:enSize_categories:10M<n<100MFormat:textModality:textLibrary:datasets
DWKPartners/dartlab-data HF Unverified

DartLab Data Structured company data from DART & EDGAR disclosure filings DART 전자공시 + EDGAR 공시 데이터 — 한국 2,700사 / 미국 970사 What is this? Pre-collected Parquet files from DartLab — a Python library that turns DART (Korea) and EDGAR (US) disclosure filings into one structured company map. 한국 DART 전자공시 시스템과 미국 SEC EDGAR에서 수집한 기업 공시 데이터입니다. This dataset is the data layer behind DartLab. When you run dartlab.Company("005930"), the library automatically downloads the… See the full description on the dataset page: https://huggingface.co/datasets/DWKPartners/dartlab-data.

Task_categories:table-Question-AnsweringTask_categories:text-ClassificationLanguage:koLanguage:enSize_categories:1K<n<10KFinance
facebook/wiki_dpr HF Unverified

This is the wikipedia split used to evaluate the Dense Passage Retrieval (DPR) model. It contains 21M passages from wikipedia along with their DPR embeddings. The wikipedia articles were split into multiple, disjoint text blocks of 100 words as passages.

Task_categories:fill-MaskTask_categories:text-GenerationTask_ids:language-ModelingTask_ids:masked-Language-ModelingAnnotations_creators:no-AnnotationLanguage_creators:crowdsourced
Snowfall0601/dartlab-data HF Unverified

DartLab Data Structured company data from DART & EDGAR disclosure filings DART 전자공시 + EDGAR 공시 데이터 — 한국 2,700사 / 미국 970사 What is this? Pre-collected Parquet files from DartLab — a Python library that turns DART (Korea) and EDGAR (US) disclosure filings into one structured company map. 한국 DART 전자공시 시스템과 미국 SEC EDGAR에서 수집한 기업 공시 데이터입니다. This dataset is the data layer behind DartLab. When you run dartlab.Company("005930"), the library automatically downloads the… See the full description on the dataset page: https://huggingface.co/datasets/Snowfall0601/dartlab-data.

Task_categories:table-Question-AnsweringTask_categories:text-ClassificationLanguage:koLanguage:enSize_categories:1K<n<10KFinance
nyu-mll/blimp HF Unverified

Dataset Card for "blimp" Dataset Summary BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars. Supported Tasks and Leaderboards More Information Needed Languages More Information… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/blimp.

Task_categories:text-ClassificationTask_ids:acceptability-ClassificationAnnotations_creators:crowdsourcedLanguage_creators:machine-GeneratedMultilinguality:monolingualSource_datasets:original
open-index/open-github HF Unverified

OpenGitHub What is it? This dataset contains every public event on GitHub: every push, pull request, issue, star, fork, code review, release, and discussion across all public repositories. GitHub is the world's largest software development platform, home to over 200 million repositories and the daily work of tens of millions of developers, from individual open-source contributors to the engineering teams behind the most widely used software on earth. The archive currently… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github.

Task_categories:text-GenerationTask_categories:text-ClassificationTask_categories:feature-ExtractionLanguage:enLanguage:mulSize_categories:100K<n<1M
BrightData/Goodreads-Books HF Unverified

Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.

Task_categories:text-ClassificationTask_categories:summarizationTask_categories:text-GenerationLanguage:enSize_categories:1M<n<10MFormat:csv
muset-ai/DeepResearch-Bench-II-Dataset HF Unverified

Task_categories:text-GenerationTask_categories:text-ClassificationLanguage:zhLanguage:enSize_categories:n<1KModality:document
Voxel51/Office-Home HF Unverified

Dataset Card for Office-Home This is a FiftyOne dataset with 15588 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = fouh.load_from_hub("Voxel51/Office-Home") # Launch the App session = fo.launch_app(dataset) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Office-Home.

Task_categories:image-ClassificationLanguage:enSize_categories:10K<n<100KFormat:imagefolderModality:imageLibrary:datasets
zai-org/LongBench HF Unverified

LongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.

Task_categories:question-AnsweringTask_categories:text-GenerationTask_categories:summarizationTask_categories:text-ClassificationLanguage:enLanguage:zh
ceval/ceval-exam HF Unverified

C-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details. Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model… See the full description on the dataset page: https://huggingface.co/datasets/ceval/ceval-exam.

Task_categories:text-ClassificationTask_categories:multiple-ChoiceTask_categories:question-AnsweringLanguage:zhSize_categories:10K<n<100KFormat:parquet
cornell-movie-review-data/rotten_tomatoes HF Unverified

Dataset Card for "rotten_tomatoes" Dataset Summary Movie Review Dataset. This is a dataset of containing 5,331 positive and 5,331 negative processed sentences from Rotten Tomatoes movie reviews. This data was first used in Bo Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.'', Proceedings of the ACL, 2005. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/cornell-movie-review-data/rotten_tomatoes.

Task_categories:text-ClassificationTask_ids:sentiment-ClassificationAnnotations_creators:crowdsourcedLanguage_creators:crowdsourcedMultilinguality:monolingualSource_datasets:original
ethz/food101 HF Unverified

Dataset Card for Food-101 Dataset Summary This dataset consists of 101 food categories, with 101'000 images. For each class, 250 manually reviewed test images are provided as well as 750 training images. On purpose, the training images were not cleaned, and thus still contain some amount of noise. This comes mostly in the form of intense colors and sometimes wrong labels. All images were rescaled to have a maximum side length of 512 pixels. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ethz/food101.

Task_categories:image-ClassificationTask_ids:multi-Class-Image-ClassificationAnnotations_creators:crowdsourcedLanguage_creators:crowdsourcedMultilinguality:monolingualSource_datasets:extended|other-Foodspotting
Horama/animal-200 HF Unverified

Horama/animal-200 Raw wildlife image collection covering 199 species (mammals, birds, reptiles), scraped from multiple web sources. Images are organized by species folder and can be used as-is for image classification (species identification) or as input for downstream annotation pipelines (object detection, etc.). For animal detection, see Horama/animal-200-detection dataset. Sources Images were collected from three web sources using dedicated scrapers: Source… See the full description on the dataset page: https://huggingface.co/datasets/Horama/animal-200.

Task_categories:image-ClassificationTask_categories:object-DetectionLanguage:enLanguage:frSize_categories:10K<n<100KModality:image
Showing 20 of 208 datasets (page 9 of 11)