Datasets

Training datasets with quantum-safe provenance

HuggingFaceFW/finepdfs_lang_classification HF Unverified

Size_categories:1M<n<10MFormat:parquetModality:tabularLibrary:datasetsLibrary:pandasLibrary:mlcroissant
abisee/cnn_dailymail HF Unverified

Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.

Task_categories:summarizationTask_ids:news-Articles-SummarizationAnnotations_creators:no-AnnotationLanguage_creators:foundMultilinguality:monolingualSource_datasets:original
jasperai/monet HF Unverified

Dataset Card for MONET MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 104.9 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with… See the full description on the dataset page: https://huggingface.co/datasets/jasperai/monet.

Task_categories:text-To-ImageTask_categories:image-Feature-ExtractionTask_categories:zero-Shot-Image-ClassificationLanguage:enSize_categories:100M<n<1BMultimodal
stanford-vision-lab/gpic HF Unverified

GPIC: A Giant Permissive Image Corpus for Visual Generation Keshigeyan&nbsp;Chandrasegaran*1,&nbsp; Kyle&nbsp;Sargent*1,&nbsp; Suchir&nbsp;Agarwal1,&nbsp; Michael&nbsp;Jang1,&nbsp; Michael&nbsp;Poli1,2,&nbsp; Juan&nbsp;Carlos&nbsp;Niebles1,4,&nbsp; Justin&nbsp;Johnson3,&nbsp; Jiajun&nbsp;Wu1,&nbsp; Li&nbsp;Fei-Fei1 1&nbsp;Stanford University&nbsp;&nbsp; 2&nbsp;Radical Numerics&nbsp;&nbsp; 3&nbsp;University of Michigan&nbsp;&nbsp; 4&nbsp;Salesforce… See the full description on the dataset page: https://huggingface.co/datasets/stanford-vision-lab/gpic.

Language:en
labelmaker/arkit_labelmaker HF Unverified

ARKit Labelmaker: A New Scale for Indoor 3D Scene Understanding [arxiv] [website] [checkpoints] [code] We complement ARKitScenes dataset with dense semantic annotations that are automatically generated at scale. This produces the first large-scale, real-world 3D dataset with dense semantic annotations. Training on this auto-generated data, we push forward the state-of-the-art performance on ScanNet and ScanNet200 with prevalent 3D semantic segmentation models.

Task_categories:image-SegmentationLanguage:enSize_categories:1K<n<10KDoi:10.57967/hf/23893D semantic segmentationIndoor 3D scene dataset
CERN/ColliderML-Release-1 HF Unverified

ColliderML: Dataset Release 1 Dataset Description This dataset contains simulated high-energy physics collision events generated using the Open Data Detector (ODD) geometry within the Key4hep and ACTS (A Common Tracking Software) frameworks, representing a generic collider detector similar to those at the HL-LHC. Dataset Summary Collision Energy: 14 TeV (proton-proton) Detector: Open Data Detector (ODD) Simulation: DD4hep + Geant4 + ACTS Format: Apache Parquet… See the full description on the dataset page: https://huggingface.co/datasets/CERN/ColliderML-Release-1.

Task_categories:otherSize_categories:10M<n<100MFormat:parquetModality:timeseriesLibrary:datasetsLibrary:dask
aline-gassenn/MedDialog-Audio HF Unverified

MedDialogue-Audio English Medical Dialogue Corpus for Speech Recognition Research. This repository contains MedDialogue-Audio, an English audio corpus designed for research in Automatic Speech Recognition (ASR) in the healthcare domain. The dataset was published in the proceedings of the 7th SBBD Dataset Showcase Workshop, and is available online at the following link: https://sol.sbc.org.br/index.php/dsw/article/view/37199 Dataset Description MedDialogue-Audio is… See the full description on the dataset page: https://huggingface.co/datasets/aline-gassenn/MedDialog-Audio.

Task_categories:automatic-Speech-RecognitionLanguage:enSize_categories:100K<n<1MDoi:10.57967/hf/5889Medical
Williamsanderson/MedQA-Darija-MultiLingual HF Unverified

MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.

Task_categories:question-AnsweringTask_categories:automatic-Speech-RecognitionTask_categories:text-To-SpeechLanguage:arLanguage:frLanguage:en
nvidia/SAGE-10k HF Unverified

SAGE-10k SAGE-10k is a large-scale interactive indoor scene dataset featuring realistic layouts, generated by the agentic-driven pipeline introduced in "SAGE: Scalable Agentic 3D Scene Generation for Embodied AI". The dataset contains 10,000 diverse scenes spanning 50 room types and styles, along with 565K uniquely generated 3D objects. 🔑 Key Features SAGE-10k integrates a wide variety of scenes, and particularly, preserves small items for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SAGE-10k.

Task_categories:text-To-3dLanguage:enSize_categories:10K<n<100KScene-GenerationInteractive-ScenesEmbodied-AI
allenai/winogrande HF Unverified

Dataset Card for "winogrande" Dataset Summary WinoGrande is a new collection of 44k problems, inspired by Winograd Schema Challenge (Levesque, Davis, and Morgenstern 2011), but adjusted to improve the scale and robustness against the dataset-specific bias. Formulated as a fill-in-a-blank task with binary options, the goal is to choose the right option for a given sentence which requires commonsense reasoning. Supported Tasks and Leaderboards More Information… See the full description on the dataset page: https://huggingface.co/datasets/allenai/winogrande.

Language:enSize_categories:10K<n<100KFormat:parquetModality:textLibrary:datasetsLibrary:pandas
InternRobotics/OmniWorld HF Unverified

[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling         🎉NEWS [2026.3.21] 🔥 OmniWorld-Game with Metric Scale is now released! Check out our latest model Pi3X (an enhanced version of Pi3), which leverages this data to achieve better performance! [2026.1.26] 🎉 OmniWorld was accepted by ICLR 2026! [2026.1.7] Update OmniWorld-Game, release RH20T-Robot, RH20T-Human, Ego-Exo4D, EgoDex, Epic-Kitchens. [2025.11.11] The OmniWorld is… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/OmniWorld.

Task_categories:text-To-VideoTask_categories:image-To-VideoTask_categories:image-To-3dTask_categories:roboticsTask_categories:otherLanguage:en
imageomics/fish-vista HF Unverified

Dataset Card for Fish-Visual Trait Analysis (Fish-Vista) Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images. See Example Code to Use the Segmentation Dataset Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset. Instructions for downloading dataset and images… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fish-vista.

Task_categories:image-ClassificationTask_categories:image-SegmentationLanguage:enSize_categories:10K<n<100KFormat:csvModality:image
bigcode/the-stack-metadata HF Unverified

Dataset Card for The Stack Metadata Changelog Release Description v1.1 This is the first release of the metadata. It is for The Stack v1.1 v1.2 Metadata dataset matching The Stack v1.2 Dataset Summary This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories. Supported Tasks and Leaderboards The main task is to recreate… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.

Task_categories:text-GenerationLanguage_creators:crowdsourcedLanguage_creators:expert-GeneratedMultilinguality:multilingualLanguage:codeSize_categories:10B<n<100B
HuggingFaceM4/the_cauldron HF Unverified

Dataset Card for The Cauldron Dataset description The Cauldron is part of the Idefics2 release. It is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2. Load the dataset To load the dataset, install the library datasets with pip install datasets. Then, from datasets import load_dataset ds = load_dataset("HuggingFaceM4/the_cauldron", "ai2d") to download and load the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/the_cauldron.

Size_categories:1M<n<10MFormat:parquetModality:imageModality:textLibrary:datasetsLibrary:dask
OpenGVLab/GUI-Odyssey HF Unverified

Dataset Card for GUI Odyssey News⭐️ A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉 👉 Please use the latest version and refer to the updated README for the most up-to-date information. We highly recommend using the new version for all training and evaluation! Repository: https://github.com/OpenGVLab/GUI-Odyssey Latest Version of Dataset: hflqf88888/GUIOdyssey Paper: https://arxiv.org/pdf/2406.08451 Introduction GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.

Language:enSize_categories:1K<n<10KFormat:jsonModality:imageModality:tabularModality:text
PleIAs/common_corpus HF Unverified

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open and permissible licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners. Common Corpus differs from existing open datasets in that it is: Truly Open: contains only data that… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/common_corpus.

Language:enLanguage:frLanguage:deLanguage:zhLanguage:itLanguage:es
annoymous-1/CC-Bench HF Unverified

CC-Bench: A Cognitive Conflict Benchmark for MLLMs in Safety-Critical Visual Inspection CC-Bench is a joint medical-industrial benchmark for evaluating whether multimodal large language models (MLLMs) remain visually grounded when plausible textual context conflicts with image evidence. The benchmark reorganizes public anomaly datasets into a unified four-way multiple-choice QA format for high-risk visual inspection. This repository currently contains: 4,282 images in total 2,157… See the full description on the dataset page: https://huggingface.co/datasets/annoymous-1/CC-Bench.

Task_categories:visual-Question-AnsweringTask_categories:image-ClassificationSize_categories:1K<n<10KFormat:jsonModality:imageModality:text
codeparrot/github-code-clean HF Unverified

The GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in almost 1TB of text data.

Size_categories:10M<n<100MModality:textLibrary:datasetsLibrary:mlcroissant
google-research-datasets/mbpp HF Unverified

Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.

Annotations_creators:crowdsourcedAnnotations_creators:expert-GeneratedLanguage_creators:crowdsourcedLanguage_creators:expert-GeneratedMultilinguality:monolingualSource_datasets:original
zekaiwang/trex_dataset HF Unverified

T-Rex Dataset A large-scale, tactile-reactive bimanual manipulation dataset, collected via teleoperation on a Dexmate Vega-1 robot with two Sharpa Wave dexterous hands. Stored as a LeRobotDataset v3.0. 🌐 Project Page · ✍️ Paper (arXiv) · 💻 Code (T-Rex) · 🚀 Dataset Quickstart · 📓 Colab notebook One episode from each of 20 motor primitives (head-camera view, cropped to the workspace), each with a different object. Teleoperation setup: Manus gloves + VIVE… See the full description on the dataset page: https://huggingface.co/datasets/zekaiwang/trex_dataset.

Task_categories:roboticsLanguage:enSize_categories:1M<n<10MFormat:parquetModality:tabularModality:text
Showing 20 of 208 datasets (page 4 of 11)