Datasets

Training datasets with quantum-safe provenance

jasperai/monet HF Unverified

Dataset Card for MONET MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 104.9 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with… See the full description on the dataset page: https://huggingface.co/datasets/jasperai/monet.

Task_categories:text-To-ImageTask_categories:image-Feature-ExtractionTask_categories:zero-Shot-Image-ClassificationLanguage:enSize_categories:100M<n<1BMultimodal
TIGER-Lab/MMLU-Pro HF Unverified

MMLU-Pro Dataset MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. |Github | 🏆Leaderboard | 📖Paper | 🚀 What's New [2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro.

Benchmark:officialTask_categories:question-AnsweringLanguage:enSize_categories:10K<n<100KFormat:parquetModality:tabular
anisoleai/fineweb-tokenized HF PQC Verified

FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.

Task_categories:text-GenerationLanguage:enSize_categories:n>1TModality:tabularModality:textTabular
mlfoundations/MINT-1T-PDF-CC-2023-06 HF PQC Verified

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-06.

Task_categories:image-To-TextTask_categories:text-GenerationLanguage:enSize_categories:100B<n<1TMultimodal
SWE-bench/SWE-bench_Multilingual HF Unverified

Language:enSize_categories:n<1KFormat:parquetModality:textLibrary:datasetsLibrary:pandas
agibot-world/AgiBotWorld-Beta HF Unverified

Key Features 🔑 1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours. 100+ real-world scenarios across 5 target domains. Cutting-edge hardware: visual tactile sensors / 6-DoF dexterous hand / mobile dual-arm robots 200+ types of tasks: Contact-rich manipulation Long-horizon planning Multi-robot collaboration 87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc. Your… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta.

Task_categories:otherTask_categories:roboticsLanguage:enSize_categories:100M<n<1BFormat:webdatasetModality:text
HuggingFaceM4/FineVision HF Unverified

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision Load the data from datasets import load_dataset, get_dataset_config_names # Get all subset names and load the first one available_subsets =… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/FineVision.

Size_categories:10M<n<100MFormat:parquetModality:imageModality:textLibrary:datasetsLibrary:dask
GokuScraper/seedance-2-prompts-datasets HF Unverified

🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.

Task_categories:text-To-VideoLanguage:enLanguage:zhSize_categories:1K<n<10KModality:imageModality:video
CohereLabs/xP3x HF Unverified

Dataset Card for xP3x Dataset Summary xP3x (Crosslingual Public Pool of Prompts eXtended) is a collection of prompts & datasets across 277 languages & 16 NLP tasks. It contains all of xP3 + much more! It is used for training future contenders of mT0 & BLOOMZ at project Aya @Cohere Labs 🧡 Creation: The dataset can be recreated using instructions available here together with the file in this repository named xp3x_create.py. We provide this version to save processing… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/xP3x.

Task_categories:otherAnnotations_creators:expert-GeneratedAnnotations_creators:crowdsourcedMultilinguality:multilingualLanguage:afLanguage:ar
HuggingFaceH4/MATH-500 HF Unverified

Dataset Card for MATH-500 This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits

Task_categories:text-GenerationLanguage:enSize_categories:n<1KFormat:jsonModality:textLibrary:datasets
Kazimir-ai/text-to-image-prompts HF Unverified

The dataset of the most popular text-to-image prompts. Dataset Details Dataset Description Curated by: kazimir.ai Funded by [optional]: [More Information Needed] Shared by [optional]: https://kazimir.ai License: apache-2.0 Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Free to use. Dataset Structure CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Kazimir-ai/text-to-image-prompts.

Language:enSize_categories:10K<n<100KFormat:csvModality:textLibrary:datasetsLibrary:pandas
mvp-lab/LLaVA-OneVision-2-Data HF Unverified

LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used in mid-training. Dataset Composition Subset Format Description mid_training_video/60s_rest/ WebDataset (.tar) 10,809 shards of ~60s video clips mid_training_video/caption_v0/split_30s.jsonl JSONL Captions for 30-second video clips mid_training_video/caption_v0/split_60s.jsonl JSONL Captions for… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.

Task_categories:video-Text-To-TextTask_categories:visual-Question-AnsweringTask_categories:image-Text-To-TextLanguage:enSize_categories:n<1KFormat:parquet
RekaAI/RekaDaily-10k-raw HF Unverified

RekaDaily-10k (raw) Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are exactly as collected — no re-encoding, no cuts, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.

Task_categories:video-ClassificationTask_categories:image-To-VideoLanguage:enSize_categories:100K<n<1MFormat:parquetModality:image
WINGNUS/ACL-OCL HF Unverified

Dataset Card for ACL Anthology Corpus This repository provides full-text and metadata to the ACL anthology collection (80k articles/posters as of September 2022) also including .pdf files and grobid extractions of the pdfs. How is this different from what ACL anthology provides and what already exists? We provide pdfs, full-text, references and other details extracted by grobid from the PDFs while ACL Anthology only provides abstracts. There exists a similar corpus… See the full description on the dataset page: https://huggingface.co/datasets/WINGNUS/ACL-OCL.

Task_categories:token-ClassificationLanguage_creators:foundMultilinguality:monolingualSource_datasets:originalLanguage:enSize_categories:10K<n<100K
SWE-bench/SWE-smith HF Unverified

SWE-smith Dataset Code • Paper • Site [12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the `SWE-bench/SWE-smith-[lang]` datasets. For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo. The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit.… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-smith.

Task_categories:text-GenerationLanguage:enSize_categories:10K<n<100KFormat:parquetModality:textLibrary:datasets
google-research-datasets/paws HF Unverified

Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.

Task_categories:text-ClassificationTask_ids:semantic-Similarity-ClassificationTask_ids:semantic-Similarity-ScoringTask_ids:text-ScoringTask_ids:multi-Input-Text-ClassificationAnnotations_creators:expert-Generated
jsulz/FIFA23 HF Unverified

About this dataset Context The datasets provided include the players data for the Career Mode from FIFA 15 to FIFA 23. The data allows multiple comparisons for the same players across the last 9 versions of the video game. Some ideas of possible analysis: Historical comparison between Messi and Ronaldo (what skill attributes changed the most during time - compared to real-life stats); Ideal budget to create a competitive team (at the level of top n teams in Europe) and… See the full description on the dataset page: https://huggingface.co/datasets/jsulz/FIFA23.

Task_categories:tabular-ClassificationTask_categories:tabular-RegressionLanguage:enSize_categories:10M<n<100MModality:tabularTabular
HuggingFaceFW/finepdfs_lang_classification HF Unverified

Size_categories:1M<n<10MFormat:parquetModality:tabularLibrary:datasetsLibrary:pandasLibrary:mlcroissant
abisee/cnn_dailymail HF Unverified

Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.

Task_categories:summarizationTask_ids:news-Articles-SummarizationAnnotations_creators:no-AnnotationLanguage_creators:foundMultilinguality:monolingualSource_datasets:original
labelmaker/arkit_labelmaker HF Unverified

ARKit Labelmaker: A New Scale for Indoor 3D Scene Understanding [arxiv] [website] [checkpoints] [code] We complement ARKitScenes dataset with dense semantic annotations that are automatically generated at scale. This produces the first large-scale, real-world 3D dataset with dense semantic annotations. Training on this auto-generated data, we push forward the state-of-the-art performance on ScanNet and ScanNet200 with prevalent 3D semantic segmentation models.

Task_categories:image-SegmentationLanguage:enSize_categories:1K<n<10KDoi:10.57967/hf/23893D semantic segmentationIndoor 3D scene dataset
Showing 20 of 270 datasets (page 4 of 14)