Model Hub
Browse PQC-verified AI models, datasets, and tools
π MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens π MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. π MINT-1T is designed to facilitate research in multimodal pretraining. π MINT-1T is created by a team from the University of Washington inβ¦ See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.
Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks designed to benchmark Large Language Models on lengthy audio inputs. π Dataset Statistics Task Statistics Task Train Val Test Total Unique Audios ASR 71,275 15,273 15,274 101,822 101,822 Temporal_Relative_QA 5,886 1,261 1,262 8,409 8,409 summary 4,366 935 937 6,238 6,238β¦ See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/Marco_Longspeech.
MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser.
Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it? Fineweb-Edu-Fortified is a dataset derived from Fineweb-Edu by applying exact-match deduplication across the whole dataset and producing an embedding for each row. The number of times the text from each row appears is also included as a count column. The embeddings were produced using TaylorAI/bge-micro Fineweb and⦠See the full description on the dataset page: https://huggingface.co/datasets/airtrain-ai/fineweb-edu-fortified.
Dataset Card for "qasc" Dataset Summary QASC is a question-answering dataset with a focus on sentence composition. It consists of 9,980 8-way multiple-choice questions about grade school science (8,134 train, 926 dev, 920 test), and comes with a corpus of 17M sentences. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of⦠See the full description on the dataset page: https://huggingface.co/datasets/allenai/qasc.
β οΈ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B β οΈ For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T. Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B. For this reason, please use ourβ¦ See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.
The CT-RATE Team organizes the VLM3D Challenge VLM3D 2026 (2nd Edition) β Challenge Finals at MICCAI 2026 VLM3D 2025 (1st Edition) β Challenge Finals at MICCAI 2025 β’ Workshop at ICCV 2025 The CT-RATE Team is developing the MR-RATE Dataset A large-scale brain MRI dataset with paired radiology reports for training 3D vision-language models. GitHub Β | Β Dataset Β | Β Metadata Dashboard Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomographyβ¦ See the full description on the dataset page: https://huggingface.co/datasets/ibrahimhamamci/CT-RATE.
MNBVC: Massive Never-ending BT Vast Chinese corpus
21cm Hydrogen Line Sky Survey Continuous 1-second-cadence power spectra of the 21cm neutral hydrogen (HI) line at 1420.405 MHz, recorded 24/7 from a fixed omnidirectional observer on the US East Coast. What this is A radio telescope pointed at the whole sky, recording one spectrum per second, indefinitely. The Earth's rotation scans the beam across the galactic plane daily, producing a natural drift scan. Every row is a self-timestamped power spectrum spanning 2β¦ See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/hi-21cm-survey.
Dataset Card for "uspto_full" More Information needed
Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here. Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation. Dataset Statistics Attribute Value Total Hours 100,405 Total Frames 10.8 billion Video Clips 2,010,759 Median Clip Length 180.0 seconds Mean⦠See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.
Ultra-FineWeb π Technical Report | π¦ UltraData Collection | π UltraData | π€ MiniCPM4 Series | π€ MiniCPM5 Series English | δΈζ π Introduction Ultra-FineWeb is a large-scale, high-quality, and efficiently-filtered dataset. We use the proposed efficient verification-based high-quality filtering pipeline to the FineWeb and Chinese FineWeb datasets (source data from Chinese FineWeb-edu-v2, which includes IndustryCorpus2, MiChao, WuDao, SkyPileβ¦ See the full description on the dataset page: https://huggingface.co/datasets/openbmb/Ultra-FineWeb.
Forecast News Deduplicated daily news corpus used by forecast-sim and future-sim. Snapshot 31,696,086 articles 3,401 daily partitions Coverage: 2016-08-26 through 2026-06-30 Snapshot published: 2026-07-13 Stored data size: approximately 168.8 GB The Parquet files are the canonical complete representation. The repository also contains daily JSONL files where available and compact headline JSON files used by article-browsing workflows. Layout Files⦠See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.
MICCAI FLARE25 Task 4: Foundation Models for 3D CT and MRI Scans (Homepage) This is the official dataset for CT image foundation model development. We provide 10,000+ CT scans for model pretraining. Downstream tasks include: Abdominal disease classification Abdominal lesion segmentation Abdominal organ segmentation Lung lesion segmentation Dataset Dataset Name Task Metric Source License Abdominal Disease Classification multi-label⦠See the full description on the dataset page: https://huggingface.co/datasets/FLARE-MedFM/FLARE-Task4-CT-FM.
Audio-FLAN Dataset (Paper) (the FULL audio files and jsonl files are still updating) An Instruction-Tuning Dataset for Unified Audio Understanding and Generation Across Speech, Music, and Sound. 1. Dataset Structure The Audio-FLAN-Dataset has the following directory structure: Audio-FLAN-Dataset/ βββ audio_files/ β βββ audio/ β β βββ 177_TAU_Urban_Acoustic_Scenes_2022/ β β βββ 179_Audioset_for_Audio_Inpainting/ β β βββ ... β βββ music/ β β ββββ¦ See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Audio-FLAN-Dataset.