Model Hub

Browse PQC-verified AI models, datasets, and tools

jat-project/jat-dataset HF Unverified

JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset = load_dataset("jat-project/jat-dataset"… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.

Task_categories:reinforcement-LearningTask_categories:text-GenerationTask_categories:question-AnsweringAnnotations_creators:foundAnnotations_creators:machine-GeneratedSource_datasets:conceptual-Captions
B
black-forest-labs/FLUX.2-klein-4B HF PQC Verified

Image-To-ImageDiffusersSafetensorsText-to-ImageImage-EditingFlux HIGH
Benjy/typed_digital_signatures HF PQC Verified

Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size: ~90,000… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.

Task_categories:image-ClassificationTask_categories:zero-Shot-Image-ClassificationTask_categories:image-Feature-ExtractionLanguage:enSize_categories:10K<n<100KModality:image
nvidia/SAGE-10k HF Unverified

SAGE-10k SAGE-10k is a large-scale interactive indoor scene dataset featuring realistic layouts, generated by the agentic-driven pipeline introduced in "SAGE: Scalable Agentic 3D Scene Generation for Embodied AI". The dataset contains 10,000 diverse scenes spanning 50 room types and styles, along with 565K uniquely generated 3D objects. 🔑 Key Features SAGE-10k integrates a wide variety of scenes, and particularly, preserves small items for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SAGE-10k.

Task_categories:text-To-3dLanguage:enSize_categories:10K<n<100KScene-GenerationInteractive-ScenesEmbodied-AI
M
MCG-NJU/videomae-base HF Unverified

Video-ClassificationTransformersPyTorchSafetensorsVideomaePretraining MEDIUM
HuggingFaceFW/FineWeb HF PQC Verified

15T token dataset of cleaned English web data. Deduplicated and filtered from CommonCrawl, outperforms C4 and RefinedWeb for LLM pretraining.

DatasetPretrainingEnglish15T tokens CRITICAL
allenai/winogrande HF Unverified

Dataset Card for "winogrande" Dataset Summary WinoGrande is a new collection of 44k problems, inspired by Winograd Schema Challenge (Levesque, Davis, and Morgenstern 2011), but adjusted to improve the scale and robustness against the dataset-specific bias. Formulated as a fill-in-a-blank task with binary options, the goal is to choose the right option for a given sentence which requires commonsense reasoning. Supported Tasks and Leaderboards More Information… See the full description on the dataset page: https://huggingface.co/datasets/allenai/winogrande.

Language:enSize_categories:10K<n<100KFormat:parquetModality:textLibrary:datasetsLibrary:pandas
L
Leo97/KoELECTRA-small-v3-modu-ner HF Unverified

Token ClassificationTransformersPyTorchTensorboardSafetensorsElectra MEDIUM
H
Helsinki-NLP/opus-mt-de-en HF Unverified

TranslationTransformersPyTorchTfRustMarian HIGH
B
bosonai/higgs-tts-2-3b-base HF Unverified

Text-To-SpeechTransformersSafetensorsHiggs_audio_v2Text-To-Audio HIGH
robbyant/mdm_depth HF Unverified

LingBot-Depth Dataset Self-curated RGB-D dataset for training LingBot-Depth, a masked depth modeling approach (arxiv:2601.17895). Each sample contains an RGB image, raw sensor depth, and ground truth depth. Total size: 2.71 TBDepth scale: millimeters (mm), stored as 16-bit PNGLicense: CC BY-NC-SA 4.0 Sub-datasets Name Description Samples RobbyReal Real-world indoor scenes captured with multiple RGB-D cameras 1,400,000 RobbyVla Real-world data collected… See the full description on the dataset page: https://huggingface.co/datasets/robbyant/mdm_depth.

Task_categories:depth-EstimationLanguage:enModality:3d3D3dDepth
mlfoundations/MINT-1T-PDF-CC-2023-40 HF PQC Verified

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-40.

Task_categories:image-To-TextTask_categories:text-GenerationLanguage:enSize_categories:100B<n<1TMultimodal
L
larryvrh/MiniMax-H3-Turbo-Lora HF Unverified

Text-To-VideoMinimax-H3Text-To-AudioAudio-VideoLoraComfyui CRITICAL
O
OpenMOSS-Team/MOSS-TTS-v1.5 HF Unverified

Text-To-SpeechSafetensorsMoss_tts_delayCustom_codeYue HIGH
allenai/sciq HF Unverified

Dataset Card for "sciq" Dataset Summary The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/sciq.

Task_categories:question-AnsweringTask_ids:closed-Domain-QaAnnotations_creators:no-AnnotationLanguage_creators:crowdsourcedMultilinguality:monolingualSource_datasets:original
W
webAI-Official/TwIL-LM3 HF Unverified

Text GenerationTransformersSafetensorsGGUFSmollm3Formal-Logic HIGH
F
fishaudio/s2-pro HF Unverified

Text-To-SpeechSafetensorsFish_qwen3_omniInstruction-FollowingMultilingualJw HIGH
Q
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign HF PQC Verified

Text-To-SpeechQwen-TtsSafetensorsQwen3_ttsAudioTts HIGH
Salesforce/GiftEvalPretrain HF Unverified

GIFT-Eval Pre-training Datasets Pretraining dataset aligned with GIFT-Eval that has 71 univariate and 17 multivariate datasets, spanning seven domains and 13 frequencies, totaling 4.5 million time series and 230 billion data points. Notably this collection of data has no leakage issue with the train/test split and can be used to pretrain foundation models that can be fairly evaluated on GIFT-Eval. 📄 Paper 🖥️ Code 📔 Blog Post 🏎️ Leader Board Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/GiftEvalPretrain.

Task_categories:time-Series-ForecastingSize_categories:1M<n<10MModality:timeseriesTimeseriesForecastingBenchmark
A
ai4bharat/indic-parler-tts HF PQC Verified

Text-To-SpeechTransformersSafetensorsParler_ttsText GenerationAnnotation HIGH
Showing 20 of 896 items (page 21 of 45)