Model Hub

Browse PQC-verified AI models, datasets, and tools

M
microsoft/speecht5_tts HF Unverified

Text-To-SpeechTransformersPyTorchSpeecht5Text-To-AudioAudio MEDIUM
GokuScraper/seedance-2-prompts-datasets HF Unverified

🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.

Task_categories:text-To-VideoLanguage:enLanguage:zhSize_categories:1K<n<10KModality:imageModality:video
G
google/pegasus-xsum HF Unverified

SummarizationTransformersPyTorchTfJAXPegasus HIGH
B
black-forest-labs/FLUX.1-Fill-dev HF PQC Verified

DiffusersSafetensorsImage GenerationFluxDiffusion-Single-FileDiffusers:FluxFillPipeline HIGH
B
bosonai/higgs-audio-v2-generation-3B-base HF PQC Verified

Text-To-SpeechTransformersSafetensorsHiggs_audio_v2Text-To-Audio HIGH
M
mattmdjaga/segformer_b2_clothes HF Unverified

Image-SegmentationTransformersPyTorchONNXSafetensorsSegformer MEDIUM
C
cagliostrolab/animagine-xl-3.1 HF PQC Verified

Text-to-ImageDiffusersSafetensorsStable-DiffusionStable-Diffusion-XlBase_model:cagliostrolab/animagine-Xl-3.0 CRITICAL
D
depth-anything/DA3METRIC-LARGE HF PQC Verified

Depth-EstimationDepth-Anything-3SafetensorsComputer-VisionMonocular-DepthMulti-View-Geometry HIGH
HuggingFaceM4/FineVision HF Unverified

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision Load the data from datasets import load_dataset, get_dataset_config_names # Get all subset names and load the first one available_subsets =… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/FineVision.

Size_categories:10M<n<100MFormat:parquetModality:imageModality:textLibrary:datasetsLibrary:dask
CohereLabs/xP3x HF Unverified

Dataset Card for xP3x Dataset Summary xP3x (Crosslingual Public Pool of Prompts eXtended) is a collection of prompts & datasets across 277 languages & 16 NLP tasks. It contains all of xP3 + much more! It is used for training future contenders of mT0 & BLOOMZ at project Aya @Cohere Labs 🧡 Creation: The dataset can be recreated using instructions available here together with the file in this repository named xp3x_create.py. We provide this version to save processing… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/xP3x.

Task_categories:otherAnnotations_creators:expert-GeneratedAnnotations_creators:crowdsourcedMultilinguality:multilingualLanguage:afLanguage:ar
L
LiheYoung/depth-anything-large-hf HF PQC Verified

Depth-EstimationTransformersSafetensorsDepth_anythingVision HIGH
Kazimir-ai/text-to-image-prompts HF Unverified

The dataset of the most popular text-to-image prompts. Dataset Details Dataset Description Curated by: kazimir.ai Funded by [optional]: [More Information Needed] Shared by [optional]: https://kazimir.ai License: apache-2.0 Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Free to use. Dataset Structure CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Kazimir-ai/text-to-image-prompts.

Language:enSize_categories:10K<n<100KFormat:csvModality:textLibrary:datasetsLibrary:pandas
mvp-lab/LLaVA-OneVision-2-Data HF Unverified

LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used in mid-training. Dataset Composition Subset Format Description mid_training_video/60s_rest/ WebDataset (.tar) 10,809 shards of ~60s video clips mid_training_video/caption_v0/split_30s.jsonl JSONL Captions for 30-second video clips mid_training_video/caption_v0/split_60s.jsonl JSONL Captions for… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.

Task_categories:video-Text-To-TextTask_categories:visual-Question-AnsweringTask_categories:image-Text-To-TextLanguage:enSize_categories:n<1KFormat:parquet
HuggingFaceH4/MATH-500 HF Unverified

Dataset Card for MATH-500 This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits

Task_categories:text-GenerationLanguage:enSize_categories:n<1KFormat:jsonModality:textLibrary:datasets
U
unsloth/LTX-2.3-GGUF HF Unverified

Image-To-VideoGgmlGGUFUnslothText-To-VideoVideo-To-Video CRITICAL
J
jameslahm/yolov10s HF Unverified

Object-DetectionYolov10SafetensorsComputer-VisionPytorch_model_hub_mixin MEDIUM
T
tencent/HunyuanImage-3.0 HF Unverified

Text-to-ImageTransformersSafetensorsHunyuan_image_3_moeText GenerationCustom_code CRITICAL
G
griko/gender_cls_svm_ecapa_voxceleb HF Unverified

Audio-ClassificationJoblibGender-ClassificationSpeaker-CharacteristicsSpeaker-RecognitionVoice-Analysis MEDIUM
SWE-bench/SWE-smith HF Unverified

SWE-smith Dataset Code • Paper • Site [12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the `SWE-bench/SWE-smith-[lang]` datasets. For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo. The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit.… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-smith.

Task_categories:text-GenerationLanguage:enSize_categories:10K<n<100KFormat:parquetModality:textLibrary:datasets
RekaAI/RekaDaily-10k-raw HF Unverified

RekaDaily-10k (raw) Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are exactly as collected — no re-encoding, no cuts, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.

Task_categories:video-ClassificationTask_categories:image-To-VideoLanguage:enSize_categories:100K<n<1MFormat:parquetModality:image
Showing 20 of 896 items (page 26 of 45)