Model Hub

Browse PQC-verified AI models, datasets, and tools

jobs-git/Zyda-2 HF Unverified

Zyda-2 Zyda-2 is a 5 trillion token language modeling dataset created by collecting open and high quality datasets and combining them and cross-deduplication and model-based quality filtering. Zyda-2 comprises diverse sources of web data, highly educational content, math, code, and scientific papers. To construct Zyda-2, we took the best open-source datasets available: Zyda, FineWeb, DCLM, and Dolma. Models trained on Zyda-2 significantly outperform identical models trained on the… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/Zyda-2.

Task_categories:text-GenerationLanguage:enSize_categories:n>1T
wikimedia/Wikipedia (Nov 2023) HF PQC Verified

Complete Wikipedia dump across all languages. Standard pretraining data source. Structured articles with metadata.

DatasetTextMultilingualKnowledge CRITICAL
I
Intel/dpt-hybrid-midas HF PQC Verified

Depth-EstimationTransformersPyTorchDptVisionModel-Index MEDIUM
B
black-forest-labs/FLUX.1-schnell HF PQC Verified

Fastest FLUX variant. 4-step generation for near-instant high-quality images. Optimized for speed.

DiffusionImage Generation12BFast HIGH
mlfoundations/MINT-1T-PDF-CC-2023-06 HF PQC Verified

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-06.

Task_categories:image-To-TextTask_categories:text-GenerationLanguage:enSize_categories:100B<n<1TMultimodal
S
shi-labs/oneformer_ade20k_swin_tiny HF Unverified

Image-SegmentationTransformersPyTorchOneformerVision MEDIUM
G
google-t5/t5-large HF Unverified

TranslationTransformersPyTorchTfJAXSafetensors HIGH
F
facebook/mms-lid-256 HF Unverified

Audio-ClassificationTransformersPyTorchSafetensorsWav2vec2Mms HIGH
SWE-bench/SWE-bench_Multilingual HF Unverified

Language:enSize_categories:n<1KFormat:parquetModality:textLibrary:datasetsLibrary:pandas
D
depth-anything/DA3NESTED-GIANT-LARGE-1.1 HF Unverified

Depth-EstimationDepth-Anything-3SafetensorsComputer-VisionMonocular-DepthMulti-View-Geometry HIGH
D
deepset/tinyroberta-squad2 HF Unverified

Question AnsweringTransformersPyTorchSafetensorsRobertaModel-Index MEDIUM
P
playgroundai/playground-v2.5-1024px-aesthetic HF PQC Verified

Text-to-ImageDiffusersSafetensorsPlaygroundDiffusers:StableDiffusionXLPipeline CRITICAL
J
joeddav/xlm-roberta-large-xnli HF Unverified

Zero-Shot ClassificationTransformersPyTorchTfSafetensorsXlm-Roberta HIGH
N
nphSi/Z-Image-Lora HF Unverified

Text-to-ImageDiffusersLoraSafetensorsZ-ImageBase_model:Tongyi-MAI/Z-Image CRITICAL
C
cross-encoder/nli-MiniLM2-L6-H768 HF Unverified

Zero-Shot ClassificationSentence-TransformersPyTorchONNXSafetensorsOpenvino HIGH
espnet/yodas-granary HF Unverified

Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.

Task_categories:automatic-Speech-RecognitionTask_categories:translationLanguage:bgLanguage:csLanguage:daLanguage:de
M
microsoft/speecht5_tts HF Unverified

Text-To-SpeechTransformersPyTorchSpeecht5Text-To-AudioAudio MEDIUM
angie-chen55/javascript-github-code HF Unverified

Size_categories:10M<n<100MFormat:parquetModality:textLibrary:datasetsLibrary:daskLibrary:polars
B
black-forest-labs/FLUX.1-Fill-dev HF PQC Verified

DiffusersSafetensorsImage GenerationFluxDiffusion-Single-FileDiffusers:FluxFillPipeline HIGH
B
bosonai/higgs-audio-v2-generation-3B-base HF PQC Verified

Text-To-SpeechTransformersSafetensorsHiggs_audio_v2Text-To-Audio HIGH
Showing 20 of 738 items (page 21 of 37)