Model Hub

Browse PQC-verified AI models, datasets, and tools

B
black-forest-labs/FLUX.2-klein-base-4B HF PQC Verified

Image-To-ImageDiffusersSafetensorsText-to-ImageImage-EditingFlux HIGH
jhu-clsp/ettin-pretraining-data HF Unverified

Ettin Pre-training Data Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite. This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description DCLM 837.2 49.1% High-quality web crawl data CC Head 356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.

Task_categories:text-GenerationTask_categories:fill-MaskTask_categories:text-ClassificationLanguage:enPretrainingLanguage-Modeling
S
speechbrain/lang-id-voxlingua107-ecapa HF Unverified

Audio-ClassificationSpeechbrainEmbeddingsLanguageIdentificationPyTorch MEDIUM
HuggingFaceFW/finepdfs_lang_classification HF Unverified

Size_categories:1M<n<10MFormat:parquetModality:tabularLibrary:datasetsLibrary:pandasLibrary:mlcroissant
S
stabilityai/stable-diffusion-3.5-medium HF Unverified

Text-to-ImageDiffusersSafetensorsStable-DiffusionDiffusers:StableDiffusion3PipelineEnglish CRITICAL
R
RunDiffusion/Juggernaut-XL-v9 HF Unverified

Text-to-ImageDiffusersStable-DiffusionStable-Diffusion-XlSdxlPhotorealistic CRITICAL
K
keremberke/yolov8m-table-extraction HF Unverified

Object-DetectionUltralyticsTensorboardV8UltralyticsplusYolov8 MEDIUM
abisee/cnn_dailymail HF Unverified

Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.

Task_categories:summarizationTask_ids:news-Articles-SummarizationAnnotations_creators:no-AnnotationLanguage_creators:foundMultilinguality:monolingualSource_datasets:original
S
stable-diffusion-v1-5/stable-diffusion-inpainting HF PQC Verified

Text-to-ImageDiffusersStable-DiffusionStable-Diffusion-DiffusersDiffusers:StableDiffusionInpaintPipeline CRITICAL
jasperai/monet HF Unverified

Dataset Card for MONET MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 104.9 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with… See the full description on the dataset page: https://huggingface.co/datasets/jasperai/monet.

Task_categories:text-To-ImageTask_categories:image-Feature-ExtractionTask_categories:zero-Shot-Image-ClassificationLanguage:enSize_categories:100M<n<1BMultimodal
stanford-vision-lab/gpic HF Unverified

GPIC: A Giant Permissive Image Corpus for Visual Generation Keshigeyan&nbsp;Chandrasegaran*1,&nbsp; Kyle&nbsp;Sargent*1,&nbsp; Suchir&nbsp;Agarwal1,&nbsp; Michael&nbsp;Jang1,&nbsp; Michael&nbsp;Poli1,2,&nbsp; Juan&nbsp;Carlos&nbsp;Niebles1,4,&nbsp; Justin&nbsp;Johnson3,&nbsp; Jiajun&nbsp;Wu1,&nbsp; Li&nbsp;Fei-Fei1 1&nbsp;Stanford University&nbsp;&nbsp; 2&nbsp;Radical Numerics&nbsp;&nbsp; 3&nbsp;University of Michigan&nbsp;&nbsp; 4&nbsp;Salesforce… See the full description on the dataset page: https://huggingface.co/datasets/stanford-vision-lab/gpic.

Language:en
labelmaker/arkit_labelmaker HF Unverified

ARKit Labelmaker: A New Scale for Indoor 3D Scene Understanding [arxiv] [website] [checkpoints] [code] We complement ARKitScenes dataset with dense semantic annotations that are automatically generated at scale. This produces the first large-scale, real-world 3D dataset with dense semantic annotations. Training on this auto-generated data, we push forward the state-of-the-art performance on ScanNet and ScanNet200 with prevalent 3D semantic segmentation models.

Task_categories:image-SegmentationLanguage:enSize_categories:1K<n<10KDoi:10.57967/hf/23893D semantic segmentationIndoor 3D scene dataset
L
LyliaEngine/Pony_Diffusion_V6_XL HF Unverified

Text-to-ImageDiffusersStable-DiffusionLoraTemplate:sd-LoraBase_model:Bakanayatsu/Pony-Diffusion-V6-XL-For-Anime HIGH
F
facebook/mask2former-swin-tiny-coco-instance HF Unverified

Image-SegmentationTransformersPyTorchSafetensorsMask2formerVision MEDIUM
M
MoritzLaurer/DeBERTa-v3-large-mnli-fever-anli-ling-wanli HF Unverified

Zero-Shot ClassificationTransformersPyTorchONNXSafetensorsDeberta-V2 HIGH
aline-gassenn/MedDialog-Audio HF Unverified

MedDialogue-Audio English Medical Dialogue Corpus for Speech Recognition Research. This repository contains MedDialogue-Audio, an English audio corpus designed for research in Automatic Speech Recognition (ASR) in the healthcare domain. The dataset was published in the proceedings of the 7th SBBD Dataset Showcase Workshop, and is available online at the following link: https://sol.sbc.org.br/index.php/dsw/article/view/37199 Dataset Description MedDialogue-Audio is… See the full description on the dataset page: https://huggingface.co/datasets/aline-gassenn/MedDialog-Audio.

Task_categories:automatic-Speech-RecognitionLanguage:enSize_categories:100K<n<1MDoi:10.57967/hf/5889Medical
Williamsanderson/MedQA-Darija-MultiLingual HF Unverified

MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.

Task_categories:question-AnsweringTask_categories:automatic-Speech-RecognitionTask_categories:text-To-SpeechLanguage:arLanguage:frLanguage:en
K
kpsss34/FHDR_Uncensored HF PQC Verified

Text-to-ImageDiffusersSafetensorsGGUFArtBase_model:black-Forest-Labs/FLUX.1-Dev CRITICAL
H
h94/IP-Adapter-FaceID HF PQC Verified

Text-to-ImageDiffusersStable-DiffusionEnglish HIGH
CERN/ColliderML-Release-1 HF Unverified

ColliderML: Dataset Release 1 Dataset Description This dataset contains simulated high-energy physics collision events generated using the Open Data Detector (ODD) geometry within the Key4hep and ACTS (A Common Tracking Software) frameworks, representing a generic collider detector similar to those at the HL-LHC. Dataset Summary Collision Energy: 14 TeV (proton-proton) Detector: Open Data Detector (ODD) Simulation: DD4hep + Geant4 + ACTS Format: Apache Parquet… See the full description on the dataset page: https://huggingface.co/datasets/CERN/ColliderML-Release-1.

Task_categories:otherSize_categories:10M<n<100MFormat:parquetModality:timeseriesLibrary:datasetsLibrary:dask
Showing 20 of 738 items (page 23 of 37)