Model Hub

Browse PQC-verified AI models, datasets, and tools

B
bigcode/starcoder2-15b HF Ollama PQC Verified

Code LLM trained on The Stack v2 with 600+ programming languages. 4x the training data of StarCoder1.

TransformerCode Generation15B HIGH
S
StanfordAIMI/stanford-deidentifier-base HF PQC Verified

Token ClassificationTransformersPyTorchBERTSequence-Tagger-ModelPubmedbert MEDIUM
R
rizvandwiki/gender-classification HF PQC Verified

Image-ClassificationTransformersPyTorchTensorboardSafetensorsVit HIGH
M
mudler/locate-anything.cpp-gguf HF Unverified

Object-DetectionGGUFLocate-Anything.cppGgmlOpen-Vocabulary-DetectionVisual-Grounding HIGH
D
dslim/bert-base-NER HF PQC Verified

Token ClassificationTransformersPyTorchTfJAXONNX HIGH
P
pyannote/voice-activity-detection HF PQC Verified

Speech RecognitionPyannote-AudioPyannotePyannote-Audio-PipelineAudioVoice MEDIUM
allenai/c4 HF PQC Verified

C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one per… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.

Task_categories:text-GenerationTask_categories:fill-MaskTask_ids:language-ModelingTask_ids:masked-Language-ModelingAnnotations_creators:no-AnnotationLanguage_creators:found
F
facebook/bart-large-cnn HF PQC Verified

SummarizationTransformersPyTorchTfJAXRust HIGH
K
k2-fsa/OmniVoice HF PQC Verified

Text-To-SpeechOmnivoiceSafetensorsZero-ShotMultilingualVoice-Cloning HIGH
A
autogluon/mitra-regressor HF Unverified

Tabular-RegressionSafetensors MEDIUM
openai/gsm8k HF PQC Verified

Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to reach the… See the full description on the dataset page: https://huggingface.co/datasets/openai/gsm8k.

Benchmark:officialBenchmark:eval-YamlTask_categories:text-GenerationAnnotations_creators:crowdsourcedLanguage_creators:crowdsourcedMultilinguality:monolingual
LLM360/TxT360 HF PQC Verified

TxT360: A Top-Quality LLM Pre-training Dataset Requires the Perfect Blend Changelog Version Details v1.1 Added new data sources: TxT360_BestOfWeb, TxT360_QA, europarl-aligned, and wikipedia_extended. Details of v1.1 Additions TxT360_BestOfWeb: This is a filtered version of the TxT360 dataset, created using the ProX document filtering model. The model is similar to the FineWeb-Edu classifier, but also assigns an additional format score that… See the full description on the dataset page: https://huggingface.co/datasets/LLM360/TxT360.

Task_categories:text-GenerationLanguage:enSize_categories:n>1T
T
theainerd/Wav2Vec2-large-xlsr-hindi HF PQC Verified

Speech RecognitionTransformersPyTorchSafetensorsWav2vec2Base_model:facebook/wav2vec2-Large-Xlsr-53 HIGH
S
Salesforce/SFR-Embedding-2_R HF PQC Verified

State-of-the-art text embedding model. Top of MTEB leaderboard with strong retrieval and clustering.

TransformerEmbeddings7BRetrieval HIGH
fineinstructions/fineinstructions_nemotron HF Unverified

✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline. The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details. Each .parquet file in the data folderhas a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/fineinstructions/fineinstructions_nemotron.

Language:enSize_categories:1B<n<10BFormat:parquetModality:tabularModality:textLibrary:datasets
L
lucas-leme/FinBERT-PT-BR HF PQC Verified

Text ClassificationTransformersPyTorchBERTPt MEDIUM
S
Systran/faster-whisper-tiny.en HF PQC Verified

Speech RecognitionCtranslate2AudioEnglish MEDIUM
D
dima806/fairface_age_image_detection HF PQC Verified

Image-ClassificationTransformersSafetensorsVitBase_model:google/vit-Base-Patch16-224-In21kBase_model:finetune:google/vit-Base-Patch16-224-In21k HIGH
F
facebook/mms-1b-all HF PQC Verified

Speech RecognitionTransformersPyTorchSafetensorsWav2vec2Mms HIGH
N
nlptown/bert-base-multilingual-uncased-sentiment HF PQC Verified

Text ClassificationTransformersPyTorchTfJAXSafetensors HIGH
Showing 20 of 896 items (page 14 of 45)