Model Hub

Browse PQC-verified AI models, datasets, and tools

D
deepset/tinyroberta-squad2 HF Unverified

Question AnsweringTransformersPyTorchSafetensorsRobertaModel-Index MEDIUM
M
MoritzLaurer/DeBERTa-v3-large-mnli-fever-anli-ling-wanli HF Unverified

Zero-Shot ClassificationTransformersPyTorchONNXSafetensorsDeberta-V2 HIGH
S
Synthefy/Nori-30M HF Unverified

Tabular-RegressionSynthefy-NoriFeatures-TransformerTabularTabular-Foundation-ModelIn-Context-Learning MEDIUM
H
hustvl/yolos-tiny HF Unverified

Object-DetectionTransformersPyTorchSafetensorsYolosVision MEDIUM
mlfoundations/MINT-1T-PDF-CC-2024-18 HF PQC Verified

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-18.

Task_categories:image-To-TextTask_categories:text-GenerationLanguage:enSize_categories:100B<n<1TMultimodal
allenai/MADLAD-400 HF PQC Verified

MADLAD-400 Dataset and Introduction MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is a document-level multilingual dataset based on Common Crawl, covering 419 languages in total. This uses all snapshots of CommonCrawl available as of August 1, 2022. The primary advantage of this dataset over similar datasets is that it is more multilingual (419 languages), it is audited and more highly filtered, and it is document-level. The main disadvantage… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MADLAD-400.

Task_categories:text-GenerationSize_categories:n>1T
M
monologg/koelectra-small-v2-distilled-korquad-384 HF Unverified

Question AnsweringTransformersPyTorchTfliteSafetensorsElectra MEDIUM
M
MoritzLaurer/deberta-v3-large-zeroshot-v2.0 HF Unverified

Zero-Shot ClassificationTransformersONNXSafetensorsDeberta-V2Text Classification HIGH
D
depth-anything/DA3NESTED-GIANT-LARGE-1.1 HF Unverified

Depth-EstimationDepth-Anything-3SafetensorsComputer-VisionMonocular-DepthMulti-View-Geometry HIGH
X
Xenova/segformer-b0-finetuned-ade-512-512 HF Unverified

Image-SegmentationTransformers.jsONNXSegformerBase_model:nvidia/segformer-B0-Finetuned-Ade-512-512Base_model:quantized:nvidia/segformer-B0-Finetuned-Ade-512-512 MEDIUM
P
PekingU/rtdetr_v2_r50vd HF Unverified

Object-DetectionTransformersSafetensorsRt_detr_v2VisionEnglish MEDIUM
Tuxifan/UbuntuIRC HF Unverified

Completely uncurated collection of IRC logs from the Ubuntu IRC channels

Task_categories:text-GenerationSize_categories:1M<n<10MFormat:textModality:textLibrary:datasetsLibrary:mlcroissant
labofsahil/pypi-packages-metadata-dataset HF Unverified

Size_categories:10M<n<100MModality:text
D
depth-anything/Depth-Anything-V2-Base-hf HF Unverified

Depth-EstimationTransformersSafetensorsDepth_anythingDepthRelative depth MEDIUM
IFM/MegaMath HF Unverified

MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/IFM/MegaMath.

Task_categories:text-GenerationLanguage:enSize_categories:100M<n<1BFormat:parquetModality:textLibrary:datasets
mlfoundations/MINT-1T-PDF-CC-2023-23 HF PQC Verified

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.

Task_categories:image-To-TextTask_categories:text-GenerationLanguage:enSize_categories:1M<n<10MFormat:webdatasetModality:image
ceval/ceval-exam HF Unverified

C-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details. Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model… See the full description on the dataset page: https://huggingface.co/datasets/ceval/ceval-exam.

Task_categories:text-ClassificationTask_categories:multiple-ChoiceTask_categories:question-AnsweringLanguage:zhSize_categories:10K<n<100KFormat:parquet
PresentBench/PresentBench HF Unverified

PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation [🌐 Homepage] [📖 Paper] [💻 Code] This repository hosts the PresentBench benchmark dataset. 📄 Abstract Slides serve as a critical medium for conveying information in presentation-oriented scenarios such as academia, education, and business. Despite their importance, creating high-quality slide decks remains time-consuming and cognitively demanding. Recent advances in generative models, such… See the full description on the dataset page: https://huggingface.co/datasets/PresentBench/PresentBench.

Task_categories:any-To-AnyTask_categories:text-GenerationLanguage:enLanguage:zhSize_categories:n<1KFormat:json
J
jonathandinu/face-parsing HF Unverified

Image-SegmentationTransformersPyTorchONNXSafetensorsSegformer HIGH
mandarjoshi/trivia_qa HF Unverified

Dataset Card for "trivia_qa" Dataset Summary TriviaqQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaqQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/mandarjoshi/trivia_qa.

Task_categories:question-AnsweringTask_ids:open-Domain-QaTask_ids:open-Domain-Abstractive-QaTask_ids:extractive-QaTask_ids:abstractive-QaAnnotations_creators:crowdsourced
Showing 20 of 898 items (page 30 of 45)