Model Hub

Browse PQC-verified AI models, datasets, and tools

amphion/Emilia-Dataset HF Unverified

Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.

Task_categories:text-To-SpeechTask_categories:automatic-Speech-RecognitionLanguage:zhLanguage:enLanguage:jaLanguage:fr
JosephusCheung/GuanacoDataset HF Unverified

Sorry, it's no longer available on Hugging Face. Please reach out to those who have already downloaded it. If you have a copy, please refrain from re-uploading it to Hugging Face. The people here don't deserve it. See also: https://twitter.com/RealJosephus/status/1779913520529707387 GuanacoDataset News: We're heading towards multimodal VQA, with blip2-flan-t5-xxl Alignment to Guannaco 7B LLM. Still under construction: GuanacoVQA weight & GuanacoVQA Dataset Notice: Effective… See the full description on the dataset page: https://huggingface.co/datasets/JosephusCheung/GuanacoDataset.

Task_categories:text-GenerationTask_categories:question-AnsweringLanguage:zhLanguage:enLanguage:jaLanguage:de
D
deepset/xlm-roberta-large-squad2 HF Unverified

Question AnsweringTransformersPyTorchSafetensorsXlm-RobertaMultilingual HIGH
N
nvidia/GR00T-N1.7-3B HF Unverified

RoboticsSafetensorsGr00tN1d7 HIGH
baber/uspto_raw HF Unverified

Size_categories:10M<n<100MFormat:parquetModality:textLibrary:datasetsLibrary:daskLibrary:mlcroissant
opendatalab/Sci-Base HF Unverified

Sci-Base: The Largest AI-Ready Scientific Foundation Dataset 🌌 The Sciverse Data Foundation Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research. Sciverse… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Sci-Base.

Language:enSize_categories:1M<n<10MFormat:parquetModality:textLibrary:datasetsLibrary:dask
cot-leaderboard/cot-eval-traces-2.0 HF Unverified

Size_categories:1M<n<10MFormat:parquetModality:textLibrary:datasetsLibrary:daskLibrary:mlcroissant
D
deepset/xlm-roberta-base-squad2 HF Unverified

Question AnsweringTransformersPyTorchSafetensorsXlm-RobertaModel-Index MEDIUM
lishaoyong/latex-formulas-80M HF Unverified

For more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository. IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set

Size_categories:10M<n<100MFormat:parquetModality:imageModality:textLibrary:datasetsLibrary:dask
hotpotqa/hotpot_qa HF Unverified

Dataset Card for "hotpot_qa" Dataset Summary HotpotQA is a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) the questions require finding and reasoning over multiple supporting documents to answer; (2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas; (3) we provide sentence-level supporting facts required for reasoning, allowingQA systems to reason… See the full description on the dataset page: https://huggingface.co/datasets/hotpotqa/hotpot_qa.

Task_categories:question-AnsweringAnnotations_creators:crowdsourcedLanguage_creators:foundMultilinguality:monolingualSource_datasets:originalLanguage:en
K
KingTechnician/videomae-small-finetuned-kinetics-xd-violence-binary HF PQC Verified

Video-ClassificationTransformersSafetensorsVideomaeGenerated_from_trainerBase_model:MCG-NJU/videomae-Small-Finetuned-Kinetics MEDIUM
MLCommons/peoples_speech HF Unverified

Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.

Task_categories:automatic-Speech-RecognitionAnnotations_creators:crowdsourcedAnnotations_creators:machine-GeneratedLanguage_creators:crowdsourcedLanguage_creators:machine-GeneratedMultilinguality:monolingual
fancyzhx/ag_news HF Unverified

Dataset Card for "ag_news" Dataset Summary AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc), xml… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.

Task_categories:text-ClassificationTask_ids:topic-ClassificationAnnotations_creators:foundLanguage_creators:foundMultilinguality:monolingualSource_datasets:original
RichardErkhov/DASP HF Unverified

Dataset Card for DASP Dataset Description The DASP (Distributed Analysis of Sentinel-2 Pixels) dataset consists of cloud-free satellite images captured by Sentinel-2 satellites. Each image represents the most recent, non-partial, and cloudless capture from over 30 million Sentinel-2 images in every band. The dataset provides a near-complete cloudless view of Earth's surface, ideal for various geospatial applications. Images were converted from JPEG2000 to JPEG-XL to… See the full description on the dataset page: https://huggingface.co/datasets/RichardErkhov/DASP.

Task_categories:image-SegmentationTask_categories:image-ClassificationTask_categories:object-DetectionTask_categories:otherModality:geospatialSatellite-Imagery
nkp37/OpenVid-1M HF Unverified

Summary This is the dataset proposed in our paper [ICLR 2025] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation. OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets. All videos in the OpenVid-1M dataset have resolutions of at least 512×512.… See the full description on the dataset page: https://huggingface.co/datasets/nkp37/OpenVid-1M.

Task_categories:text-To-VideoTask_categories:image-To-VideoLanguage:enSize_categories:1M<n<10MFormat:csvModality:tabular
T
timpal0l/mdeberta-v3-base-squad2 HF Unverified

Question AnsweringTransformersPyTorchSafetensorsDeberta-V2Deberta HIGH
H
HumanCompatibleAI/ppo-seals-CartPole-v0 HF Unverified

Reinforcement-LearningStable-Baselines3Seals/CartPole-V0Deep-Reinforcement-LearningModel-Index MEDIUM
Skylion007/openwebtext HF Unverified

Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset files: 13.51 GB Size of the… See the full description on the dataset page: https://huggingface.co/datasets/Skylion007/openwebtext.

Task_categories:text-GenerationTask_categories:fill-MaskTask_ids:language-ModelingTask_ids:masked-Language-ModelingAnnotations_creators:no-AnnotationLanguage_creators:found
L
LG-AI-Research/EXAONE-Tabular HF Unverified

Tabular-ClassificationTabularTabular-RegressionIn-Context-LearningFoundation-ModelPyTorch MEDIUM
W
Wan-AI/Wan2.1-T2V-14B HF Unverified

Text-To-VideoDiffusersSafetensorsT2vVideo generationEnglish HIGH
Showing 20 of 898 items (page 37 of 45)