Model Hub

Browse PQC-verified AI models, datasets, and tools

fancyzhx/ag_news HF Unverified

Dataset Card for "ag_news" Dataset Summary AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc), xml… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.

Task_categories:text-ClassificationTask_ids:topic-ClassificationAnnotations_creators:foundLanguage_creators:foundMultilinguality:monolingualSource_datasets:original
P
PekingU/rtdetr_r18vd_coco_o365 HF PQC Verified

Object-DetectionTransformersSafetensorsRt_detrVisionEnglish MEDIUM
zekaiwang/trex_dataset HF Unverified

T-Rex Dataset A large-scale, tactile-reactive bimanual manipulation dataset, collected via teleoperation on a Dexmate Vega-1 robot with two Sharpa Wave dexterous hands. Stored as a LeRobotDataset v3.0. 🌐 Project Page Β· ✍️ Paper (arXiv) Β· πŸ’» Code (T-Rex) Β· πŸš€ Dataset Quickstart Β· πŸ““ Colab notebook One episode from each of 20 motor primitives (head-camera view, cropped to the workspace), each with a different object. Teleoperation setup: Manus gloves + VIVE… See the full description on the dataset page: https://huggingface.co/datasets/zekaiwang/trex_dataset.

Task_categories:roboticsLanguage:enSize_categories:1M<n<10MFormat:parquetModality:tabularModality:text
A
abhishtagatya/hubert-base-960h-itw-deepfake HF Unverified

Audio-ClassificationTransformersTensorboardSafetensorsHubertDeepfake MEDIUM
Zyphra/Zyda-2 HF Unverified

Zyda-2 Zyda-2 is a 5 trillion token language modeling dataset created by collecting open and high quality datasets and combining them and cross-deduplication and model-based quality filtering. Zyda-2 comprises diverse sources of web data, highly educational content, math, code, and scientific papers. To construct Zyda-2, we took the best open-source datasets available: Zyda, FineWeb, DCLM, and Dolma. Models trained on Zyda-2 significantly outperform identical models trained on the… See the full description on the dataset page: https://huggingface.co/datasets/Zyphra/Zyda-2.

Task_categories:text-GenerationLanguage:enSize_categories:n>1T
P
philschmid/bart-large-cnn-samsum HF Unverified

SummarizationTransformersPyTorchBartText2text-GenerationSagemaker HIGH
mlfoundations/MINT-1T-PDF-CC-2023-14 HF PQC Verified

πŸƒ MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens πŸƒ MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. πŸƒ MINT-1T is designed to facilitate research in multimodal pretraining. πŸƒ MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.

Task_categories:image-To-TextTask_categories:text-GenerationLanguage:enSize_categories:1M<n<10MFormat:webdatasetModality:image
S
Synthefy/Nori HF Unverified

Tabular-RegressionSynthefy-NoriFeatures-TransformerTabularTabular-Foundation-ModelIn-Context-Learning MEDIUM
ATH-MaaS/Marco_Longspeech HF Unverified

Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks designed to benchmark Large Language Models on lengthy audio inputs. πŸ“Š Dataset Statistics Task Statistics Task Train Val Test Total Unique Audios ASR 71,275 15,273 15,274 101,822 101,822 Temporal_Relative_QA 5,886 1,261 1,262 8,409 8,409 summary 4,366 935 937 6,238 6,238… See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/Marco_Longspeech.

Task_categories:automatic-Speech-RecognitionTask_categories:audio-ClassificationTask_categories:text-GenerationLanguage:enLanguage:zhSize_categories:10K<n<100K
schwein69/hagrid-subset HF Unverified

HaGRID Gesture Recognition Subset Dataset Description A curated subset of the HaGRID (Hand Gesture Recognition Image Dataset) containing 24 gesture classes for training gesture recognition models. Dataset Summary Total Images: 19,200 Gesture Classes: 24 Samples per Class: 800 Image Format: JPEG Average Image Size: ~302 KB Splits Split Images Percentage Train 14,592 76% Val 1,728 9% Test 2,880 15% Gesture Classes call… See the full description on the dataset page: https://huggingface.co/datasets/schwein69/hagrid-subset.

Task_categories:image-ClassificationTask_categories:object-DetectionSize_categories:10K<n<100KGesture-RecognitionComputer-VisionHand-Gestures
airtrain-ai/fineweb-edu-fortified HF Unverified

Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it? Fineweb-Edu-Fortified is a dataset derived from Fineweb-Edu by applying exact-match deduplication across the whole dataset and producing an embedding for each row. The number of times the text from each row appears is also included as a count column. The embeddings were produced using TaylorAI/bge-micro Fineweb and… See the full description on the dataset page: https://huggingface.co/datasets/airtrain-ai/fineweb-edu-fortified.

Task_categories:text-GenerationLanguage:enSize_categories:100M<n<1BFormat:parquetModality:tabularModality:text
D
depth-anything/DA3-LARGE HF Unverified

Depth-EstimationDepth-Anything-3SafetensorsComputer-VisionMonocular-DepthMulti-View-Geometry HIGH
liwu/MNBVC HF Unverified

MNBVC: Massive Never-ending BT Vast Chinese corpus

Task_categories:text-GenerationTask_categories:fill-MaskTask_ids:language-ModelingTask_ids:masked-Language-ModelingAnnotations_creators:otherLanguage_creators:other
phanerozoic/hi-21cm-survey HF Unverified

21cm Hydrogen Line Sky Survey Continuous 1-second-cadence power spectra of the 21cm neutral hydrogen (HI) line at 1420.405 MHz, recorded 24/7 from a fixed omnidirectional observer on the US East Coast. What this is A radio telescope pointed at the whole sky, recording one spectrum per second, indefinitely. The Earth's rotation scans the beam across the galactic plane daily, producing a natural drift scan. Every row is a self-timestamped power spectrum spanning 2… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/hi-21cm-survey.

Task_categories:time-Series-ForecastingSize_categories:1M<n<10MFormat:parquetModality:tabularModality:textLibrary:datasets
L
LiheYoung/depth_anything_vitb14 HF Unverified

Depth-EstimationTransformersPyTorchDepth_anything MEDIUM
ibrahimhamamci/CT-RATE HF Unverified

The CT-RATE Team organizes the VLM3D Challenge VLM3D 2026 (2nd Edition) β†’ Challenge Finals at MICCAI 2026 VLM3D 2025 (1st Edition) β†’ Challenge Finals at MICCAI 2025 β€’ Workshop at ICCV 2025 The CT-RATE Team is developing the MR-RATE Dataset A large-scale brain MRI dataset with paired radiology reports for training 3D vision-language models. GitHub Β  | Β  Dataset Β  | Β  Metadata Dashboard Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimhamamci/CT-RATE.

Task_categories:image-To-TextTask_categories:text-To-ImageTask_categories:image-ClassificationTask_categories:question-AnsweringTask_categories:visual-Question-AnsweringTask_categories:zero-Shot-Classification
chenxran/uspto_full HF Unverified

Dataset Card for "uspto_full" More Information needed

Size_categories:1M<n<10MFormat:parquetModality:textLibrary:datasetsLibrary:pandasLibrary:mlcroissant
ILSVRC/imagenet-1k HF Unverified

Dataset Card for ImageNet Dataset Summary ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are… See the full description on the dataset page: https://huggingface.co/datasets/ILSVRC/imagenet-1k.

Task_categories:image-ClassificationTask_ids:multi-Class-Image-ClassificationAnnotations_creators:crowdsourcedLanguage_creators:crowdsourcedMultilinguality:monolingualSource_datasets:original
FLARE-MedFM/FLARE-Task4-CT-FM HF Unverified

MICCAI FLARE25 Task 4: Foundation Models for 3D CT and MRI Scans (Homepage) This is the official dataset for CT image foundation model development. We provide 10,000+ CT scans for model pretraining. Downstream tasks include: Abdominal disease classification Abdominal lesion segmentation Abdominal organ segmentation Lung lesion segmentation Dataset Dataset Name Task Metric Source License Abdominal Disease Classification multi-label… See the full description on the dataset page: https://huggingface.co/datasets/FLARE-MedFM/FLARE-Task4-CT-FM.

Task_categories:image-ClassificationTask_categories:image-SegmentationLanguage:enMedical
HKUSTAudio/Audio-FLAN-Dataset HF Unverified

Audio-FLAN Dataset (Paper) (the FULL audio files and jsonl files are still updating) An Instruction-Tuning Dataset for Unified Audio Understanding and Generation Across Speech, Music, and Sound. 1. Dataset Structure The Audio-FLAN-Dataset has the following directory structure: Audio-FLAN-Dataset/ β”œβ”€β”€ audio_files/ β”‚ β”œβ”€β”€ audio/ β”‚ β”‚ └── 177_TAU_Urban_Acoustic_Scenes_2022/ β”‚ β”‚ └── 179_Audioset_for_Audio_Inpainting/ β”‚ β”‚ └── ... β”‚ β”œβ”€β”€ music/ β”‚ β”‚ └──… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Audio-FLAN-Dataset.

Task_categories:text-To-SpeechTask_categories:text-To-AudioTask_categories:automatic-Speech-RecognitionLanguage:enLanguage:zhSize_categories:10M<n<100M
Showing 20 of 738 items (page 27 of 37)