Model Hub

Browse PQC-verified AI models, datasets, and tools

locuslab/TOFU HF Unverified

TOFU: Task of Fictitious Unlearning 🍢 The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set. Quick Links Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/locuslab/TOFU.

Task_categories:question-AnsweringTask_ids:closed-Domain-QaAnnotations_creators:machine-GeneratedLanguage_creators:machine-GeneratedMultilinguality:monolingualSource_datasets:original
bluuebunny/arxiv_metadata_by_year HF Unverified

Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/arxiv_metadata_by_year.

Language:enSize_categories:1M<n<10MFormat:parquetModality:textLibrary:datasetsLibrary:dask
roneneldan/TinyStories HF Unverified

Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.

Task_categories:text-GenerationLanguage:enSize_categories:1M<n<10MFormat:parquetModality:textLibrary:datasets
open-index/arctic HF Unverified

Arctic Shift Reddit Archive Every Reddit comment and submission since 2005, organized as monthly Parquet shards What is it? The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02. Right now the archive has 15.7B items (12.9B comments, 2.8B submissions) in 1.3 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly… See the full description on the dataset page: https://huggingface.co/datasets/open-index/arctic.

Task_categories:text-GenerationTask_categories:text-ClassificationTask_categories:feature-ExtractionLanguage:enSize_categories:1B<n<10BReddit
PEBE691445/sat-image-boundingbox-sft-full HF Unverified

NU-TONIC raw SFT Full Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters. Provenance Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2) Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8. Labels:… See the full description on the dataset page: https://huggingface.co/datasets/PEBE691445/sat-image-boundingbox-sft-full.

Task_categories:image-Text-To-TextLanguage:enSize_categories:1M<n<10MModality:geospatialSatelliteLand-Cover
HuggingFaceM4/FineVisionMax HF Unverified

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision The version in this repository concatenated all the configs in the original dataset and then shuffled them. This is done to facilitate streaming the data directly from the hub! Load… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/FineVisionMax.

Task_categories:image-Text-To-TextLanguage:enLanguage:zhSize_categories:10M<n<100MFormat:parquetFormat:optimized-Parquet
bio-nlp-umass/MedThinkVQA HF Unverified

MedThinkVQA MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning. Links GitHub: https://github.com/benluwang/MedThinkVQA Leaderboard: https://benluwang.github.io/MedThinkVQA/ Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.

Task_categories:question-AnsweringTask_categories:text-GenerationLanguage:enSize_categories:1K<n<10KFormat:parquetModality:image
facebook/belebele HF Unverified

The Belebele Benchmark for Massively Multilingual NLU Evaluation Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. This dataset enables the evaluation of mono- and multi-lingual models in high-, medium-, and low-resource languages. Each question has four multiple-choice answers and is linked to a short passage from the FLORES-200 dataset. The human annotation procedure was carefully curated to create questions that discriminate… See the full description on the dataset page: https://huggingface.co/datasets/facebook/belebele.

Task_categories:question-AnsweringTask_categories:zero-Shot-ClassificationTask_categories:text-ClassificationTask_categories:multiple-ChoiceLanguage:afLanguage:am
DeepAuto-AI/MacroLens HF Unverified

MacroLens A benchmarking corpus for contextual financial reasoning under macroeconomic scenarios across 4,416 U.S. small- and micro-cap equities (2021-01-04 — 2026-03-31). MacroLens unifies seven tasks over a single point-in-time panel: contextual time-series forecasting, public valuation, financial-statement generation, scenario-conditioned return forecasting, private-company valuation, generator evaluation from natural-language descriptions, and real-estate valuation. Task… See the full description on the dataset page: https://huggingface.co/datasets/DeepAuto-AI/MacroLens.

Task_categories:time-Series-ForecastingTask_categories:tabular-RegressionTask_categories:text-GenerationTask_categories:question-AnsweringLanguage:enSize_categories:1M<n<10M
M
monologg/koelectra-base-v3-finetuned-korquad HF Unverified

Question AnsweringTransformersPyTorchSafetensorsElectra MEDIUM
fixie-ai/covost2 HF Unverified

This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer. The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger. As such, not all the data is included: Only the validation and test subsets are available. From the XX_EN subsets, only fr, es, and zh-CN are included.

Size_categories:1M<n<10MFormat:parquetModality:audioModality:textLibrary:datasetsLibrary:dask
nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1 HF Unverified

Introduction TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments. Project page: https://dreamdojo-world.github.io/ Paper: https://arxiv.org/abs/2602.06949 Code: https://github.com/NVIDIA/DreamDojo How to Use Check out https://github.com/NVIDIA/DreamDojo Citation @article{gao2026dreamdojo, title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.

Size_categories:1M<n<10MFormat:parquetModality:tabularModality:videoLibrary:datasetsLibrary:dask
facebook/voxpopuli HF Unverified

Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.

Task_categories:automatic-Speech-RecognitionMultilinguality:multilingualLanguage:enLanguage:deLanguage:frLanguage:es
Y
yatharth97/T5-base-10K-summarization HF Unverified

SummarizationTransformersSafetensorsT5Text2text-GenerationGenerated_from_trainer MEDIUM
leosltl/Android-in-the-Wild HF Unverified

Android in the Wild (AITW) This is a mirror of Google's Android in the Wild (AITW) dataset, re-hosted on Hugging Face for easier community access. Original Source Paper: Android in the Wild: A Large-Scale Dataset for Android Device Control Original Repository: google-research/google-research/tree/master/android_in_the_wild Dataset Description Android in the Wild (AITW) is a large-scale dataset for Android device control. It contains human demonstrations of… See the full description on the dataset page: https://huggingface.co/datasets/leosltl/Android-in-the-Wild.

Task_categories:image-ClassificationTask_categories:visual-Question-AnsweringSize_categories:100M<n<1BAndroidMobileUi-Automation
ILSVRC/imagenet-1k HF Unverified

Dataset Card for ImageNet Dataset Summary ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are… See the full description on the dataset page: https://huggingface.co/datasets/ILSVRC/imagenet-1k.

Task_categories:image-ClassificationTask_ids:multi-Class-Image-ClassificationAnnotations_creators:crowdsourcedLanguage_creators:crowdsourcedMultilinguality:monolingualSource_datasets:original
CohereLabs/aya_collection HF Unverified

This dataset is uploaded in two places: here and additionally here as 'Aya Collection Language Split.' These datasets are identical in content but differ in structure of upload. This dataset is structured by folders split according to dataset name. The version here instead divides the Aya collection into folders split by language. We recommend you use the language split version if you are only interested in downloading data for a single or smaller set of languages, and this version if you… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection.

Task_categories:text-ClassificationTask_categories:summarizationTask_categories:translationLanguage:aceLanguage:afrLanguage:amh
M
michiyasunaga/BioLinkBERT-large HF Unverified

Text ClassificationTransformersPyTorchBERTFeature ExtractionExbert MEDIUM
natgillin/translations-raw HF Unverified

natgillin/translations-raw Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines. 31,663 parquet files (1566.8 GB) 49 language pairs under data/<src-tgt>/ Schema: 5 columns — see below Read-only for downstream pipelines. Do not delete or modify. Schema Each parquet has 5 columns: column type description source string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.

Language:multilingualSize_categories:1M<n<10MFormat:parquetModality:textLibrary:datasetsLibrary:dask
fpvlabs/stera-10m HF Unverified

Stera-10M Visualizer: https://platform.fpvlabs.ai/dataset/stera-10m/viz Dataset Summary Stera-10M is an open egocentric multimodal dataset for embodied AI, robotics, world models, and spatial intelligence, captured end-to-end on commodity iPhone Pro hardware through the open Stera platform. It contains 200 hours of synchronized first-person recordings across 500+ sessions from 20 contributors in 20+ unique environments, with 10 million RGB frames, LiDAR depth, ARKit… See the full description on the dataset page: https://huggingface.co/datasets/fpvlabs/stera-10m.

Task_categories:roboticsTask_categories:video-ClassificationTask_categories:image-To-TextTask_categories:depth-EstimationLanguage:enSize_categories:1M<n<10M
Showing 20 of 898 items (page 36 of 45)