Model Hub

Browse PQC-verified AI models, datasets, and tools

anonnnnnsub/neuripsED_2701 HF Unverified

PURGE: Partition-Aware Unlearning for Removing Spurious-Correlation Generated Errors PURGE is a partitioning strategy applied to existing public datasets (MSCOCO 2017) that separates object-relevant evidence from spurious background cues in LVLMs. This repository hosts the resulting preprocessed retain/forget partitions for direct reuse. NeurIPS 2026 Evaluations & Datasets Track, Submission #2701. Hosted under an anonymous account for double-blind review; will be transferred… See the full description on the dataset page: https://huggingface.co/datasets/anonnnnnsub/neuripsED_2701.

Task_categories:image-ClassificationSize_categories:10K<n<100K
codraja2006/tomato-leaves-dataset HF Unverified

Tomato Leaves Dataset Overview This dataset contains images of tomato leaves categorized into different classes based on the type of disease or health condition. The dataset is divided into training, validation, and test sets, with a ratio of 8:1:1. The classes include various diseases as well as healthy leaves. The dataset includes both augmented and non-augmented images. Dataset Structure The dataset is organized into three main splits: train validation test… See the full description on the dataset page: https://huggingface.co/datasets/codraja2006/tomato-leaves-dataset.

Task_categories:feature-ExtractionTask_categories:image-ClassificationLanguage:enSize_categories:n<1KModality:imageTomato
ChengyouJia/agentic-critic-dataset HF Unverified

Agentic Critic Dataset High-quality AIGC images with rich metadata for aesthetic evaluation. Metadata Fields Each entry in metadata.jsonl contains: prompt: Positive prompt negative_prompt: Negative prompt model: Model name and hash sampler: Sampling method steps: Generation steps cfg_scale: CFG scale seed: Random seed stats: Engagement metrics image_path: Relative path to image Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/ChengyouJia/agentic-critic-dataset.

Task_categories:image-ClassificationTask_categories:text-To-ImageSize_categories:n<1KAigcCivitaiAesthetic
M
MCG-NJU/videomae-large HF Unverified

Video-ClassificationTransformersPyTorchSafetensorsVideomaePretraining MEDIUM
axel-riben/arcdataset-brutalism-extension HF Unverified

Architectural Styles Dataset (Curated and Extended) Dataset Summary A curated and extended version of dumitrux's Architectural Styles Dataset. The original dataset covered 25 architectural styles; 630 images were removed by automated filters (duplicates, low-resolution), leaving 9,483 images. A 26th class, Brutalism, was added from 284 manually curated Wikimedia Commons photographs, bringing the total to 9,767 images across 26 classes. Intended use: training and… See the full description on the dataset page: https://huggingface.co/datasets/axel-riben/arcdataset-brutalism-extension.

Task_categories:image-ClassificationSize_categories:n<1KFormat:imagefolderModality:imageLibrary:datasetsLibrary:mlcroissant
anilbhujel/Gilt_posture_dataset HF Unverified

Gilt Posture Recognition Dataset Each RGB image has a matching depth image (same filename, .png extension). YOLO-format label files correspond to each image. 🐷 Annotated Postures Five postures are labeled using YOLO bounding boxes: Class Name Class ID feeding 0 lateral_lying 1 sitting 2 standing 3 sternal_lying 4 📊 Class Distribution Below is a histogram showing the distribution of posture classes across the dataset:… See the full description on the dataset page: https://huggingface.co/datasets/anilbhujel/Gilt_posture_dataset.

Task_categories:object-DetectionTask_categories:image-ClassificationTask_ids:multi-Class-Image-ClassificationAnnotations_creators:expert-AnnotatedMultilinguality:monolingualLanguage:en
Ahnuf/Military_Aircraft_Detection_Classification_Image_Dataset HF Unverified

Military Aircraft Detection & Classification Dataset 88 Classes with Advanced Background Suppression Overview This dataset is a professionally curated resource for training high-performance object detection and image classification models such as YOLOv11.It contains 88 distinct military aircraft classes and is explicitly designed for real-world deployment, where false positives from civilian aircraft, birds, and small drones are common. To address this, the… See the full description on the dataset page: https://huggingface.co/datasets/Ahnuf/Military_Aircraft_Detection_Classification_Image_Dataset.

Task_categories:object-DetectionTask_categories:image-ClassificationMilitaryAircraftAerospaceYolo
Butterfree/IndoLepAtlas HF Unverified

IndoLepAtlas — Indian Lepidoptera & Host Plants Dataset A large-scale computer vision dataset of Indian butterflies, moths, and their larval host plants. Sourced from ifoundbutterflies.org with public CC-licensed photographs. Inspired by: iNaturalist | Domain: Indian Wildlife & Biodiversity Dataset Overview Butterflies Host Plants Total Species ~967 ~127 ~1,094 Images ~60,000 ~700 ~60,700 Source ifoundbutterflies.org ifoundbutterflies.org —… See the full description on the dataset page: https://huggingface.co/datasets/Butterfree/IndoLepAtlas.

Task_categories:image-ClassificationLanguage:enSize_categories:n<1KModality:imageModality:textWildlife
coastalcph/lex_glue HF Unverified

Dataset Card for "LexGLUE" Dataset Summary Inspired by the recent widespread use of the GLUE multi-task benchmark NLP dataset (Wang et al., 2018), the subsequent more difficult SuperGLUE (Wang et al., 2019), other previous multi-task NLP benchmarks (Conneau and Kiela, 2018; McCann et al., 2018), and similar initiatives in other domains (Peng et al., 2019), we introduce the Legal General Language Understanding Evaluation (LexGLUE) benchmark, a benchmark dataset to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/coastalcph/lex_glue.

Task_categories:question-AnsweringTask_categories:text-ClassificationTask_ids:multi-Class-ClassificationTask_ids:multi-Label-ClassificationTask_ids:multiple-Choice-QaTask_ids:topic-Classification
krithik274/NOAA-PIFSC-ESD-CORAL-Bleaching-Dataset HF Unverified

Dataset Card for NOAA-ESD-CORAL-Bleaching Classification Dataset v1 Overview For the development of machine learning models to classify coral health, specifically identifying healthy hard coral (CORAL) and bleached hard coral (CORAL_BL).This dataset contains underwater imagery collected by NOAA's Ecosystem Sciences Division (ESD) and other benthic surveys. Labels Label Name Functional Group CORAL Healthy Hard Coral Hard Coral CORAL_BL Bleached… See the full description on the dataset page: https://huggingface.co/datasets/krithik274/NOAA-PIFSC-ESD-CORAL-Bleaching-Dataset.

Task_categories:image-ClassificationLanguage:enModality:imageCoralBleachingCoral-Reef
ruggsea/infini-news-corpus HF Unverified

INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.

Task_categories:text-GenerationTask_categories:text-ClassificationTask_categories:text-RetrievalAnnotations_creators:machine-GeneratedMultilinguality:multilingualSource_datasets:original
P
Prior-Labs/TabPFN-v2-reg HF Unverified

Tabular-RegressionTabpfn MEDIUM
annoymous-1/CC-Bench HF Unverified

CC-Bench: A Cognitive Conflict Benchmark for MLLMs in Safety-Critical Visual Inspection CC-Bench is a joint medical-industrial benchmark for evaluating whether multimodal large language models (MLLMs) remain visually grounded when plausible textual context conflicts with image evidence. The benchmark reorganizes public anomaly datasets into a unified four-way multiple-choice QA format for high-risk visual inspection. This repository currently contains: 4,282 images in total 2,157… See the full description on the dataset page: https://huggingface.co/datasets/annoymous-1/CC-Bench.

Task_categories:visual-Question-AnsweringTask_categories:image-ClassificationSize_categories:1K<n<10KFormat:jsonModality:imageModality:text
zeio/mediach HF Unverified

Mediach This dataset contains images and videos from russian imageboard 2ch. This dataset contains media files attached to the first post in discussions on the image board. There is another dataset with thread texts, which can be used for solving media label prediction task. Some of the files include NSFW content Prerequisites You need to have lfs and xet installed in your system: sudo emerge --ask dev-vcs/git-lfs curl -sSfL https://hf.co/git-xet/install.sh | sh… See the full description on the dataset page: https://huggingface.co/datasets/zeio/mediach.

Task_categories:image-ClassificationTask_categories:video-ClassificationTask_categories:image-To-TextTask_ids:multi-Label-Image-ClassificationTask_ids:multi-Class-Image-ClassificationTask_ids:image-Captioning
zalando-datasets/fashion_mnist HF Unverified

Dataset Card for FashionMNIST Dataset Summary Fashion-MNIST is a dataset of Zalando's article images—consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes. We intend Fashion-MNIST to serve as a direct drop-in replacement for the original MNIST dataset for benchmarking machine learning algorithms. It shares the same image size and structure of training and testing… See the full description on the dataset page: https://huggingface.co/datasets/zalando-datasets/fashion_mnist.

Task_categories:image-ClassificationTask_ids:multi-Class-Image-ClassificationAnnotations_creators:expert-GeneratedLanguage_creators:foundMultilinguality:monolingualSource_datasets:original
schwein69/hagrid-subset HF Unverified

HaGRID Gesture Recognition Subset Dataset Description A curated subset of the HaGRID (Hand Gesture Recognition Image Dataset) containing 24 gesture classes for training gesture recognition models. Dataset Summary Total Images: 19,200 Gesture Classes: 24 Samples per Class: 800 Image Format: JPEG Average Image Size: ~302 KB Splits Split Images Percentage Train 14,592 76% Val 1,728 9% Test 2,880 15% Gesture Classes call… See the full description on the dataset page: https://huggingface.co/datasets/schwein69/hagrid-subset.

Task_categories:image-ClassificationTask_categories:object-DetectionSize_categories:10K<n<100KGesture-RecognitionComputer-VisionHand-Gestures
deepguess/tornet-temporal HF Unverified

TorNet-Temporal: Temporal Dual-Pol NEXRAD Radar for Tornado Detection A large-scale dataset of storm-centered NEXRAD WSR-88D radar sequences for tornado detection and prediction, featuring 24-channel dual-polarimetric data across variable-length temporal sequences. Dataset Summary 24,862 storm events from NEXRAD Level-II radar archives (2013-2022) 8-22 consecutive radar scans per event (~4-5 min cadence, ~45-90 min total; median 13 frames) 24 channels: 6 dual-pol radar… See the full description on the dataset page: https://huggingface.co/datasets/deepguess/tornet-temporal.

Task_categories:image-ClassificationTask_categories:video-ClassificationSize_categories:10K<n<100KWeatherRadarTornado
Forithmus/MR-RATE-atlas HF Unverified

MR-RATE: A Vision-Language Foundation Model and Dataset for Magnetic Resonance Imaging This is the MR-RATE-atlas repository, part of the MR-RATE dataset release. It contains atlas-registered MRI volumes in which all imaging sequences within each study have been spatially normalized to a standard atlas-space. For full dataset details, native-space MRI volumes, radiology reports, metadata, and data splits, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/Forithmus/MR-RATE-atlas.

Task_categories:image-To-TextTask_categories:text-To-ImageTask_categories:image-ClassificationTask_categories:question-AnsweringTask_categories:visual-Question-AnsweringTask_categories:zero-Shot-Classification
nguha/legalbench HF Unverified

Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.

Task_categories:text-ClassificationTask_categories:question-AnsweringLanguage:enSize_categories:10K<n<100KFormat:csvModality:tabular
imageomics/fish-vista HF Unverified

Dataset Card for Fish-Visual Trait Analysis (Fish-Vista) Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images. See Example Code to Use the Segmentation Dataset Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset. Instructions for downloading dataset and images… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fish-vista.

Task_categories:image-ClassificationTask_categories:image-SegmentationLanguage:enSize_categories:10K<n<100KFormat:csvModality:image
Showing 20 of 898 items (page 40 of 45)