Model Hub

Browse PQC-verified AI models, datasets, and tools

annoymous-1/CC-Bench HF Unverified

CC-Bench: A Cognitive Conflict Benchmark for MLLMs in Safety-Critical Visual Inspection CC-Bench is a joint medical-industrial benchmark for evaluating whether multimodal large language models (MLLMs) remain visually grounded when plausible textual context conflicts with image evidence. The benchmark reorganizes public anomaly datasets into a unified four-way multiple-choice QA format for high-risk visual inspection. This repository currently contains: 4,282 images in total 2,157… See the full description on the dataset page: https://huggingface.co/datasets/annoymous-1/CC-Bench.

Task_categories:visual-Question-AnsweringTask_categories:image-ClassificationSize_categories:1K<n<10KFormat:jsonModality:imageModality:text
codeparrot/github-code-clean HF Unverified

The GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in almost 1TB of text data.

Size_categories:10M<n<100MModality:textLibrary:datasetsLibrary:mlcroissant
J
John6666/amanatsu-illustrious-v11-sdxl HF PQC Verified

Text-to-ImageDiffusersSafetensorsStable-DiffusionStable-Diffusion-XlAnime HIGH
mlfoundations/MINT-1T-HTML HF Unverified

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.

Task_categories:image-To-TextTask_categories:text-GenerationLanguage:enSize_categories:100M<n<1BFormat:parquetModality:text
J
John6666/obsession-illustriousxl-v10-sdxl HF PQC Verified

Text-to-ImageDiffusersSafetensorsStable-DiffusionStable-Diffusion-XlAnime HIGH
J
jonathandinu/face-parsing HF Unverified

Image-SegmentationTransformersPyTorchONNXSafetensorsSegformer HIGH
google-research-datasets/mbpp HF Unverified

Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.

Annotations_creators:crowdsourcedAnnotations_creators:expert-GeneratedLanguage_creators:crowdsourcedLanguage_creators:expert-GeneratedMultilinguality:monolingualSource_datasets:original
T
TahaDouaji/detr-doc-table-detection HF PQC Verified

Object-DetectionTransformersPyTorchONNXSafetensorsDetr MEDIUM
C
cagliostrolab/animagine-xl-3.1 HF PQC Verified

Text-to-ImageDiffusersSafetensorsStable-DiffusionStable-Diffusion-XlBase_model:cagliostrolab/animagine-Xl-3.0 CRITICAL
X
Xenova/segformer-b0-finetuned-ade-512-512 HF Unverified

Image-SegmentationTransformers.jsONNXSegformerBase_model:nvidia/segformer-B0-Finetuned-Ade-512-512Base_model:quantized:nvidia/segformer-B0-Finetuned-Ade-512-512 MEDIUM
echodict/KakologArchives_duplicate HF Unverified

ニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA などの家電への対応が軒並み終了する中、当時の生の声が詰まった約11年分の過去ログも同時に失われることとなってしまいました。 そこで 5ch の DTV 板の住民が中心となり、旧ニコニコ実況が終了するまでに11年分の全チャンネルの過去ログをアーカイブする計画が立ち上がりました。紆余曲折あり Nekopanda 氏が約11年分のラジオや BS も含めた全チャンネルの過去ログを完璧に取得してくださったおかげで、11年分の過去ログが電子の海に消えていく事態は回避できました。しかし、旧 API が廃止されてしまったため過去ログを API… See the full description on the dataset page: https://huggingface.co/datasets/echodict/KakologArchives_duplicate.

Task_categories:text-ClassificationLanguage:ja
labofsahil/pypi-packages-metadata-dataset HF Unverified

Size_categories:10M<n<100MModality:text
isaacus/open-australian-legal-corpus HF Unverified

Open Australian Legal Corpus ‍⚖️ The Open Australian Legal Corpus by Isaacus, a foundational legal AI research company, is the first and only multijurisdictional open corpus of Australian legislative and judicial documents. Comprised of 229,122 texts totalling over 60 million lines and 1.4 billion tokens, the Corpus includes every in force statute and regulation in the Commonwealth, New South Wales, Queensland, Western Australia, South Australia, Tasmania and Norfolk Island, in… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/open-australian-legal-corpus.

Task_categories:text-GenerationTask_categories:fill-MaskTask_categories:text-RetrievalTask_ids:language-ModelingTask_ids:masked-Language-ModelingTask_ids:document-Retrieval
G
google/pegasus-xsum HF Unverified

SummarizationTransformersPyTorchTfJAXPegasus HIGH
mvp-lab/LLaVA-OneVision-1.5-Instruct-Data HF Unverified

LLaVA-OneVision-1.5 Instruction Data Paper | Code 📌 Introduction This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the development of LLaVA-OneVision-1.5. LLaVA-OneVision-1.5 is a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. This meticulously curated 22M instruction dataset (LLaVA-OneVision-1.5-Instruct) is part of a comprehensive and… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Instruct-Data.

Task_categories:image-Text-To-TextLanguage:enSize_categories:10M<n<100MModality:imageModality:textMultimodal
C
cross-encoder/nli-deberta-v3-xsmall HF Unverified

Zero-Shot ClassificationSentence-TransformersPyTorchONNXSafetensorsDeberta-V2 HIGH
F
facebook/mask2former-swin-large-cityscapes-semantic HF Unverified

Image-SegmentationTransformersPyTorchSafetensorsMask2formerVision HIGH
nebius/SWE-rebench-V2-PRs HF Unverified

SWE-rebench-V2-PRs Dataset Summary SWE-rebench-V2-PRs is a large-scale dataset of real-world GitHub pull requests collected across multiple programming languages, intended for training and evaluating code-generation and software-engineering agents. The dataset contains 126,300 samples covering Go, Python, JavaScript, TypeScript, Rust, Java, C, C++, Julia, Elixir, Kotlin, PHP, Scala, Clojure, Dart, OCaml, and other languages. For log parser functions, base Dockerfiles, and… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-V2-PRs.

Task_categories:text-GenerationLanguage:enSize_categories:100K<n<1MFormat:parquetModality:textLibrary:datasets
O
openvla/openvla-7b HF PQC Verified

RoboticsTransformersSafetensorsOpenvlaFeature ExtractionVla HIGH
W
Wan-AI/Wan2.1-T2V-1.3B-Diffusers HF Unverified

Text-To-VideoDiffusersSafetensorsVideoVideo-GenerationDiffusers:WanPipeline HIGH
Showing 20 of 738 items (page 25 of 37)