Model Hub
Browse PQC-verified AI models, datasets, and tools
đ LLaVA-One-Vision-1.5-Mid-Training-85M Dataset is being uploaded đ Upload Status All Completed: ImageNet-21kăLAIONCNăDataComp-1BăZero250MăCOYO700MăSA-1BăMINTăObelics đ Cite If you find LLaVA-One-Vision-1.5-Mid-Training-85M useful in your research, please consider to cite the following related papers: @misc{an2025llavaonevision15fullyopenframework, title={LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training}⌠See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M.
Further cleaning done. Please look through the dataset and ensure that I didn't miss anything. Update: Confirmed working method for training the model: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c Two choices: Removes instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/blob/main/ShareGPT_V3_unfiltered_cleaned_split_no_imsorry.json Has instances of "I'm sorry, but":⌠See the full description on the dataset page: https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered.
LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used in mid-training. Dataset Composition Subset Format Description mid_training_video/60s_rest/ WebDataset (.tar) 10,809 shards of ~60s video clips mid_training_video/caption_v0/split_30s.jsonl JSONL Captions for 30-second video clips mid_training_video/caption_v0/split_60s.jsonl JSONL Captions for⌠See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.
GPIC: A Giant Permissive Image Corpus for Visual Generation Keshigeyan Chandrasegaran*1, Kyle Sargent*1, Suchir Agarwal1, Michael Jang1, Michael Poli1,2, Juan Carlos Niebles1,4, Justin Johnson3, Jiajun Wu1, Li Fei-Fei1 1 Stanford University 2 Radical Numerics 3 University of Michigan 4 Salesforce⌠See the full description on the dataset page: https://huggingface.co/datasets/stanford-vision-lab/gpic.
SAGE-10k SAGE-10k is a large-scale interactive indoor scene dataset featuring realistic layouts, generated by the agentic-driven pipeline introduced in "SAGE: Scalable Agentic 3D Scene Generation for Embodied AI". The dataset contains 10,000 diverse scenes spanning 50 room types and styles, along with 565K uniquely generated 3D objects. đ Key Features SAGE-10k integrates a wide variety of scenes, and particularly, preserves small items for⌠See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SAGE-10k.
Dataset Card for Fish-Visual Trait Analysis (Fish-Vista) Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images. See Example Code to Use the Segmentation Dataset Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset. Instructions for downloading dataset and images⌠See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fish-vista.