Model Hub
Browse PQC-verified AI models, datasets, and tools
Google's Gemma 2 instruction-tuned model. Lightweight yet performant with knowledge distillation from larger models.
NVIDIA's optimized Llama 3.1 70B. Custom alignment for helpfulness with strong benchmark performance.
Human-generated, human-annotated conversation trees. 91K messages across 35+ languages. RLHF training data.
Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far larger… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.