Model Hub
Browse PQC-verified AI models, datasets, and tools
Largest open code dataset. 67.5TB of permissively licensed source code across 600+ programming languages from Software Heritage.
Code LLM trained on The Stack v2 with 600+ programming languages. 4x the training data of StarCoder1.
Dataset Card for The Stack Metadata Changelog Release Description v1.1 This is the first release of the metadata. It is for The Stack v1.1 v1.2 Metadata dataset matching The Stack v1.2 Dataset Summary This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories. Supported Tasks and Leaderboards The main task is to recreate… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.