Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Datasets

Datasets

Open any dataset and read its schema, its real rows and its column statistics — then search inside it. 245 curated, and far more beyond them.

Request a public dataset

Name one and we will usually open it straight away — paste a link, an owner/name, or just describe it.

245 datasets

ADSKAILab

ABC-1M

One million CAD-quality 3D shapes drawn from the ABC dataset — the foundation training corpus for the Make-A-Shape and WaLa generative models.

EngineeringMaterials Science

CAD Geometry Corpus

ADSKAILab

codeparrot_megatron

Megatron-formatted CodeParrot release used for large-scale code language-model pretraining experiments at Autodesk AI Lab. ## Models (25)

EngineeringScientific Reasoning

Code Pretraining

ADSKAILab

LLM-narrative-planning-taskset

Narrative planning task set for evaluating LLM planning and reasoning over multi-step design and engineering scenarios.

EngineeringScientific Reasoning

Planning Benchmark

ADSKAILab

Zero-To-CAD-100k

Curated 100K-example subset of Zero-To-CAD — useful for benchmarking and lightweight fine-tuning of CAD-from-image models.

EngineeringScientific Reasoning

CAD Vision-Language Corpus

ADSKAILab

Zero-To-CAD-1m

1M paired image-and-CAD-program examples for training vision-language models that synthesise parametric CAD from images.

EngineeringScientific Reasoning

CAD Vision-Language Corpus

Ahmad0067

MedSynth

Realistic synthetic medical dialogue–SOAP note pairs generated to support training and evaluation of clinical documentation models without exposing real patient data.

Medicine

Clinical NLP

AI-MO

aimo-validation-aime

AIME I/II problems reformatted for AIMO challenge validation — 15-question integer-answer format, covering competition math at difficulty levels 5–9.

MathematicsBenchmark

Competition Math

AI-MO

aimo-validation-amc

AMC 10/12 competition problems reformatted for AIMO challenge validation, covering algebra, geometry, and number theory at difficulty levels 1–5.

MathematicsBenchmark

Competition Math

AI-MO

aimo-validation-math-level-4

Level-4 MATH benchmark problems (pre-calculus difficulty) used for AIMO challenge validation and fine-grained model evaluation.

MathematicsBenchmark

Math Problems

AI-MO

aimo-validation-math-level-5

Level-5 MATH benchmark problems (highest difficulty) used for AIMO challenge validation and measuring the ceiling of model mathematical reasoning.

MathematicsBenchmark

Math Problems

AI-MO

aops_raw

Raw problem posts and discussion threads from the Art of Problem Solving forums, spanning AMC, AIME, and international olympiad competitions.

Mathematics

Competition Math

AI-MO

CombiBench

Combinatorics problems drawn from AMC, AIME, and olympiad competitions, formalised for benchmarking discrete-mathematics reasoning in language models.

MathematicsBenchmark

Combinatorics

AI-MO

GeometryLeanBench

Geometry theorem proving problems formalised in Lean 4, covering Euclidean, affine, and metric geometry for automated reasoning evaluation.

MathematicsBenchmark

Theorem Proving

AI-MO

Kimina-Prover-Promptset

Prompt-set for training and evaluating Kimina, a Lean 4 theorem prover that uses reinforcement learning over formal mathematical proofs.

MathematicsScientific Reasoning

Theorem Proving

AI-MO

minif2f_test

Test set for miniF2F formal mathematics benchmark.

MathematicsBenchmark

Theorem Proving

AI-MO

NuminaMath-1.5

860K+ competition math problems from 17 sources with verified solutions — the training backbone of the gold-medal solution at the 2024 AI Mathematical Olympiad.

Mathematics

Math Problems

AI-MO

NuminaMath-CoT

NuminaMath with Chain-of-Thought reasoning annotations.

MathematicsScientific Reasoning

Math Problems

AI-MO

NuminaMath-LEAN

Mathematical problems formalized in LEAN proof assistant.

Mathematics

Theorem Proving

AI-MO

NuminaMath-TIR

NuminaMath with Tool-Integrated Reasoning annotations.

MathematicsScientific Reasoning

Math Problems

AI-MO

olympiads

Olympiad-level mathematical problems collected from international and national competitions, formatted for training and evaluating mathematical reasoning models. ## Models (3)

MathematicsScientific Reasoning

Math Reasoning Corpus

AI-MO

olympiads-ref

Extended reference set of olympiad problems with verified step-by-step solutions, used for Chain-of-Thought and formal reasoning training.

MathematicsScientific Reasoning

Competition Math

AI-MO

olympiads-ref-base

Canonical reference set of international and national mathematical olympiad problems, used as the base for downstream NuminaMath training splits.

Mathematics

Competition Math

Aignostics

OpenTME

Pre-analyzed H&E whole-slide images from TCGA across breast, bladder, colorectal, liver, and lung cancers — cell-level annotations and tumour-microenvironment spatial features generated by Atlas H&E-TME.

MedicineBiology

Digital Pathology

allenai

olmoearth-paper-embeddings

Precomputed embeddings released alongside the OlmoEarth paper — drop-in feature set for evaluating Earth-observation foundation-model performance on downstream tasks without re-running the backbone.

Earth ScienceClimate

Earth-Observation Embeddings