Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Catalog

Library

Find the model or dataset for the question you have.

Ask in plain language and we will work out where the answer lives — or filter the index yourself. Almost everything here runs on our servers, from this page, with no setup.

Models
297
Datasets
245
Organisations
128

Labs, institutes and companies

Ready to run
24

Including 6 that run free and instantly

Looking for a particular value, not a particular dataset?Search schemas, rows and column statistics inside the data itself.

Things that work

Filtering

26 entries in Mathematics

Showing 26 of 26

aimo-validation-aimeAI-MOAIME I/II problems reformatted for AIMO challenge validation — 15-question integer-answer format, covering competition math at difficulty levels 5–9.Competition MathBrowse
aimo-validation-amcAI-MOAMC 10/12 competition problems reformatted for AIMO challenge validation, covering algebra, geometry, and number theory at difficulty levels 1–5.Competition MathBrowse
aimo-validation-math-level-4AI-MOLevel-4 MATH benchmark problems (pre-calculus difficulty) used for AIMO challenge validation and fine-grained model evaluation.Math ProblemsBrowse
aimo-validation-math-level-5AI-MOLevel-5 MATH benchmark problems (highest difficulty) used for AIMO challenge validation and measuring the ceiling of model mathematical reasoning.Math ProblemsBrowse
aops_rawAI-MORaw problem posts and discussion threads from the Art of Problem Solving forums, spanning AMC, AIME, and international olympiad competitions.Competition MathBrowse
BixBenchfuturehouseBenchmark with 205 reproducible research questions paired with data capsules for AI evaluation.Research BenchmarkBrowse
ChemBenchjablonkagroupManually curated benchmark of 3,000+ chemistry and materials science questions across spectroscopy, reactivity, synthesis, and property prediction for evaluating LLMs.Chemistry BenchmarkBrowse
CombiBenchAI-MOCombinatorics problems drawn from AMC, AIME, and olympiad competitions, formalised for benchmarking discrete-mathematics reasoning in language models.CombinatoricsBrowse
equational-theories-benchmarkSAIRfoundationFull benchmark suite of equational theory problems spanning algebraic structures, designed to evaluate formal reasoning capabilities of AI models.Mathematical ReasoningBrowse
equational-theories-selected-problemsSAIRfoundationCurated selection of equational theory problems for benchmarking LLM mathematical reasoning and automated theorem proving.Mathematical ReasoningBrowse
ether0-benchmarkfuturehouseChemistry reasoning benchmark covering SMILES-based tasks including reaction prediction, retrosynthesis, and molecular property estimation for evaluating chemistry LLMs.Chemistry BenchmarkBrowse
GeometryLeanBenchAI-MOGeometry theorem proving problems formalised in Lean 4, covering Euclidean, affine, and metric geometry for automated reasoning evaluation.Theorem ProvingBrowse
Kimina-Prover-PromptsetAI-MOPrompt-set for training and evaluating Kimina, a Lean 4 theorem prover that uses reinforcement learning over formal mathematical proofs.Theorem ProvingBrowse
lab-benchfuturehouseLanguage Agent Biology Benchmark - 8 categories of scientific research tasks including cloning, figures, and protocols.Research BenchmarkBrowse
MetaMathQAmeta-mathMathematical question-answering dataset for training and evaluating math reasoning.Math ProblemsBrowse
minif2f_testAI-MOTest set for miniF2F formal mathematics benchmark.Theorem ProvingBrowse
NuminaMath-1.5AI-MO860K+ competition math problems from 17 sources with verified solutions — the training backbone of the gold-medal solution at the 2024 AI Mathematical Olympiad.Math ProblemsBrowse
NuminaMath-CoTAI-MONuminaMath with Chain-of-Thought reasoning annotations.Math ProblemsBrowse
NuminaMath-LEANAI-MOMathematical problems formalized in LEAN proof assistant.Theorem ProvingBrowse
NuminaMath-TIRAI-MONuminaMath with Tool-Integrated Reasoning annotations.Math ProblemsBrowse
olympiadsAI-MOOlympiad-level mathematical problems collected from international and national competitions, formatted for training and evaluating mathematical reasoning models. ## Models (3)Math Reasoning CorpusBrowse
olympiads-refAI-MOExtended reference set of olympiad problems with verified step-by-step solutions, used for Chain-of-Thought and formal reasoning training.Competition MathBrowse
olympiads-ref-baseAI-MOCanonical reference set of international and national mathematical olympiad problems, used as the base for downstream NuminaMath training splits.Competition MathBrowse
principia-benchfacebookCurated benchmark of challenging STEM problems requiring multi-step reasoning, quantitative analysis, and domain knowledge across natural sciences.STEM BenchmarkBrowse
principia-collectionfacebookLarge-scale STEM reasoning dataset from Meta covering mathematics, physics, chemistry, and biology problems for training and evaluating scientific reasoning in language models.STEM ReasoningBrowse
spiqagoogleScientific Paper Image Question Answering benchmark requiring multimodal reasoning over figures, charts, and diagrams from research papers across scientific domains.Scientific BenchmarkBrowse