Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Catalog

Library

Find the model or dataset for the question you have.

Ask in plain language and we will work out where the answer lives — or filter the index yourself. Almost everything here runs on our servers, from this page, with no setup.

Models
297
Datasets
245
Organisations
128

Labs, institutes and companies

Ready to run
24

Including 6 that run free and instantly

Looking for a particular value, not a particular dataset?Search schemas, rows and column statistics inside the data itself.

Things that work

Filtering

48 entries in Scientific Reasoning

Showing 48 of 48

aimo-validation-aimeAI-MOAIME I/II problems reformatted for AIMO challenge validation — 15-question integer-answer format, covering competition math at difficulty levels 5–9.Competition MathBrowse
aimo-validation-amcAI-MOAMC 10/12 competition problems reformatted for AIMO challenge validation, covering algebra, geometry, and number theory at difficulty levels 1–5.Competition MathBrowse
aimo-validation-math-level-4AI-MOLevel-4 MATH benchmark problems (pre-calculus difficulty) used for AIMO challenge validation and fine-grained model evaluation.Math ProblemsBrowse
aimo-validation-math-level-5AI-MOLevel-5 MATH benchmark problems (highest difficulty) used for AIMO challenge validation and measuring the ceiling of model mathematical reasoning.Math ProblemsBrowse
AirfRANS_originalPLAID-datasetsOriginal AirfRANS airfoil RANS simulation dataset — graph-structured CFD over NACA airfoils for benchmarking physics-informed and graph neural networks.Aerodynamics CFDBrowse
BioMysteryBench-fullAnthropicFull BioMysteryBench evaluation set — challenging biology problems used to probe expert-level scientific reasoning in frontier models.Biology BenchmarkBrowse
BioMysteryBench-previewAnthropicPreview slice of BioMysteryBench — challenging, expert-curated biology problems for evaluating AI scientific reasoning capability.Biology BenchmarkBrowse
bioreason-pro-sft-reasoning-datawanglabReasoning trace dataset used to supervised-fine-tune BioReason-Pro — multimodal biological problems with rationales over genomic variants and pathway data.Biological Reasoning CorpusBrowse
BixBenchfuturehouseBenchmark with 205 reproducible research questions paired with data capsules for AI evaluation.Research BenchmarkBrowse
ChemBenchjablonkagroupManually curated benchmark of 3,000+ chemistry and materials science questions across spectroscopy, reactivity, synthesis, and property prediction for evaluating LLMs.Chemistry BenchmarkBrowse
chempile-captionjablonkagroupImage-to-text dataset of chemistry figures (molecular structures, reaction schemes, plots) with expert captions for training multimodal chemistry models.Chemistry CaptioningBrowse
chempile-codejablonkagroupCurated chemistry-relevant code (RDKit, ASE, simulation tooling) drawn from The Stack — supports training models that can read and write computational chemistry workflows.Chemistry Code CorpusBrowse
chempile-educationjablonkagroupEducational chemistry corpus — multiple-choice and open-ended items spanning introductory through graduate chemistry for assessing model educational capability.Chemistry Education CorpusBrowse
chempile-instructionjablonkagroupInstruction-tuning corpus for chemistry — curated Q&A and dialogue traces drawn from chemical literature and educational sources for training chemistry-specialist LLMs.Chemistry Instruction CorpusBrowse
chempile-liftjablonkagroupChemPile-LIFT — large-scale language-modelling dataset combining curated chemistry literature and structured chemical knowledge for foundation-model pretraining.Chemistry PretrainingBrowse
chempile-reasoningjablonkagroupMulti-step chemistry reasoning corpus — open-domain QA, NLI, and multiple-choice items with chains of reasoning for training and evaluating chemical reasoning models.Chemistry Reasoning CorpusBrowse
codeparrot_megatronADSKAILabMegatron-formatted CodeParrot release used for large-scale code language-model pretraining experiments at Autodesk AI Lab. ## Models (25)Code PretrainingBrowse
CombiBenchAI-MOCombinatorics problems drawn from AMC, AIME, and olympiad competitions, formalised for benchmarking discrete-mathematics reasoning in language models.CombinatoricsBrowse
dataset_metallicglass_rc_llmfoundry-mlLLM-extracted critical cooling rate data for metallic glasses — text-mined complement to the structured Rc dataset.Metallic GlassBrowse
equational-theories-benchmarkSAIRfoundationFull benchmark suite of equational theory problems spanning algebraic structures, designed to evaluate formal reasoning capabilities of AI models.Mathematical ReasoningBrowse
equational-theories-selected-problemsSAIRfoundationCurated selection of equational theory problems for benchmarking LLM mathematical reasoning and automated theorem proving.Mathematical ReasoningBrowse
ether0-benchmarkfuturehouseChemistry reasoning benchmark covering SMILES-based tasks including reaction prediction, retrosynthesis, and molecular property estimation for evaluating chemistry LLMs.Chemistry BenchmarkBrowse
frontierscienceopenaiFrontier science evaluation benchmark probing model capabilities on expert-level reasoning across natural sciences — designed to surface what AI systems can and cannot do at the research frontier.Scientific Reasoning BenchmarkBrowse
genomic-niahHuggingFaceBioGenomic Needle-in-a-Haystack — long-context evaluation for genomic LLMs, measuring retrieval and reasoning over long DNA sequences.Long-Context BenchmarkBrowse
GeometryLeanBenchAI-MOGeometry theorem proving problems formalised in Lean 4, covering Euclidean, affine, and metric geometry for automated reasoning evaluation.Theorem ProvingBrowse
healthbenchopenaiRealistic multi-turn health conversations graded against physician-written rubrics across multiple axes (accuracy, completeness, communication) — an open evaluation benchmark for AI assistants in medicine.Medical BenchmarkBrowse
healthbench-professionalopenaiProfessional-graded subset of HealthBench: physician evaluators score model responses to clinically realistic conversations, targeting expert-level health assessment.Medical BenchmarkBrowse
keggwanglabKEGG pathway entries paired with variant annotations for training and evaluating multimodal biological reasoning models (used by the BioReason work).Biological ReasoningBrowse
Kimina-Prover-PromptsetAI-MOPrompt-set for training and evaluating Kimina, a Lean 4 theorem prover that uses reinforcement learning over formal mathematical proofs.Theorem ProvingBrowse
lab-benchfuturehouseLanguage Agent Biology Benchmark - 8 categories of scientific research tasks including cloning, figures, and protocols.Research BenchmarkBrowse
LLM-narrative-planning-tasksetADSKAILabNarrative planning task set for evaluating LLM planning and reasoning over multi-step design and engineering scenarios.Planning BenchmarkBrowse
MedDialogOpenMedDoctor-patient medical dialogue dataset for training and evaluating clinical conversation models — covers triage, symptom checking, and diagnostic reasoning.Medical DialogueBrowse
medical-o1-reasoning-SFTFreedomIntelligenceMedical chain-of-thought reasoning dataset (o1-style) for supervised fine-tuning of medical LLMs — one of the most-liked medical training corpora on Hugging Face (1000+ likes).Medical Reasoning CorpusBrowse
medical-o1-verifiable-problemFreedomIntelligenceVerifiable medical reasoning problems with checker functions — supports RL/reward-model training for medical-LLM alignment beyond static SFT.Medical RL Reward DataBrowse
Medical-Reasoning-SFT-MegaOpenMedLarge supervised fine-tuning corpus for clinical reasoning — multi-step medical question-answer chains with rationales for training instruction-following medical LLMs.SFT Reasoning CorpusBrowse
minif2f_testAI-MOTest set for miniF2F formal mathematics benchmark.Theorem ProvingBrowse
miriad-4.4Mmiriad4.4M-example medical reasoning subset of MIRIAD — earlier release used for benchmarking medical instruction-tuning workflows.Medical Reasoning CorpusBrowse
miriad-5.8Mmiriad5.8M-example medical instruction-tuning and reasoning corpus curated from clinical literature for training healthcare LLMs at scale.Medical Reasoning CorpusBrowse
NuminaMath-CoTAI-MONuminaMath with Chain-of-Thought reasoning annotations.Math ProblemsBrowse
NuminaMath-TIRAI-MONuminaMath with Tool-Integrated Reasoning annotations.Math ProblemsBrowse
olympiadsAI-MOOlympiad-level mathematical problems collected from international and national competitions, formatted for training and evaluating mathematical reasoning models. ## Models (3)Math Reasoning CorpusBrowse
olympiads-refAI-MOExtended reference set of olympiad problems with verified step-by-step solutions, used for Chain-of-Thought and formal reasoning training.Competition MathBrowse
peS2oallenaiApproximately 40M cleaned, filtered, and formatted open-access academic papers derived from S2ORC — a large multi-domain pretraining corpus for science-aware language models, spanning biology, chemistry, engineering, computer science, and physics.Pretraining CorpusBrowse
principia-benchfacebookCurated benchmark of challenging STEM problems requiring multi-step reasoning, quantitative analysis, and domain knowledge across natural sciences.STEM BenchmarkBrowse
principia-collectionfacebookLarge-scale STEM reasoning dataset from Meta covering mathematics, physics, chemistry, and biology problems for training and evaluating scientific reasoning in language models.STEM ReasoningBrowse
spiqagoogleScientific Paper Image Question Answering benchmark requiring multimodal reasoning over figures, charts, and diagrams from research papers across scientific domains.Scientific BenchmarkBrowse
Zero-To-CAD-100kADSKAILabCurated 100K-example subset of Zero-To-CAD — useful for benchmarking and lightweight fine-tuning of CAD-from-image models.CAD Vision-Language CorpusBrowse
Zero-To-CAD-1mADSKAILab1M paired image-and-CAD-program examples for training vision-language models that synthesise parametric CAD from images.CAD Vision-Language CorpusBrowse