Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Catalog

Library

Find the model or dataset for the question you have.

Ask in plain language and we will work out where the answer lives — or filter the index yourself. Almost everything here runs on our servers, from this page, with no setup.

Models
297
Datasets
245
Organisations
128

Labs, institutes and companies

Ready to run
24

Including 6 that run free and instantly

Looking for a particular value, not a particular dataset?Search schemas, rows and column statistics inside the data itself.

Things that work

Filtering

67 entries in Medicine

Showing 48 of 67

AlphaFoldDBLiteFoldAlphaFoldDB prediction index — open database of predicted protein 3D structures with confidence scores, providing structural coverage for known protein sequences at scale.Predicted StructuresBrowse
B3DBmaomlabBlood-Brain Barrier Database (B3DB) — curated permeability measurements for compounds, supporting CNS drug-discovery ML benchmarks.BBB PermeabilityBrowse
BioMysteryBench-fullAnthropicFull BioMysteryBench evaluation set — challenging biology problems used to probe expert-level scientific reasoning in frontier models.Biology BenchmarkBrowse
BioMysteryBench-previewAnthropicPreview slice of BioMysteryBench — challenging, expert-curated biology problems for evaluating AI scientific reasoning capability.Biology BenchmarkBrowse
camelyon16-featuresowkinPre-extracted features from the CAMELYON16 breast cancer lymph node metastasis detection challenge, enabling efficient benchmarking of MIL methods.Computational PathologyBrowse
clinvar-vepHuggingFaceBioClinVar variants annotated with VEP (Variant Effect Predictor) — formatted for benchmarking genomic-LLM pathogenicity classification.Clinical VariantsBrowse
CryptoCENmaomlabCryptoCEN — Cryptococcus coexpression network dataset for fungal pathogen biology and drug-target prioritisation.Coexpression NetworkBrowse
CT_DeepLesion-MedSAM2wanglabCT volumes from the DeepLesion benchmark with mask annotations restructured for training and evaluating MedSAM2, the universal medical image segmentation foundation model.Medical ImagingBrowse
CycPepMPDBLiteFoldCycPeptMPDB — public dataset of experimentally measured membrane permeability for cyclic peptides, the standard resource for peptide-permeability modelling.Peptide PropertiesBrowse
DisProtLiteFoldDisProt — manually curated database of intrinsically disordered proteins and regions with experimental evidence and functional annotations.Protein AnnotationBrowse
drug-target-activityeve-bioDrug-target interaction measurements for 1,397 FDA-approved small molecule drugs.Drug DiscoveryBrowse
DynaCLR-databiohubTraining and evaluation data for DynaCLR — dynamic contrastive learning of cell embeddings from live-cell microscopy time series, including infection-status labels.MicroscopyBrowse
EmeraldBaytahoebioEmerald Bay — single-cell perturbation dataset of 1.8M+ transcriptomic profiles spanning 52 cell lines × 91 drug treatments (plus combinations), generated on Tahoe’s MOSAIC high-throughput platform with paired transcriptional and drug-phenotype readouts over a five-day culture.Single-Cell PerturbationBrowse
ether0-benchmarkfuturehouseChemistry reasoning benchmark covering SMILES-based tasks including reaction prediction, retrosynthesis, and molecular property estimation for evaluating chemistry LLMs.Chemistry BenchmarkBrowse
finemed-frdoctolib-lab21.1M documents and 19.2B words of French medical text curated from FineWeb-2, FinePDFs, and FineWiki — annotated for medical subdomain, educational quality, and medical-term density; the pretraining corpus for DoctoBERT and DoctoModernBERT.Medical Pretraining CorpusBrowse
finemed-rephrased-frdoctolib-labLLM-rephrased variant of FineMed-fr (13.6M documents) — synthetically rewritten French medical text used alongside the original corpus to pretrain DoctoBERT and DoctoModernBERT. ## Models (102)Medical Pretraining CorpusBrowse
FireProtDBLiteFoldFireProtDB — curated thermostability mutation data for proteins, supporting protein-engineering and stability-prediction modelling.Protein StabilityBrowse
foundry_moses_v1-1foundry-mlFoundry mirror of MOSES — molecular sets benchmark for evaluating generative chemistry models on drug-like molecule generation.Molecular GenerationBrowse
healthbenchopenaiRealistic multi-turn health conversations graded against physician-written rubrics across multiple axes (accuracy, completeness, communication) — an open evaluation benchmark for AI assistants in medicine.Medical BenchmarkBrowse
healthbench-professionalopenaiProfessional-graded subset of HealthBench: physician evaluators score model responses to clinically realistic conversations, targeting expert-level health assessment.Medical BenchmarkBrowse
her2-challenge-2026owkinHER2 scoring challenge dataset with H&E-stained whole-slide images for evaluating AI-based HER2 status prediction in breast cancer.Computational PathologyBrowse
hestMahmoodLabHEST-1k — 1,276 spatial-transcriptomic profiles each linked and aligned to a Whole Slide Image (pixel size <1.15 µm/px), the largest paired histology + spatial-transcriptomics resource on the Hub.Spatial TranscriptomicsBrowse
hest-benchMahmoodLabHEST-Bench — companion benchmark to HEST-1k for evaluating spatial-transcriptomics and pathology foundation models on aligned histology / expression tasks.Pathology BenchmarkBrowse
HumanProteinAtlasLiteFoldHuman Protein Atlas — gene and protein expression across tissues, cells, organs, pathology, and subcellular locations for human proteins.Protein AnnotationBrowse
IEDBLiteFoldIEDB assay export — experimentally characterised B-cell, T-cell, and MHC-binding epitopes across infectious, allergic, autoimmune, and transplant settings.ImmunologyBrowse
MedDialogOpenMedDoctor-patient medical dialogue dataset for training and evaluating clinical conversation models — covers triage, symptom checking, and diagnostic reasoning.Medical DialogueBrowse
medical-o1-reasoning-SFTFreedomIntelligenceMedical chain-of-thought reasoning dataset (o1-style) for supervised fine-tuning of medical LLMs — one of the most-liked medical training corpora on Hugging Face (1000+ likes).Medical Reasoning CorpusBrowse
medical-o1-verifiable-problemFreedomIntelligenceVerifiable medical reasoning problems with checker functions — supports RL/reward-model training for medical-LLM alignment beyond static SFT.Medical RL Reward DataBrowse
Medical-Reasoning-SFT-MegaOpenMedLarge supervised fine-tuning corpus for clinical reasoning — multi-step medical question-answer chains with rationales for training instruction-following medical LLMs.SFT Reasoning CorpusBrowse
MedSynthAhmad0067Realistic synthetic medical dialogue–SOAP note pairs generated to support training and evaluation of clinical documentation models without exposing real patient data.Clinical NLPBrowse
MegaScale-Tsuboyama2023LiteFoldMegaScale (Tsuboyama et al., 2023) — large-scale experimental measurements of protein stability across hundreds of thousands of mutations for stability prediction modelling.Protein StabilityBrowse
miriad-4.4Mmiriad4.4M-example medical reasoning subset of MIRIAD — earlier release used for benchmarking medical instruction-tuning workflows.Medical Reasoning CorpusBrowse
miriad-5.8Mmiriad5.8M-example medical instruction-tuning and reasoning corpus curated from clinical literature for training healthcare LLMs at scale.Medical Reasoning CorpusBrowse
nct-crc-heowkinColorectal cancer tissue classification dataset with H&E-stained patches across 9 tissue classes, widely used for benchmarking pathology models.Computational PathologyBrowse
OASopigObserved Antibody Space: a curated database of over one billion antibody sequences from immune repertoire sequencing studies, the standard resource for antibody ML.Antibody SequencesBrowse
Octant_CYP_inhibition_reactivity_blog_releaseopenadmetOctant CYP inhibition and chemical reactivity dataset measuring cytochrome P450 activity across a diverse compound library for ADMET modelling.Drug DiscoveryBrowse
openadmet-expansionrx-challenge-dataopenadmetFull ExpansionRx challenge dataset of RNA-targeted small-molecule compounds with measured ADMET properties for open pharmacokinetics benchmarking.Drug DiscoveryBrowse
openadmet-expansionrx-challenge-train-dataopenadmetTraining data for the OpenADMET ExpansionRx ADMET prediction challenge.Drug DiscoveryBrowse
opengenome2arcinstituteCurated collection of prokaryotic and eukaryotic genomic sequences for training and benchmarking large-scale biological foundation models.GenomicsBrowse
OpenTMEAignosticsPre-analyzed H&E whole-slide images from TCGA across breast, bladder, colorectal, liver, and lung cancers — cell-level annotations and tumour-microenvironment spatial features generated by Atlas H&E-TME.Digital PathologyBrowse
parse-de-rhaistertahoebioDifferential-expression summary statistics for the Parse PBMC cytokine screen (12 donors × 18 immune cell types × 90 cytokines), derived from the Arc Institute ST-HVG-Parse dataset — per-gene fold changes, significance, and pseudobulk deltas used to train Rhaister.Cytokine Differential ExpressionBrowse
Patho-BenchMahmoodLabPatho-Bench — benchmark designed to evaluate patch and slide encoder foundation models for whole-slide images across cancer subtyping, biomarker prediction, and survival tasks.Pathology BenchmarkBrowse
PATHOS-PLM-EMBEDDINGSDSIMBPATHOS protein language-model embeddings — precomputed feature representations supporting pathogenicity prediction and downstream macromolecular ML on protein sequences.Protein EmbeddingsBrowse
PDBLiteFoldHub-packaged index over the Protein Data Bank mmCIF entries — the global archive of experimentally-determined 3D structures of biological macromolecules.Protein StructureBrowse
Perturb-SapiensarcinstituteLarge-scale human single-cell perturbation dataset used in the STACK foundation-model lineage — paired baseline and perturbed expression profiles for genetic perturbation screens.Single-Cell PerturbationBrowse
plism-dataset-tilesowkinLarge-scale histopathology tile dataset for benchmarking robustness of pathology foundation models across staining and scanner variability.Computational PathologyBrowse
replogle-nadig-de-rhaistertahoebioDifferential-expression summary statistics for the Replogle-Nadig genome-scale CRISPR screen (4 cell lines × ~2,000 gene knockdowns: HepG2, Jurkat, K562, RPE1) — per-gene fold changes, significance, and pseudobulk deltas used to train Rhaister.CRISPR Differential ExpressionBrowse
Replogle-Nadig-PreprintarcinstituteReplogle-Nadig single-cell perturbation dataset (preprint release) — Perturb-seq screens used in the STATE single-cell embedding work for perturbation-response modelling.Single-Cell PerturbationBrowse

19 more behind this view