Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Catalog
  2. Datasets
  3. HuggingFaceBio/carbon-pretraining-corpus

HuggingFaceBio

carbon-pretraining-corpus

Carbon pretraining corpus — the curated DNA sequence dataset used to train the Carbon-500M / 3B / 8B genomic foundation models.

Original source
Rows
179,874,547
On disk
497 GB
Downloads
3.8k

Last 30 days

Updated
Jun 18

Explore

Read the real rows without downloading anything

Reading rows…

LiveRead from HuggingFaceBio/carbon-pretraining-corpus at the moment you asked. Nothing is cached or stored — every row above came from that request.

Splits

5
  • eukaryote_generator/train
  • eukaryote_generator_10B_subset/train
  • mrna_evo2/train
  • mrna_splice_evo2/train
  • prokaryote_evo2/train

Query support

  • Row preview
  • Paginated browse
  • Full-text search
  • SQL filter
  • Column statistics

Provenance

Licence
other
Likes
30

Fields

BiologyGenomics
text-generation