Breaking
Through
The
Data
Wall
Basecamp Research’s proprietary biological data powers a new class of frontier biological foundation models
Breaking
Through The
Data Wall
Basecamp Research’s proprietary biological data powers a new class of frontier biological foundation models
Breaking Through
The Data Wall
Basecamp Research’s proprietary biological data powers a new class of frontier biological foundation models
The data wall holding back AI in biology
Today’s biological datasets are narrow, repetitive, and heavily biased toward a small number of well studied organisms. Over 68% of data comes from just 5 species, 70% originates from only 10 countries, and the most widely used database in AI model training is growing by less than 10% per year.
Much of biology's diversity remains entirely absent from the data used to train AI models. This is the data wall, and it has created a fundamental bottleneck on what biological models can learn, predict, and design, resulting in plateauing levels of AI model output.
Basecamp Research’s data is intricately connected: a vast web of biological networks capturing relationships across environments, ecosystems, and evolutionary history.
By creating a global network of biodiversity partners, we have created the first global supply chain of primary evolutionary data spanning 31 countries across six continents. This allowed us to create BaseData™; the largest and most diverse evolutionary genetic dataset created to date, spanning over 10 billion novel genes from over 1 million new species, breaking the data wall holding back AI in biology.

We used this immense corpus of novel biology to train EDEN, our biological foundation model family, and in doing so revealed new scaling laws: as biological datasets grow larger and richer, AI capabilities jump, opening the door to systems that can design new medicines across multiple diseases. We’re now expanding our dataset with The Trillion Gene Atlas, expanding our current dataset by another 100x, unlocking a new generation of biological intelligence.

BaseData™ isn’t scraped from public data, it’s built from the ground up through field exploration in some of the world’s most extreme and biodiverse environments. From deep-sea vents to polar ice shelves, our global network uncovers entirely new species and functions, expanding biology’s known frontier with every expedition.
The Trillion Gene Atlas builds on these collaborations, adding hundreds of new partners to our network worldwide.

The Trillion Gene Atlas
The Trillion Gene Atlas is one of the most ambitious biological data initiatives ever undertaken, a voyage of evolutionary discovery led by Basecamp Research and its expanding global network of biodiversity partners. The goal is to expand known evolutionary genetic diversity by 1000-fold across more than 100 million species worldwide. This will generate the vast, contextualised training data required for AI systems to learn from evolution.
BaseData™ broke through the Data Wall. The Trillion Gene Atlas is designed to extend that breakthrough far beyond anything known so far.
The Trillion Gene Atlas is a collaborative project uniting Basecamp Research with hundreds of biodiversity partners worldwide, alongside Anthropic, Ultima Genomics, and PacBio, and powered by NVIDIA AI infrastructure. Together, we are building the biological foundation and biodiversity value chain needed to design new biology at scale.

Advisors to The Trillion Gene Atlas
The Trillion Gene Atlas is guided by an Advisory Board comprising leaders behind some of the largest genomics programmes in history who have advanced genomic sequencing and curation, built community standards, and shaped frameworks for data equity, Indigenous data governance, and access and benefit-sharing.


A global data network built through partnership
At the core of our data are our biodiversity partners. Every sample is collected under informed-consent and benefit-sharing agreements so that the countries and communities who steward this biodiversity share in the value it creates, with each sequence traceable to one of hundreds of country-specific permits. This allows a portion of revenue to be fed back to the country and community where the data was originally sourced.
Setting a standard of data provenance that sets a benchmark that the rest of the field has yet to match.
Our partners contribute invaluable local knowledge and stewardship of biodiversity while we support their priorities through both monetary and non-monetary benefits that continue throughout the life of the collaboration. This collaborative approach is fundamental to enabling EDEN's capabilities.
Read more about EDEN and how we built BaseData.

Leading the way in ABS regulation
We actively contribute to discussions shaping the frameworks and policies that govern access to genetic resources, and the sharing of benefits arising from their use. Members of our team participate in processes under the Convention on Biological Diversity and the Informal Advisory Committee on Capacity Building for the Implementation of the Nagoya Protocol, the Multilateral Mechanism and its Cali Fund Steering Committee, and engage with the UK Government's FCDO on implementation of the BBNJ Treaty. This policy engagement helps ensure that our partnerships remain aligned with evolving international biodiversity governance frameworks.
Keep reading
Contact Us
We collaborate with the top teams across biopharma and are building biodiversity partnerships across the globe. Reach out— we'd love to hear from you.
Looking to join the team? Click here