A former Adaptive Biotechnologies co-founder is building a marketplace for proprietary scientific datasets AI developers can train on without seeing raw data.
A former Adaptive Biotechnologies co-founder is building a marketplace for proprietary scientific datasets AI developers can train on without seeing raw data.

Harell Data Corp., founded by Adaptive Biotechnologies co-founder Harlan Robins, raised $15 million from Fuse and Cercano Management to build a secure platform connecting AI modelers with proprietary scientific training datasets.
"The goal of the company is to connect AI modelers to proprietary training sets to enable the solution of challenging scientific problems," Robins said. "At present, the entities that generate proprietary training data using their own technology and expense do not have a good way to commercialize the data. So they effectively sit on it."
The Bellevue-based startup's platform lets data creators host proprietary datasets in a secure cloud environment, allowing machine learning groups to train models without the raw data ever leaving the system. Data owners earn a direct share of compute revenue generated during training runs, rather than waiting years for speculative downstream drug royalties. The 11-person team is split between Bellevue and Palo Alto, with Rakesh Nair as chief technology officer, Saray Covey as head of operations, and Analise Polsky as head of sales.
The funding arrives as AI developers face a widening gap between publicly available training data and the high-cost, proprietary datasets needed for breakthroughs in computational medicine and materials science. Adaptive Biotechnologies, which trades at a $3.9 billion market value, is among the initial data partners, along with Seattle-based A-Alpha Bio.
High-profile successes like AlphaFold, DeepMind's protein structure prediction system, thrived because they drew on decades of publicly available, experimentally derived protein structures. But in many critical areas of biology and medicine, generating high-quality datasets requires millions of dollars and years of labor. The entities that own this data currently have no secure, profitable way to share it, so it remains sitting in isolated silos.
Harell Data's model addresses this by creating a direct revenue stream for data owners. Instead of licensing data outright or waiting for downstream royalties, data creators earn a share of the compute revenue generated when AI developers train on their datasets. As models improve through training on these rich datasets, the intrinsic value of the underlying data compounds.
While the startup is initially targeting computational medicine — the area Robins knows best from his years at Adaptive Biotechnologies — he emphasizes that the data-silo problem spans scientific disciplines, from materials science to imaging.
"If the business works right, we should be able to enable solutions to really important problems," he said.
Early access for the platform launched this week with initial datasets from A-Alpha Bio and Adaptive Biotechnologies. Adaptive, which earlier this year spun out Digital Biotechnologies to develop clinical DNA sequencing technology, continues to work with Robins as a consultant on scientific strategy.
The company's name carries a personal story: Robins asked his 8-year-old son, Ellis, for a name suggestion, and within seconds the boy proposed "Harell" — a combination of "Harlan" and "Ellis."
"No offense to the large cap cloud compute companies, but it sounds better to me than any of their names, and things seem to have worked out OK for them so I went with it," Robins said.
The funding round, first reported by the Timmerman Report, adds to a growing wave of venture capital flowing into AI-for-science infrastructure. As AI modelers exhaust publicly available training data, platforms that unlock proprietary datasets could become critical infrastructure for the next wave of drug discovery and materials innovation. For investors tracking the AI-in-science theme, the round signals that data ownership is emerging as a distinct monetization layer in the AI value chain, separate from model development and compute provisioning.
This article is for informational purposes only and does not constitute investment advice.