The experiments you already ran
Every lab already owns the thing it most needs and cannot use: its own measured history — too big to hold in one head, too lossy once it has been reduced to papers. A data space turns that history into objects, and then lets a model trained on measured biology extend them, on your own machine, with nothing uploaded.
Open your lab's data drive and count the folders. One per dataset, most named after the month they arrived. The RNA-seq from the knockdown line. The methylation array a collaborator sent two years ago. Last year's proteomics, the run that never made it into a paper. The screen a rotation student finished the week before they left.
Each one was analyzed once — carefully, by someone competent, at the moment it landed — and then it stopped. A figure went into a slide, the slide went into a talk, the folder went quiet.
Every lab already owns the thing it most needs and cannot use: its own measured history — too big to hold in one head, too lossy once it has been reduced to papers.
Two ways a result stops being usable
It is too big to hold in one head. A counts matrix and a set of methylation β-values are not comparable by looking. Five datasets side by side is not a reading task; it is a project — someone has to reopen each one, remember what was in it, and reconcile it by hand. So comparison across datasets is work you schedule, which means it is mostly work that doesn't happen.
It is too lossy once it is a paper. A publication carries the one comparison that survived review. The other eleven things that dataset showed — the pathway that moved but not enough, the gene that went up in one line and down in another, the effect that was real but off-topic — were true then and are nowhere now.
Between the file, which nobody can hold, and the paper, which holds almost nothing, sits the only complete copy: someone's memory. The map of which experiments spoke to each other, what was already tried, where two results disagree, what the old screen was actually for. It is the most valuable object in the building, it has never been written down, and it walks out the door with the person who has it.
Three questions a lab answers from memory
Every group asks these constantly and answers all three by thinking hard and hoping:
- Have we seen this before?
- Where do we disagree with ourselves?
- What did the array from two years ago know about the run that finished this week?
Two stop needing memory the moment your results become objects. The third needs something else, and it is the reason the rest of this post exists.
Results as objects
A data space is a place you put datasets so they can be in the same room.
You drop files in — a counts table, a differential-expression result, an array, a screen, a whole project. ARC reads each one on your machine and writes back a Dataset Card: a structured record of what that dataset is and what it showed. Organism, tissue, condition, the comparison that was run, the design, the sample count. Then the findings — the top features with their direction and effect size, the enriched processes, the insights worth keeping.
The space shown here is a worked example, built so the mechanics can be shown on datasets we are free to publish — read the biology as shape rather than as findings. Everything the interface reports about the run is real: it was a real run against the model, and the numbers are the ones it produced.
The card is metadata, not data, and two things follow from that.
It is small — small enough that a hundred of them fit in one view, which is the entire point, since the question you couldn't ask was never about one dataset. And it is portable in a way the data is not. A card can be held in a graph, synced to a team, or handed to a collaborator while the ten gigabytes it describes stay in the folder they have always been in.
Cards then link where they genuinely share biology, and the graph draws itself.
That answers the first question outright. An entity several datasets name becomes a hub they all hang off, so have we seen this before stops being a memory exercise: the gene from this week's analysis is already a node, already touching the datasets that named it. In the space above, one entity is named by four separate datasets across three different assay types.
It answers the second question too, because direction travels with the finding. A hub carries what each dataset reported about it, so up in four, down in one is a property of the graph rather than something a person has to reconstruct — and that disagreement is a real scientific object, one that was invisible while those datasets lived in different folders. Names are reconciled as they arrive, so the way your lab wrote things down in 2019 and the way it writes them down now stop being two vocabularies.
The unit has shifted. Not the dataset, which is too big to reason over, and not the paper, which is too lossy — the unit is the finding. Once findings are objects, the relationships between experiments become something a machine can work out instead of something a person has to remember.
The question the words cannot answer
Look again at the three datasets sitting alone in that graph.
They are not unrelated to the others. One is a methylation array of the same tissue, one is a plasma panel from the same patients, and one is a screen in the cell line the whole program is built on. They are alone because a data space links datasets that name the same thing, and they name nothing in common.
That is the honest ceiling of everything described so far: a graph built from your words can only ever restate what you already wrote down. In biology that ceiling is low, and it is lowest exactly where the value is:
- A methylation array names CpG probes. An RNA-seq names genes. Same tissue, same disease, same lab, and not one shared string between them.
- A mouse study names Trp53. A human study names TP53.
- A screen names a phenotype. Nothing in it is a molecule at all.
Reaching for a language model helps, and then stops helping, at a second ceiling one layer up. A general model can connect two of your genes because someone once wrote a sentence connecting them. That is genuinely useful, and it is bounded by what has been published — which is the wrong bound when the most interesting thing in the building is an unpublished result in a file nobody outside your group has opened.
To connect the array to the sequencing run, you need something that knows biology as measurement rather than as text.
Discover
Discover hands the entities in your space to Archimedes — Heureka's own multi-omics model, trained on hundreds of thousands of real biological samples rather than on the literature about them — and draws what comes back as a separate layer over your graph, never into it.
This is the right moment to be precise about what moves. Your datasets do not go anywhere; they are read on your machine and they stay there. What leaves is the question — a list of entity names — and the app shows you that list, grouped and named, before you start. A run uses the model and nothing else: no literature, no outside databases, so every relationship it returns is attributable to one source.
There is the third question, answered. The array that shared no vocabulary with anything is now joined to the liver RNA-seq and to the mouse study, through the genes its probes sit in — a hop no amount of string matching could have made, because nothing in either file ever wrote it down.
Why the headline number is small on purpose
The receipt leads with bridges: pairs of your datasets that had no link before and have one now. Not relationships added, not "the graph grew 40%." A growth number rewards padding, which is the exact failure this feature could have — so the metric was chosen to fall when a run pads itself. And the pairs are named, so the number is checkable rather than believable.
That last part is worth reading closely, because it is the posture the whole feature rests on. The run's own notes, verbatim:
LRAT: neighbor list thematically unrelated (brain/eye tissue markers, distances >0.51); no retinol/stellate-cell signal recovered despite LRAT being the canonical HSC marker — dropped.
CYP2E1, EPCAM, ADIPOQ, MKI67: neighbor lists coherent but confirmed only expected tissue-identity panels — omitted as obvious rather than informative.
An agent declining to fill its own scoreboard: reporting that a marker it expected to work didn't, and discarding correct results for being unsurprising. Relationships derived from non-human symbols are flagged on the same receipt as hypotheses, because that is what they are. If the model can't be reached, the pre-flight count is labeled an estimate instead of presented as fact. If a run is cut short, what it found is kept and labeled partial.
Every link is inspectable down to its evidence.
Convergence from two directions is the strongest signal available here, and you can see the convergence rather than take it on trust. This is what makes a result usable: not that a system asserted it, but that you can follow it back and decide for yourself.
It compounds
A space is not a report you generate once.
Runs accumulate. A second run picks up where the first stopped and never re-asks anything already covered, including entities that came back empty, so adding one dataset costs one dataset's worth of work rather than the whole graph again. On a shared space, a colleague's run merges with yours instead of replacing it, and the layer records who ran what — which is also what makes the map survive the person who made it.
And the graph gets better without being asked. Adding a dataset can turn a relationship the model already returned into a bridge, because a pair that was connected only through the model is now connected through your own new result too. That reclassification is pure local computation — no new run, nothing spent. Every experiment you do makes the previous ones more valuable. This is the only asset in the building with that property.
Why this hasn't existed
Both halves exist in the world. What hasn't existed is a place they can meet, because meeting has always required one of them to move.
Lab data platforms hold the record — samples, results, permissions, an audit trail — and hold it well. They have no model of measured biology to extend that record with, and reaching them at all means the record lives on their infrastructure.
Frontier assistants are extraordinary, and they are trained on text. They will tell you what has been written about your entities, which is a real capability with a hard edge: the thing you most want connected is the thing nobody has published. They also want your unpublished data uploaded before the useful part begins, which for many groups is a legal review, an IRB conversation, or a flat no.
Biomedical knowledge graphs are built from the literature and public databases. They are rich, and your results are not in them — and cannot be, without you sending them.
A very good bioinformatician can do any one of these comparisons, and do it better than any of the above. What no person can do is all of them, across every dataset the lab has ever produced, standing, forever.
A data space is the one place a lab's unpublished evidence and a model of biology at large can meet without either one having to move. Putting them in the same graph required a local-first record and a proprietary model of measured biology to be built by the same people, deliberately, for this.
Where this goes
Today a data space is something you read. You open it, you find that two experiments run fourteen months apart are connected through a gene neither of them was about, and you go and look.
The direction is a space you ask. Every dataset a lab produces from here deposits a card, every card thickens the graph, and the question stops being what does this dataset say and becomes what does everything we have ever measured say, together.
The end state is not a better graph. It is a lab that chooses its next experiment against everything it has ever measured, rather than against what someone else remembered to publish.