Build the typed example¶
This tutorial joins three synthetic samples to two people. It demonstrates the source inventory, typed mappings, exclusions, provenance, BioCypher export and source drift.
Set up¶
Use Python 3.12, uv, Git
and Node.js on PATH. Clone the example from the release matching the package:
git clone --depth 1 --branch biotope-v0.10.0 https://github.com/biocypher/biotope.git biotope-example
cd biotope-example
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python 'biotope[graph]==0.10.0'
source .venv/bin/activate
If Node.js is unavailable, install pyright[nodejs] in this environment as
shown in installation.
Copy the example to a temporary project and initialize metadata tracking:
example=$(mktemp -d)
cp -R examples/typed_graph/. "$example/"
cd "$example"
biotope init . --no-git --no-prompt
The example supplies a completed graph/ workspace, its reviewed Croissant
description in metadata/study.jsonld and a project.yaml with its purpose.
Initialization preserves that file. In a new project, biotope graph scaffold
creates an empty layout to complete.
Register and generate sources¶
The fixture's description covers both CSV record sets:
biotope source register metadata/study.jsonld --name study --reason "Reviewed synthetic fixture-v1"
biotope source generate .biotope/datasets/study.jsonld --out graph/sources
biotope map inspect .biotope/datasets/study.jsonld
The example already contains a package per record set,
graph/sources/study/people/ and graph/sources/study/samples/, so generation
reports both as current and writes nothing except their contract revisions in
.biotope/contracts/. In a new project it would scaffold each package: a
schema.py to edit, a SOURCE registration and a placeholder loader. It never
rewrites an existing file.
Read these files to follow the transformation:
| File | Role |
|---|---|
graph/sources/study/samples/schema.py |
Declares the source fields, their types and the reviewed revision. |
graph/sources/study/samples/loader.py |
Reads CSV and decodes the score as a float. |
graph/sources/__init__.py |
Selects every source of the inventory; excludes none. |
graph/mappings/sample/__init__.py |
Doubles the score and constructs namespaced graph IDs. |
graph/pipelines/build_graph.py |
Joins samples to people and records exclusions. |
The SOURCES, TOPOLOGY and MAPPINGS registries select the definitions used by
the pipeline. See the authoring guide for the full layout.
Check and build¶
biotope graph check
biotope graph build --out graph/build/first
biotope graph build --out graph/build/second
Expect three domain nodes and two domain edges: Ada, samples s1 and s2, and
their relations. The sample scores are 3.0 and 5.0. Checks warn about the
illustrative example: namespace.
Sample s3 has no matching person; Bea has no selected sample. Both exclusions
appear in run.json. Ada's provenance includes her people row and the two
contributing sample rows.
Review the result¶
Inside graph/build/first/:
- Read
run.jsonfor completion, scope, policies, counts, exclusions and audits. - Read
schema_config.yamlfor each exported label (Sample,Person,FromPerson) with its concept ID and the concept and property descriptions. - Pair each
biocypher/*-header.csvwith its*-part000.csvto inspect exported values. - Follow a row's
biotope_provenance_idintoprovenance.jsonfor its mappings and source locations.
Compare graph_digest in both run records; it should match despite different
output paths and timestamps. Without --out, biotope graph build writes
graph/build/ and replaces it on the next successful build. It refuses a
graph/build/ that holds other directories, such as the two builds above, so
delete them first.
To view the topology with build observations:
Open graph/metagraph.html in a browser. The viewer works offline.
Change a source contract¶
Change the score's declared type in metadata/study.jsonld from sc:Float to
sc:Integer, then register the revised description and reconcile the sources:
biotope source register metadata/study.jsonld --name study --reason "Score reviewed as an integer" --replace
biotope source generate .biotope/datasets/study.jsonld --out graph/sources
biotope graph check
Generation reports graph/sources/study/samples as drifted and never rewrites the
schema; people stays current. The check fails with source.drift for
study/samples, naming the change:
field samples/score: dataType "sc:Float" -> "sc:Integer"
changed /recordSet/0/field/3/dataType: "sc:Float" -> "sc:Integer"
Review what the change means for the schema, the loader and the mappings. Here the
loader keeps parsing the score as a float, so only the acknowledgement remains:
set __source_digest__ in graph/sources/study/samples/schema.py to the revision
the finding names, then run biotope graph check again. It passes.
The temporary project can be retained for experimentation. graph/build/ and
graph/metagraph.html are derived outputs; preserve curated metadata, the
.biotope/contracts/ history and authored code if you want to reproduce a run.