Typed graph projects¶
Croissant describes sources; Python defines the graph, its loaders and its transformations. Biotope inventories every described source, checks the definitions and exports validated graph objects through BioCypher. Start with the runnable tutorial or use this guide to author a project.
Create the workspace¶
Install biotope[graph] using the installation guide, then run
from your project root:
This creates graph/ with empty registries, path constants, dependencies, a
generated source inventory and a pipeline to complete; its README describes the
layout. Scaffolding works independently of init and add; it refuses an existing
graph/ directory or symlink.
Keep raw data, purpose files and managed metadata at the project level, and graph
code under graph/. Use PROJECT_ROOT and GRAPH_ROOT from graph/paths.py and
package-relative imports for portability.
| Artifact | Ownership and purpose |
|---|---|
.biotope/datasets/<manifest>.jsonld |
Effective source description, reviewed by the project |
.biotope/contracts/ |
Recorded contract revisions, written by source generate |
graph/standardization.py |
Shared terms that fields of several sources bind to |
graph/sources/<manifest>/__init__.py |
Generated root collecting that manifest's packages |
graph/sources/<manifest>/<source>/schema.py |
Record class and field bindings; scaffolded once, then authored |
graph/sources/<manifest>/<source>/__init__.py |
SOURCE registration, created when absent |
graph/sources/<manifest>/<source>/loader.py |
Authored physical reader, created as a placeholder when absent |
graph/sources/inventory.py |
Generated INVENTORY of every source package |
graph/sources/__init__.py |
Authored SOURCES selection and reasoned EXCLUDED_SOURCES |
graph/alignment/ |
Cross-source identity resolution, run before mappings |
graph/topology/<concept>/ |
Node and relation dataclasses, collected in TOPOLOGY |
graph/mappings/<concept>/ |
Typed transformations, collected in MAPPINGS |
graph/pipelines/ |
compose.py stage sequence; build_graph.py PIPELINE |
graph/build/, graph/metagraph.html |
The current build and topology viewer; derived output |
For earlier projects, see migration.
Describe and inventory the sources¶
Use biotope add <path> to describe selected local data, then
biotope map inspect <manifest> to review the resulting metadata. These operations
have different access requirements: baking reads source bytes; inspection reads
metadata only.
Write evidence-based corrections in a separate Croissant file under
.biotope/reviews/ and register the complete description:
biotope source register .biotope/reviews/study.jsonld --name study \
--reason "Reviewed source headers and field types"
biotope source generate .biotope/datasets/study.jsonld --out graph/sources
Use --replace when replacing an existing description. Registration reports
removed keys and identifiers but does not merge documents or judge the review
reason. Reconcile annotations and structural changes before replacement.
Curated or annotated manifests are protected from add --rebake and add --force.
To refresh one, bake separately and review the differences:
biotope add raw/study --bake-to .biotope/reviews/study-fresh.jsonld
# Reconcile this file with the existing curated description before registering it.
biotope source register .biotope/reviews/study-fresh.jsonld --name study \
--reason "Reconciled new data with reviewed metadata" --replace
Here study must be the name of the managed description being replaced.
--bake-to takes one input and a new output outside that input and
.biotope/datasets/. It does not update tracking. Source locations in the fresh
description retain their original data root.
Baker leaves files it cannot parse, such as papers and notes, unclaimed. Add a
FileObject with a stable @id, its contentUrl and sha256 for each such file
the graph may read; it then becomes a source like any table.
Generation gives every top-level record set, and every file that no record set
reads, its own package. It creates missing files only and never rewrites an
existing schema, registration or loader. The generated inventory collects every
package; in graph/sources/__init__.py, select each source or list it in
EXCLUDED_SOURCES with a reason. biotope graph check reports a source that is
neither, a selected placeholder loader and a described input without a package.
See generation details for naming,
ownership, conflicts and removed sources.
Each schema starts from the description: known scalar, nested and repeated types
become Python declarations, fields are nullable unless curated metadata asserts
otherwise, and unknown types and arrays stay opaque. The schema is then yours to
edit. Bind each Croissant field once with source_field, declare source-local
missing-value tokens and aliases, and bind fields to shared terms from
standardization.py only where sources share a meaning:
@dataclass(frozen=True, kw_only=True)
class Samples:
__record_set__: ClassVar[str] = "samples"
__source_digest__: ClassVar[str] = "232b7860…"
__missing_values__: ClassVar[frozenset[str]] = frozenset({"", "na"})
sample_id: str # binds samples/sample_id
score: float = source_field("raw_score", term=SCORE)
tissue: str | None = source_field(None, default=None) # not supplied here
Review source drift¶
__source_digest__ records the contract revision you reviewed. When the
description changes, biotope graph check reports source.drift for each affected
schema, with a field-level summary and every JSON-pointer difference against the
recorded revision. Review the change against the schema and loader, update them,
then set __source_digest__ to the revision the finding names. Unaffected sources
stay current. biotope source generate records each revision it sees; --check
writes nothing.
Define topology and identity¶
Use frozen dataclasses with stable schema_id values. Give each node identifier
its own NewType over str, and use those types for relation endpoints:
from dataclasses import dataclass, field
from typing import ClassVar, NewType
PersonId = NewType("PersonId", str)
@dataclass(frozen=True)
class Person:
"""One person referenced by at least one selected sample."""
schema_id: ClassVar[str] = "study:person"
id: PersonId
name: str = field(metadata={"description": "Name as recorded by the source."})
Mint namespaced IDs such as study-a:person-17 explicitly. Static ID types prevent
accidental endpoint swaps; they do not establish uniqueness or shared biological
identity. Runtime checks reject conflicting objects with the same ID and resolve
edges against the declared endpoint types.
Concept docstrings and property descriptions are exported in schema_config.yaml,
and graph check warns about concepts and properties without one. Optional
display_name: ClassVar[str] values provide short diagram labels without changing
concept IDs. The biotope: namespace is reserved for system metadata.
Graph properties support strings, booleans, integers, finite floats, nullable
scalars and string lists. Lists cannot contain nulls. For the Neo4j import files,
strings are one line and list items contain no ; or line break; a build stops at
the first graph object that breaks this and names its mapping, concept and
property. Convert richer source values deliberately in the mappings.
Implement loaders and mappings¶
A loader returns Iterable[SourceRecord[Row]]. Each record carries a typed value
and Evidence(artifact, version, record_set, location). Use a checksum or known
release version where available; label unverified fingerprints honestly.
Read and decode files inside loader functions using established libraries, and
delete the placeholder's # biotope:placeholder first line once the loader is
implemented. Source validation performs no coercion and applies no declared alias
or missing-value token, so loaders must handle missing tokens, aliases, number
parsing and source-specific representations. Loaders preserve rows:
they do not filter, impute or deduplicate. Include the source location in decoding
errors.
Mapping parameters are named SourceRecord[...] values, accessed through .value.
The return annotation declares the output dataclasses. Mapping(name=..., function=...) registers the function, with optional requirements and evidence.
The tutorial project shows a complete implementation.
A mapping can return topology objects directly. Introduce an intermediate
dataclass when sources need a shared normalized shape, a stage requires buffering,
or several mappings reuse a transformation. Intermediates have no schema_id.
Keep file access in loaders and transformations in mappings.
Inside the pipeline:
context.load(...)validates source records as they are read.context.apply(mapping, ...)returns typed records for another transformation.context.map(mapping, ...)collects outputs that declare graph identities.context.exclude(...)records excluded records under a declared policy.context.record_audit(...)records a stage's grain, selection rule and named counts.
Ordinary Python owns joins, indexes, aggregation and resource lifetimes. State the join keys, cardinality, unmatched policy and output grain. Account for omitted records through exclusions or audits. Every graph object enters through a registered mapping; there is no direct emission API.
Resolve cross-source identity in graph/alignment/ before any mapping admits a
record: gather every source's identifiers first, so the result does not depend on
read order, and keep each resolution's contributing rows as evidence. Write the
stage sequence in pipelines/compose.py, then complete the Pipeline in
pipelines/build_graph.py with scope, policies, settings and code paths; the
scaffold already wires SOURCES, INVENTORY, EXCLUDED_SOURCES and TERMS. Bind purpose requirements with entity:<exact intent text> or
relation:<exact intent text> keys, or record a reason in deferrals.
Imports must contain declarations only. Source access and execution belong inside explicitly called functions. Biotope imports project Python; this is an authoring contract, not an execution sandbox.
Describe what the graph means¶
The consumer receives the graph and graph/build/, not the project. Interpretation
therefore travels in three places: concept and property descriptions in
schema_config.yaml, the pipeline's scope, and its policies, one per exclusion
reported through context.exclude(...). State in them what each value measures,
against which reference, when identifiers can be joined, and what the selection
left out. Record open scientific questions in graph/ASSUMPTIONS.md.
Separate selection from query filters. A graph filtered to strict_p < 0.05 can
describe measurements that passed that selection. It cannot identify all
measurements passing another adjustment merely because the retained rows also
contain an open_p column. If two readings need different populations, admit
their union and let the query choose.
Biotope checks structure, not science. Definition checks assess declarations; integrity and quality checks assess emitted objects. None of them can see a record that should have been emitted but was not. Before relying on a claim, compare the graph with expectations read independently from the sources or a published result, never from the pipeline's own selection, and compare missing and unexpected records by namespaced ID.
Check, execute and review¶
Run from the project root:
graph check checks the source inventory, drift, definitions, bindings and Python
types without invoking loaders. graph metagraph imports topology independently
and writes graph/metagraph.html. graph build checks and executes the pipeline
once, validates the graph and replaces graph/build/ after success; a failed
rebuild keeps the previous build and writes graph/build/last_failure.json.
Use biotope graph quality when you want to execute and assess without exporting;
it prints the assessment. Running quality and build separately executes the
pipeline twice. Counts, connectivity and missing-property observations are
advisory; definition, execution and integrity failures block a build.
Review run.json, schema_config.yaml, provenance.json and exported values.
Export labels are the concept's local name, such as Sample for example:sample,
widened only when two concepts collide. Every exported object carries
biotope_provenance_id, an index into provenance.json.
Use biotope graph metagraph --report graph/build/run.json to add build
observations to the viewer. See technical notes for report
limits, source attribution, reproducibility and export formats.