makeprov: Pythonic Provenance Tracking
makeprov is a small library for recording W3C PROV/JSON-LD provenance
around Python functions that read and write files: which inputs produced
which outputs, when, with what code and environment. A decorator wraps a
function, tracks the files it declares as inputs/outputs, and writes a
provenance record after each call. A minimal make-style dependency
resolver and optional bridges from Snakemake and ReproZip are included, but
the core contract of the library is the provenance record — not workflow
orchestration, which tools like Snakemake already do well.
Features
Decorator-based rules that infer dependencies from
InPath/OutPathparameters and write a PROV/JSON-LD record after every call.A clean
Plan → Run → Artifactmodel: the script at a commit is aprov:Plan, the runtime and the user are agents, and the two are tied together byprov:qualifiedAssociation/prov:hadPlan.ArtifactReflets a run cite external entities — a dataset IRI, an object-store key, a model checkpoint — without makeprov copying their metadata.Provenance write failures are fatal by default (
ProvenanceConfig(strict=True)), so a rule can’t silently “succeed” with no record of what it did.Resolve templated targets (
results/{sample}.txt) viaparse-style patterns, and a small dependency resolver (build/build_all) for chaining rules.Serialize provenance as JSON-LD, or as RDF/TriG when
rdflibis installed (pip install "makeprov[rdf]").Optional Snakemake bridge that turns
--d3dagand--detailed-summaryoutput into PROV JSON-LD artifacts ready for inclusion in Snakemake HTML reports.Converts an existing ReproZip trace (
reprounzip graph --json) into the same PROV model viamakeprov-reprozip.
Installation
You can install the module directly from PyPI:
pip install makeprov
Optional extras add RDF/TriG export, CLI subcommand support, or the Snakemake bridge:
pip install "makeprov[rdf]" # rdflib + pyshacl for RDF/TriG export
pip install "makeprov[cli]" # defopt, needed for makeprov.main()
pip install "makeprov[snakemake]" # the makeprov-snakemake bridge
Usage
Here’s an example of how to use this package in your Python scripts:
from makeprov import rule, InPath, OutPath, build
@rule()
def process_data(
sample: int | None = None,
input_file: InPath = InPath('data/{sample:d}.txt'),
output_file: OutPath = OutPath('results/{sample:d}.txt')
):
with input_file.open('r') as infile, output_file.open('w') as outfile:
data = infile.read()
outfile.write(data.upper())
if __name__ == '__main__':
# Build a specific templated target and its prerequisites
from makeprov import build
build('results/1.txt')
# Or expose rules via a command line interface
import defopt
defopt.run(process_data)
You can execute examples/example.py via the CLI like so:
python examples/example.py build-all
# Or set configuration through the CLI
python examples/example.py build-all --conf='{"base_iri": "http://mybaseiri.org/", "prov_dir": "my_prov_directory"}' --force --input_file input.txt --output_file final_output.txt
# Or set configuration through a TOML file
python examples/example.py build-all -c @my_config.toml
# Inspect dependency resolution without executing rules
python examples/example.py --explain results/1.txt
python examples/example.py --to-dot results/1.txt
For directory outputs, nested/merged provenance, streaming, opt-in rule metadata, Snakemake integration, and other advanced topics, see the full usage guide and configuration reference.
The provenance model
makeprov keeps PROV’s distinction between the plan (the recipe) and the agent (whoever carried it out):
run.py @ git SHA a prov:Plan, schema:SoftwareSourceCode
CPython 3.11 a prov:Agent, prov:SoftwareAgent
you (opt-in) a prov:Agent, schema:Person
train-20260910T…-c2f6dc7b a prov:Activity
prov:used dataset-X, the Python environment
prov:wasAssociatedWith runtime, person
prov:qualifiedAssociation [ prov:agent person ; prov:hadPlan run.py ]
results/model.txt a prov:Entity
prov:wasGeneratedBy train-20260910T…-c2f6dc7b
dct:identifier sha256:…
Keeping the plan and the agent apart is what makes the graph mappable onto
Workflow Run RO-Crate,
whose instrument (the software that was run) and agent (a Person or
Organization) are separate slots:
makeprov / PROV-O |
Process Run Crate |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
The schema:Person agent is off by default: provenance documents are
routinely committed and published, and a name and email address are personal
data you should choose to publish rather than emit by accident. Turn it on with
ProvenanceConfig(record_user=True), or --record-user on the Snakemake
bridge. Without it, the qualified association names the runtime as the
responsible agent.
Note that WRROC is Schema.org-native and defines no normative PROV-O mapping;
the table above is a practical alignment, not an OWL equivalence. makeprov’s
own vocabulary stays prov:/schema: — RO-Crate and OpenLineage are intended
as adapters over this model rather than changes to it.
Planned structure vs. observed execution
By default a document is purely retrospective: it records the activities that ran. A rule that was already up to date contributes nothing, because asserting an execution that did not happen would be worse than saying nothing.
Set emit_plan_graph = true (or pass --plan-graph) to additionally emit the
prospective structure — each rule as a prov:Plan in its own right, linked to
the plans it depends on by dct:requires:
run.py#rule-transform a prov:Plan
dct:requires run.py#rule-extract
dct:source run.py
The activity’s prov:hadPlan then points at the specific rule rather than the
whole script. The Snakemake bridge does the same, collapsing the job DAG’s
edges to rule-level dct:requires edges.
Forge profiles
When no base_iri is set, makeprov derives one from the git remote. The
supported hosts are declared in forges.toml —
GitHub, GitLab, Bitbucket, Forgejo/Gitea/Codeberg and SourceHut — each giving
the permalink layout for that host:
[[forge]]
name = "gitlab"
hosts = ["gitlab.com"]
blob = "{repo}/-/blob/{revision}/"
Point forge_profiles at your own TOML file to add self-hosted instances;
entries there are matched first, so they can also override a built-in host.
SSH and scp-style remotes (git@host:owner/repo.git) are understood, and any
credentials embedded in a remote URL are stripped before it reaches a document.
Referencing things that aren’t local files
ArtifactRef describes an entity a run consumed or produced. It is either
local (makeprov stats and hashes it) or external (makeprov records the IRI
and never touches the filesystem):
from makeprov import ArtifactRef, OutPath, rule
@rule()
def train(
dataset: ArtifactRef = ArtifactRef.external(
"https://example.org/datasets/train-v17",
types=("prov:Entity", "schema:Dataset"),
digest="sha256:...",
),
model: OutPath = OutPath("models/m.pkl"),
):
...
The external object keeps its own detailed metadata; makeprov only records that this run used its stable IRI. External refs take no part in staleness checks, since they have no local mtime to compare.
More examples
examples/complex_example.py— a CSV-to-RDF workflow that aggregates multiple inputs and embeds anrdflib.Graphresult directly into the provenance dataset.examples/merge_outdir_example.py— bundling nested provenance and directory outputs withmerge=TrueandOutDir/InDir.examples/context_demo_example.py— pinning a base IRI and isolating rules and buffers in their ownSession.
Walkthroughs of these, plus streaming/recovery mode and opt-in rule metadata, are in the usage guide.
Snakemake workflows
Install the snakemake extra (pip install "makeprov[snakemake]") to get the
makeprov-snakemake command, which shells out to Snakemake and converts its
job DAG and --detailed-summary metadata into a PROV document. See the
Snakemake integration guide
for the CLI flags and an example report() wiring.
makeprov-snakemake --prov-path prov/snakemake -- --snakefile Snakefile --nolock
ReproZip traces
makeprov-reprozip converts an existing reprounzip graph --json file into
PROV/JSON-LD or TriG — observed file accesses and process relationships only,
with no ReproZip runtime dependency. See the
ReproZip conversion guide.
makeprov-reprozip graph.json --output prov/command --base-iri https://example.org/my-experiment/
Configuration
You can customize the provenance tracking with the following options:
base_iri(str): Base IRI for new resourcesprov_dir(str): Directory for writing PROV.json-ldor.trigfilesforce(bool): Force running of dependenciesdry_run(bool): Only check workflow, don’t run anythingstrict(bool, defaultTrue): Raisemakeprov.ProvenanceWriteErrorif a rule’s provenance record fails to write, instead of only logging a warning. A rule that produces a result but no provenance record is treated as a failure by default; setstrict=Falseto opt out per-rule or globally.run_id(str | None): Adopt an externally supplied run identity, such as a CI job id. When unset, each run gets a fresh unique id.record_user(bool, defaultFalse): Record the invoking user, taken fromgit config user.name/user.email, as aschema:Personagent. Off by default so personal data isn’t published by accident. CLI:--record-user.emit_plan_graph(bool, defaultFalse): Also emit prospective structure — the rule dependency graph asprov:Plannodes linked bydct:requires. Off by default, so a document describes only what actually ran. CLI:--plan-graph.forge_profiles(str | None): TOML file of extra forge profiles, for self-hosted git hosts. CLI:--forge-profiles.
See CHANGELOG.md for what changed in past releases, including
the 0.7 provenance-model rework.
Scoped spans and cached downloads
Use makeprov.span(label, prov_path=None, frame=None, context=None) as a
context manager or decorator to bracket a chunk of work in its own provenance
buffer, and makeprov.CachedDownload to lazily fetch and record provenance for
a remote resource cached locally. See the
usage guide for examples of
both.
Documentation
Build the Sphinx docs locally (including autosummary API stubs) with the docs extra so that the CLI dependencies needed for imports are available:
pip install -e ".[docs]"
python docs/build.py
Contributing
Contributions are welcome! Please open an issue or submit a pull request.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Contents
- Home
- Usage guide
- Defining rules
- Building dependency graphs
- Command-line entry point
- Streaming input and output
- Tracking outputs within directories
- Merging provenance across nested rules
- Opt-in rule metadata
- Scoped spans and explicit outputs
- Environment: declared vs. resolved
- Caching remote downloads
- Multi-file RDF export example
- Pinning context and isolating sessions
- Controlling provenance framing
- Configuration
- Provenance output
- Snakemake integration
- ReproZip graph conversion
- API reference
- Changelog