A practical, end-to-end walkthrough for putting your datasets and ML pipeline
under version control with DVC (Data Version Control), and for using the
data_versioning toolkit in this product alongside it. Everything here uses
anonymized placeholders (your-org, s3://your-org-dvc-store); substitute your
own bucket and paths.
Git is excellent for code and terrible for data: a 4 GB parquet file does not
belong in a git history. DVC solves this by storing a tiny pointer file
(*.dvc) in git while the real bytes live in a content-addressable cache that is
pushed to remote storage (S3, GCS, Azure Blob, SSH, or a shared filesystem).
The mental model:
| Concern | Lives in git | Lives in the DVC remote |
|---|---|---|
Code, dvc.yaml | yes | no |
params.yaml, metrics | yes | no |
*.dvc pointer files | yes | no |
| Datasets, models | no (just the pointer) | yes (keyed by content hash) |
The data_versioning package complements DVC in two places:
(CI scratch, notebooks, ad-hoc audits), hashing.py + snapshots.py give you
the same content-addressable "did this dataset change?" check with zero
infrastructure.
lineage.py, pipeline.py, and reproducibility.py letyou generate dvc.yaml programmatically, render the pipeline as a graph, and
capture the run environment that DVC does not record on its own.
DVC is distributed on PyPI. Install the base package plus the extra for your
remote backend:
pip install "dvc[s3]" # Amazon S3 / S3-compatible (MinIO, Cloudflare R2)
pip install "dvc[gs]" # Google Cloud Storage
pip install "dvc[azure]" # Azure Blob Storage
pip install "dvc[ssh]" # any host you can reach over SSHVerify the install and pin the major version in your project so teammates and CI
agree:
dvc --version # expect 3.x
echo "dvc>=3,<4" >> requirements.txtRun dvc init inside an existing git repository — DVC layers on top of git:
git init # if you have not already
dvc init
git add .dvc .dvcignore
git commit -m "chore: initialize DVC"dvc init creates a .dvc/ directory (config + local cache) and a .dvcignore
file (same idea as .gitignore, but for DVC's own scans).
The remote is where dataset bytes are pushed. Add one and mark it the default
with -d:
# S3 (or any S3-compatible endpoint)
dvc remote add -d storage s3://your-org-dvc-store/ml
dvc remote modify storage region eu-north-1
# Google Cloud Storage
dvc remote add -d storage gs://your-org-dvc-store/ml
# Azure Blob
dvc remote add -d storage azure://your-org-container/ml
# SSH / on-prem
dvc remote add -d storage ssh://[email protected]/srv/dvcCommit the remote config (it contains no secrets — credentials come from your
environment / cloud SDK):
git add .dvc/config
git commit -m "chore: configure default DVC remote"Credentials never go in git. DVC reads AWS/GCP/Azure credentials from the
usual environment variables,
~/.aws/credentials, or instance metadata. Keepsecrets in your secret manager, exactly as the anonymization rules require.
Hand a file or directory to dvc add. DVC moves the bytes into its cache,
writes a small *.dvc pointer, and updates .gitignore so the raw data is not
accidentally committed:
dvc add data/raw
git add data/raw.dvc data/.gitignore
git commit -m "data: track raw dataset with DVC"
dvc push # upload the bytes to the remoteThe pointer file is what your teammates pull — dvc pull fetches the matching
bytes by hash.
Before and after dvc add, you can sanity-check the dataset's content hash with
the bundled hasher (no DVC call needed):
from data_versioning import hash_directory
print(hash_directory("data/raw").short_hash) # stable, deterministic idIf that short hash matches what a teammate sees, you are looking at byte-identical
data — independent confirmation that dvc pull gave you the right version.
dvc.yaml)A pipeline turns ad-hoc scripts into a reproducible DAG. See dvc/dvc.yaml in
this product for a complete four-stage example. The shape of one stage:
stages:
prepare:
cmd: python src/prepare.py
deps:
- src/prepare.py
- data/raw
params:
- prepare.seed
- prepare.test_size
outs:
- data/preparedYou can also generate dvc.yaml programmatically with pipeline.py, which
keeps stage definitions in code and avoids hand-editing YAML:
from data_versioning import Pipeline, Output
pipe = Pipeline()
pipe.add_stage(
name="train",
cmd="python src/train.py",
deps=["src/train.py", "data/features"],
params=["train.n_estimators", "train.max_depth"],
outs=[Output("models/model.pkl", cache=True)],
metrics=[Output("metrics.json", cache=False)],
)
pipe.write("dvc.yaml") # emits valid YAML, no PyYAML dependencydvc repro # run only the stages whose inputs changed
dvc dag # print the pipeline graph in your terminal
dvc metrics show # show metrics from metrics.json
dvc params diff # show which params changed vs the last commit
dvc metrics diff # show how metrics moved vs the last commitdvc repro is the heart of reproducibility: because every stage declares its
deps, params, and outs, DVC knows exactly what is stale and reruns the minimum.
# pull code + data to the exact versions recorded in this commit
git pull
dvc pull
# ... make changes (edit code, bump a param, refresh data) ...
dvc repro # rebuild what changed
git add dvc.lock metrics.json params.yaml
git commit -m "exp: tune max_depth=8"
dvc push # upload any new/changed artifacts
git pushdvc.lock is the auto-generated record of the exact hashes used in the last
successful run — always commit it. Together, git + dvc.lock + the remote let
anyone recreate your result byte-for-byte.
A minimal CI job that verifies the pipeline still reproduces:
# .github/workflows/dvc.yml (illustrative)
steps:
- uses: actions/checkout@v4
- run: pip install "dvc[s3]" -r requirements.txt
- run: dvc pull # fetch data by hash
- run: dvc repro --dry # fail if anything is stale/undeclared
- run: dvc metrics diff --targets metrics.jsonFor pull requests, dvc metrics diff main posts a clean before/after table so
reviewers can see the effect of a change on model quality.
| Symptom | Likely cause / fix |
|---|---|
dvc push uploads nothing | Bytes already in the remote (deduplicated by hash) — expected. |
dvc pull fails with "missing in remote" | Someone forgot to dvc push; or you are pointed at a stale remote. |
dvc repro reruns everything | A deps path changed mtime/content, or a tracked dir is non-deterministic. |
| Output "changed" every run | Stage writes nondeterministic bytes (timestamps, unsorted order). |
| Huge git history | A dataset was git add-ed instead of dvc add-ed. Untrack it. |
For the "changed every run" class of problems, hash the output directory with
data_versioning.hash_directory twice in a row: if the tree_hash differs, the
nondeterminism is in *your* code, not DVC. See reproducibility-guide.md for the
usual culprits and fixes.
reproducibility-guide.md for seeds, environment capture, and thereproducibility checklist.
python examples/version_dataset.py to watch snapshots + diffs in action.python examples/reproduce_experiment.py to see a generated dvc.yaml,a lineage graph, and an environment-drift report.
Questions: [email protected].
Get the full ML Data Versioning and unlock everything.
Get the complete guide with every chapter unlocked, including code samples, diagrams, and best practices.
Access all interactive tools with complete data, all workload profiles, and the full scenario library.
Downloadable source code, configuration files, and working examples from every chapter.
Free updates for life. Every new chapter, tool, and improvement included.