Contents

Chapter 1

DVC Setup Guide

A practical, end-to-end walkthrough for putting your datasets and ML pipeline

under version control with DVC (Data Version Control), and for using the

data_versioning toolkit in this product alongside it. Everything here uses

anonymized placeholders (your-org, s3://your-org-dvc-store); substitute your

own bucket and paths.


1. Why DVC (and what this toolkit adds)

Git is excellent for code and terrible for data: a 4 GB parquet file does not

belong in a git history. DVC solves this by storing a tiny pointer file

(*.dvc) in git while the real bytes live in a content-addressable cache that is

pushed to remote storage (S3, GCS, Azure Blob, SSH, or a shared filesystem).

The mental model:

ConcernLives in gitLives in the DVC remote
Code, dvc.yamlyesno
params.yaml, metricsyesno
*.dvc pointer filesyesno
Datasets, modelsno (just the pointer)yes (keyed by content hash)

The data_versioning package complements DVC in two places:

  • Before you adopt DVC, or in environments where a full remote is overkill

(CI scratch, notebooks, ad-hoc audits), hashing.py + snapshots.py give you

the same content-addressable "did this dataset change?" check with zero

infrastructure.

  • Alongside DVC, lineage.py, pipeline.py, and reproducibility.py let

you generate dvc.yaml programmatically, render the pipeline as a graph, and

capture the run environment that DVC does not record on its own.


2. Install

DVC is distributed on PyPI. Install the base package plus the extra for your

remote backend:

bash
pip install "dvc[s3]"      # Amazon S3 / S3-compatible (MinIO, Cloudflare R2)
pip install "dvc[gs]"      # Google Cloud Storage
pip install "dvc[azure]"   # Azure Blob Storage
pip install "dvc[ssh]"     # any host you can reach over SSH

Verify the install and pin the major version in your project so teammates and CI

agree:

bash
dvc --version          # expect 3.x
echo "dvc>=3,<4" >> requirements.txt

3. Initialize a repository

Run dvc init inside an existing git repository — DVC layers on top of git:

bash
git init                 # if you have not already
dvc init
git add .dvc .dvcignore
git commit -m "chore: initialize DVC"

dvc init creates a .dvc/ directory (config + local cache) and a .dvcignore

file (same idea as .gitignore, but for DVC's own scans).


4. Configure a remote

The remote is where dataset bytes are pushed. Add one and mark it the default

with -d:

bash
# S3 (or any S3-compatible endpoint)
dvc remote add -d storage s3://your-org-dvc-store/ml
dvc remote modify storage region eu-north-1

# Google Cloud Storage
dvc remote add -d storage gs://your-org-dvc-store/ml

# Azure Blob
dvc remote add -d storage azure://your-org-container/ml

# SSH / on-prem
dvc remote add -d storage ssh://[email protected]/srv/dvc

Commit the remote config (it contains no secrets — credentials come from your

environment / cloud SDK):

bash
git add .dvc/config
git commit -m "chore: configure default DVC remote"

Credentials never go in git. DVC reads AWS/GCP/Azure credentials from the

usual environment variables, ~/.aws/credentials, or instance metadata. Keep

secrets in your secret manager, exactly as the anonymization rules require.


5. Track a dataset

Hand a file or directory to dvc add. DVC moves the bytes into its cache,

writes a small *.dvc pointer, and updates .gitignore so the raw data is not

accidentally committed:

bash
dvc add data/raw
git add data/raw.dvc data/.gitignore
git commit -m "data: track raw dataset with DVC"
dvc push                 # upload the bytes to the remote

The pointer file is what your teammates pull — dvc pull fetches the matching

bytes by hash.

Cross-check with this toolkit

Before and after dvc add, you can sanity-check the dataset's content hash with

the bundled hasher (no DVC call needed):

python
from data_versioning import hash_directory

print(hash_directory("data/raw").short_hash)   # stable, deterministic id

If that short hash matches what a teammate sees, you are looking at byte-identical

data — independent confirmation that dvc pull gave you the right version.


6. Define a pipeline (dvc.yaml)

A pipeline turns ad-hoc scripts into a reproducible DAG. See dvc/dvc.yaml in

this product for a complete four-stage example. The shape of one stage:

yaml
stages:
  prepare:
    cmd: python src/prepare.py
    deps:
      - src/prepare.py
      - data/raw
    params:
      - prepare.seed
      - prepare.test_size
    outs:
      - data/prepared

You can also generate dvc.yaml programmatically with pipeline.py, which

keeps stage definitions in code and avoids hand-editing YAML:

python
from data_versioning import Pipeline, Output

pipe = Pipeline()
pipe.add_stage(
    name="train",
    cmd="python src/train.py",
    deps=["src/train.py", "data/features"],
    params=["train.n_estimators", "train.max_depth"],
    outs=[Output("models/model.pkl", cache=True)],
    metrics=[Output("metrics.json", cache=False)],
)
pipe.write("dvc.yaml")          # emits valid YAML, no PyYAML dependency

7. Reproduce and inspect

bash
dvc repro                 # run only the stages whose inputs changed
dvc dag                   # print the pipeline graph in your terminal
dvc metrics show          # show metrics from metrics.json
dvc params diff           # show which params changed vs the last commit
dvc metrics diff          # show how metrics moved vs the last commit

dvc repro is the heart of reproducibility: because every stage declares its

deps, params, and outs, DVC knows exactly what is stale and reruns the minimum.


8. The everyday workflow

bash
# pull code + data to the exact versions recorded in this commit
git pull
dvc pull

# ... make changes (edit code, bump a param, refresh data) ...

dvc repro                 # rebuild what changed
git add dvc.lock metrics.json params.yaml
git commit -m "exp: tune max_depth=8"
dvc push                  # upload any new/changed artifacts
git push

dvc.lock is the auto-generated record of the exact hashes used in the last

successful run — always commit it. Together, git + dvc.lock + the remote let

anyone recreate your result byte-for-byte.


9. CI integration

A minimal CI job that verifies the pipeline still reproduces:

yaml
# .github/workflows/dvc.yml (illustrative)
steps:
  - uses: actions/checkout@v4
  - run: pip install "dvc[s3]" -r requirements.txt
  - run: dvc pull                       # fetch data by hash
  - run: dvc repro --dry                # fail if anything is stale/undeclared
  - run: dvc metrics diff --targets metrics.json

For pull requests, dvc metrics diff main posts a clean before/after table so

reviewers can see the effect of a change on model quality.


10. Troubleshooting

SymptomLikely cause / fix
dvc push uploads nothingBytes already in the remote (deduplicated by hash) — expected.
dvc pull fails with "missing in remote"Someone forgot to dvc push; or you are pointed at a stale remote.
dvc repro reruns everythingA deps path changed mtime/content, or a tracked dir is non-deterministic.
Output "changed" every runStage writes nondeterministic bytes (timestamps, unsorted order).
Huge git historyA dataset was git add-ed instead of dvc add-ed. Untrack it.

For the "changed every run" class of problems, hash the output directory with

data_versioning.hash_directory twice in a row: if the tree_hash differs, the

nondeterminism is in *your* code, not DVC. See reproducibility-guide.md for the

usual culprits and fixes.


Next steps

  • Read reproducibility-guide.md for seeds, environment capture, and the

reproducibility checklist.

  • Run python examples/version_dataset.py to watch snapshots + diffs in action.
  • Run python examples/reproduce_experiment.py to see a generated dvc.yaml,

a lineage graph, and an environment-drift report.

Questions: [email protected].

Chapter 2
🔒 Available in full product

Reproducibility Guide

Chapter 3
🔒 Available in full product

DVC Setup Guide

Chapter 4
🔒 Available in full product

Reproducibility Guide

You’ve reached the end of the free preview

Get the full ML Data Versioning and unlock everything.

All Chapters

Get the complete guide with every chapter unlocked, including code samples, diagrams, and best practices.

Full Tool Suite

Access all interactive tools with complete data, all workload profiles, and the full scenario library.

Source Files

Downloadable source code, configuration files, and working examples from every chapter.

Lifetime Updates

Free updates for life. Every new chapter, tool, and improvement included.

Buy Now — $29 →
📦 Free sample included — download another copy or visit the store for the full product.
ML Data Versioning v1.0.0 — Free Preview