This guide takes you from "nothing installed" to a working MLflow tracking
server with a Postgres metadata store and an S3-compatible artifact store, then
shows how to point the kit at it. Read it once end-to-end; afterwards the
Model Registry Workflow guide covers promotion.
Every MLflow setup is made of three independent pieces. Getting them straight
saves hours of confusion:
| Piece | What it stores | This kit uses |
|---|---|---|
| Tracking server | The REST API and Web UI | mlflow server on port 5000 |
| Backend store | Runs, params, metrics, registry metadata | PostgreSQL |
| Artifact store | Models, plots, datasets, large files | MinIO (S3-compatible) |
A common beginner setup collapses all three into the local filesystem
(file:./mlruns). That is fine for a laptop spike but it cannot be shared, has
no model registry over HTTP, and corrupts easily under concurrent writes. The
bundled docker/docker-compose.yml gives you the real, shareable topology with
one command.
# From the product root
docker compose -f docker/docker-compose.yml up -d
# Watch the tracking server come up (first boot pip-installs deps)
docker compose -f docker/docker-compose.yml logs -f mlflowWhen the logs show Listening at: http://0.0.0.0:5000, open:
minioadmin / minioadmin)The stack wires everything together for you:
--backend-store-uri.--artifacts-destination (bucket s3://mlflow).--serve-artifacts, so **clients never need MinIOcredentials** -- they upload/download artifacts *through* the tracking server
using the mlflow-artifacts:/ scheme. This is the recommended modern pattern
and the reason your training code only needs one URL.
Tear it down (keeping data in named volumes) with:
docker compose -f docker/docker-compose.yml down
# add -v to also delete the postgres_data / minio_data volumesThe kit centralises configuration in mlflow_starter.config.MLflowConfig. The
idiomatic flow is "read environment, then apply":
from mlflow_starter import MLflowConfig, configure
config = MLflowConfig.from_env() # reads MLFLOW_TRACKING_URI, tags, etc.
config.experiment_name = "churn-modeling"
configure(config) # exports env + sets the MLflow client URIsSet the environment once in your shell (or a .env consumed by your process
manager):
export MLFLOW_TRACKING_URI=http://localhost:5000
export MLFLOW_EXPERIMENT_NAME=churn-modeling
export MLFLOW_DEFAULT_TAGS='{"team": "ml-platform", "project": "churn"}'MLFLOW_DEFAULT_TAGS is a JSON object string; the kit parses it and applies
those tags to every run opened through experiment_run.
No server running? Use a local file store. The example scripts already do this
automatically, but you can force it explicitly:
config = MLflowConfig(tracking_uri="file:./mlruns", experiment_name="quick-test")
configure(config)Everything in tracking.py works against a file store except the *registry*
features, which require a database-backed server (Postgres here). That is the
single biggest reason to run the Docker stack.
MinIO speaks the S3 API, so moving to AWS S3, GCS (via the S3 interop), or any
S3-compatible object store is a configuration change, not a code change. For
real AWS S3 you simply drop the custom endpoint:
config = MLflowConfig(
tracking_uri="https://internal.docs.example.com",
artifact_location="s3://acme-ml-artifacts/mlflow/",
s3_endpoint_url=None, # None => default AWS endpoints
)
config.apply()MLflowConfig.apply() only exports credentials that are actually set, so it
will never overwrite real IAM-role credentials with empty values. On AWS,
prefer an instance/role profile over static keys and leave
aws_access_key_id / aws_secret_access_key unset.
A 20-second smoke test that proves the whole chain (server + DB + artifacts) is
healthy:
import mlflow
from mlflow_starter import MLflowConfig, configure, experiment_run, log_metrics
configure(MLflowConfig.from_env())
with experiment_run("smoke-test") as run:
log_metrics({"answer": 42})
mlflow.log_text("hello from the smoke test", "notes.txt")
print("run id:", run.info.run_id)If this run appears in the UI and notes.txt is downloadable from its
Artifacts tab, your backend store and artifact store are both wired correctly.
If the run shows but the artifact 404s, your artifact store is misconfigured
(see Troubleshooting).
Connection refused to localhost:5000. The server container is still
pip-installing on first boot. Watch docker compose logs -f mlflow and wait for
the Listening at line.
Runs appear but artifacts fail to upload / 404. You are almost certainly
mixing the proxied-artifact and direct-S3 patterns. With this kit's stack the
server runs --serve-artifacts, so clients must use the tracking URL only --
do not also set MLFLOW_S3_ENDPOINT_URL in your *client* environment, or
the client will try to reach MinIO directly at a hostname it cannot resolve.
psycopg2 / database errors. The Postgres container failed its healthcheck
before MLflow started. Run docker compose ps and confirm postgres is
healthy; if it is restarting, you likely have a stale postgres_data volume
from an older Postgres major version -- remove it with down -v.
Registry calls raise RestException: ... not supported. You are pointed at
a file: store. Stage transitions and the model registry require the
database-backed server. Switch MLFLOW_TRACKING_URI to http://localhost:5000.
Stale model in serving. Loading models:/ always resolves
the *current* Production version. If a server keeps an old model, it cached it
at startup -- restart the scoring container after a promotion, or use aliases
plus a redeploy (see the registry workflow guide).
Before this leaves your laptop:
1. Replace minioadmin / mlflow passwords; move them into a secrets manager.
2. Put the tracking server behind TLS and an auth proxy (MLflow has no built-in
auth in the open-source server).
3. Pin every image tag (postgres:16.3, a dated minio release) instead of
latest for reproducible rebuilds.
4. Back up the Postgres database -- it is the source of truth for your entire
experiment and model history.
5. Bake the MLflow server dependencies into a custom image so restarts do not
re-run pip install.
Get the full MLflow Starter Kit and unlock everything.
Get the complete guide with every chapter unlocked, including code samples, diagrams, and best practices.
Access all interactive tools with complete data, all workload profiles, and the full scenario library.
Downloadable source code, configuration files, and working examples from every chapter.
Free updates for life. Every new chapter, tool, and improvement included.