Invalid data causes crashes, corrupted databases, and security vulnerabilities.
A structured validation layer catches problems at the boundary — before bad data
propagates through your system.
External Input
│
▼
┌─────────────┐
│ Type Check │ Are fields the right Python types?
├─────────────┤
│ Schema Check │ Do values satisfy constraints (length, range, pattern)?
├─────────────┤
│ Business │ Do cross-field rules hold (dates, dependencies)?
│ Rules │
├─────────────┤
│ File Check │ Are uploaded files valid (size, format, extension)?
└─────────────┘
│
▼
Application
Combine validators for defence in depth:
from src.pipeline import ValidationPipeline
from src.validators.type_validator import TypeValidator
from src.validators.schema_validator import SchemaValidator
from src.validators.business_rules import BusinessRuleValidator
pipeline = ValidationPipeline(mode="collect_all")
pipeline.add(TypeValidator({"name": str, "age": int}))
pipeline.add(SchemaValidator.from_yaml("configs/schemas/user_schema.yaml"))
pipeline.add(BusinessRuleValidator.from_yaml("configs/rules/business_rules.yaml"))
errors = pipeline.run(incoming_data)
if errors:
# Return 422 with error details
...| Mode | Behaviour |
|---|---|
collect_all | Run every validator, return all errors |
short_circuit | Stop at the first failing validator |
Use collect_all for API responses (show all problems at once).
Use short_circuit for internal pipelines (fail fast, save work).
Inherit from Validator and implement validate:
from src.validators.base import ValidationError, Validator
class NotEmptyValidator(Validator):
def validate(self, data: dict) -> list[ValidationError]:
errors = []
for key, value in data.items():
if isinstance(value, str) and not value.strip():
errors.append(
ValidationError(field=key, message="Must not be empty", code="empty")
)
return errorsfields:
field_name:
type: str | int | float | bool | list | dict
required: true | false
min: 0 # numeric minimum
max: 100 # numeric maximum
min_length: 1 # string/list minimum length
max_length: 255 # string/list maximum length
pattern: "^regex$" # regex pattern for strings
choices: # allowed values
- option_a
- option_bThree built-in reporters:
| Reporter | Best For |
|---|---|
TextReporter | CLI output, logs |
JsonReporter | API responses |
SummaryReporter | User-facing reports |
1. Validate at the boundary — API endpoints, file uploads, CLI inputs.
2. Use schemas for structure — YAML schemas keep validation rules out of code.
3. Collect all errors — Users prefer fixing everything in one pass.
4. Include field paths — Dot-notation paths help locate the problem.
5. Separate concerns — Type checks, schema, and business rules are distinct layers.
6. Test your rules — Validation logic is business logic; it deserves tests.
Get the full Data Validation Toolkit and unlock everything.
Get the complete guide with every chapter unlocked, including code samples, diagrams, and best practices.
Access all interactive tools with complete data, all workload profiles, and the full scenario library.
Downloadable source code, configuration files, and working examples from every chapter.
Free updates for life. Every new chapter, tool, and improvement included.