CI guardrails and policy linting
Dumpling ships with a built-in policy linter (lint-policy) and a reference
GitHub Actions workflow so you can catch anonymization policy regressions in
pull requests before they reach production.
dumpling lint-policy
dumpling lint-policy # auto-discover config
dumpling lint-policy --config .dumplingconf # explicit config path
dumpling lint-policy --allow-noop # treat missing config as empty (no violations)
The command loads your configuration, runs a set of policy checks, prints any violations to stderr, and exits:
| Exit code | Meaning |
|---|---|
0 | No violations found |
1 | One or more violations found |
Checks performed
| Code | Severity | Description |
|---|---|---|
empty-rules-table | warning | A [rules] entry has no column rules. Likely a stale or incomplete config section. |
empty-column-cases-table | warning | A [column_cases] entry has no column cases. |
unsalted-hash | warning | A hash strategy is used with no salt (neither per-column salt nor global salt). Unsalted hashes are reversible via precomputed lookup tables for low-entropy inputs (names, emails, common IDs). |
inconsistent-domain-strategy | error | The same domain name is used with two or more different strategies. This breaks referential integrity: a domain shared between incompatible generators (for example faker with different faker targets, or faker vs hash) cannot maintain a single stable mapping. |
uncovered-sensitive-column | error | A column listed in [sensitive_columns] has no matching anonymization rule or case. The column will pass through unmodified, making the sensitive declaration misleading. |
invalid-regex-predicate | error | A regex / iregex / not_regex / not_iregex pattern in row_filters or column_cases is missing, uses unsupported Rust regex features (look-around, backreferences, possessive quantifiers), or fails to compile. Prefer not_like / not_ilike / not_regex / not_iregex, or a default scrub plus keep, instead of negative lookahead. |
Recommended CI setup
Minimal (policy lint only)
# .github/workflows/policy-lint.yml
name: Policy Lint
on:
pull_request:
push:
branches: [main]
jobs:
policy-lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
- run: cargo build --release --locked
- run: ./target/release/dumpling lint-policy
This single job will block merges whenever a policy violation is introduced.
Production-ready: lint + strict coverage + PII scan
Combine lint-policy with Dumpling's other CI gates for defence in depth:
name: Anonymization CI
on:
pull_request:
push:
branches: [main]
jobs:
policy-lint:
name: Policy lint
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
- run: cargo build --release --locked
- name: Lint anonymization policy
run: ./target/release/dumpling lint-policy
anonymize-and-scan:
name: Anonymize + residual PII scan
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
- run: cargo build --release --locked
- name: Anonymize with strict coverage enforcement
run: |
./target/release/dumpling \
--strict-coverage \
--scan-output \
--fail-on-findings \
--report report.json \
-i dump.sql \
-o sanitized.sql
- name: Upload anonymization report
if: always()
uses: actions/upload-artifact@v4
with:
name: anonymization-report
path: report.json
Gating on report diff against a baseline
To detect increases in risk findings across PRs, store a baseline report as a CI artifact on your main branch and compare against it in PRs:
- name: Download baseline report
uses: dawidd6/action-download-artifact@v6
with:
workflow: ci.yml
branch: main
name: anonymization-report
path: baseline/
continue-on-error: true # first run has no baseline yet
- name: Fail if findings increased
run: |
BASELINE=$(jq '.output_scan.total_findings // 0' baseline/report.json 2>/dev/null || echo 0)
CURRENT=$(jq '.output_scan.total_findings' report.json)
echo "Baseline findings: $BASELINE Current findings: $CURRENT"
if [ "$CURRENT" -gt "$BASELINE" ]; then
echo "ERROR: residual PII findings increased from $BASELINE to $CURRENT"
exit 1
fi
Audit evidence
Keep report.json next to the sanitized dump (or keep the report alone when using --no-seal). Together they answer what policy and Dumpling version transformed which input, without reconstructing the run from CI logs.
Default (seal on the dump). The first line of the SQL output and seal_sha256 in the sidecar share the same digest:
SEAL=$(sed -n '1s/.*sha256=//p' sanitized.sql | tr -d '[:space:]')
REPORT=$(jq -r '.seal_sha256' report.json)
test -n "$SEAL" && test "$SEAL" = "$REPORT"
--no-seal (streaming / no comment on the SQL). There is no seal line to grep. Confirm the sidecar still has a 64-character digest and that flags.no_seal is true:
jq -e '.flags.no_seal == true and (.seal_sha256 | length == 64)' report.json
Useful fields for a compliance review:
| Field | Stable across re-runs? | Meaning |
|---|---|---|
dumpling_version | yes (same binary) | Crate semver that produced the dump |
seal_sha256 | yes (same policy + transform options) | Same value as dump-seal sha256= (recorded even with --no-seal) |
config_source / config_sha256 | yes (unchanged file) | Path and SHA-256 of the loaded config bytes |
input_sha256 / output_sha256 | yes (same streams) | SHA-256 of the SQL Dumpling read/wrote (output_sha256 omitted for --check) |
flags / outcomes | yes (same CLI / result) | Gate flags (strict_coverage, fail_on_findings, no_seal, …) and pass/fail |
run_id / started_at | no | Per-invocation identity (RFC 3339 UTC) |
input_sha256 is over the SQL byte stream Dumpling actually processed (decoded pg_restore output, decompressed gzip, and so on), not necessarily the raw archive file on disk.
Example production invocation (file output with dump seal):
dumpling \
--strict-coverage \
--scan-output \
--fail-on-findings \
--report report.json \
-i dump.sql \
-o sanitized.sql
Streaming without a dump-seal comment:
cat dump.sql | dumpling --no-seal --report report.json > sanitized.sql
Archive report.json with the sanitized SQL. To confirm a later re-run used the same policy, compare seal_sha256 (and config_sha256); do not expect run_id or started_at to match. Full field notes: JSON report (audit sidecar).
Tips
- Run
dumpling lint-policylocally before opening a PR to catch violations early:cargo run -- lint-policy. - Treat
error-severity violations as mandatory fixes;warning-severity violations are advisory but should be reviewed. - If you intentionally use
hashwithout a salt (e.g. for non-sensitive low-cardinality fields), add asalt = "${ENV_VAR}"at the global level to suppress theunsalted-hashwarning globally. - In hardened security profile environments, a global
saltis required anyway (--security-profile hardenedwill error without it), sounsalted-hashwarnings become informational. - Prefer
keepundercolumn_cases(ornot_ilike/not_iregex) for allowlist-style exceptions; do not rely on regex look-around — it is rejected at config load and byinvalid-regex-predicate.