refactor(consistency): blobs_consistency reads only the database

`blobs_consistency` probed `blob_exists` once per row and, under
`?deep=true`, read and re-hashed every blob. `backend_consistency`
already reports the same `blob_missing_from_backend` from its
merge-join — so the probe was duplicated work that found strictly less
(a DB walk cannot see backend-only orphans by construction) at N round
-trips instead of one enumeration. Every scheduled sweep paid for it.

All three physical checks move to `backend_consistency`:

* `blob_missing_from_backend` was already there; the duplicate is gone.
* `blob_corrupted` / `blob_unreadable` hook the matched arm of the
  merge-join, which holds exactly the key pairs worth reading. Guarded
  by `in_range` so a pair past the horizon is not read twice, and
  `params.deep` is persisted on a fresh run and read back on resume so
  a paused deep scan does not silently continue shallow.

Deep mode belongs there because it is backend work end to end: the
only DB input is the hash. Keeping it in `blobs_consistency` forced
that tenant to carry a backend for one flag.

What remains is the half that needs no backend: `refcount_mismatch`
and its repair. The constructor drops from five parameters to two —
no backend, no storage_entries, no storage_path_fallback — and
`?storage=<name>` / `?deep=true` are now inert there, which the
job description says outright.

`affected_files` is needed by both tenants, so it moves to a shared
`blob_diagnostics` module rather than being copied.
`PROBED_STORAGE_PARAM` moves to `backend_consistency`: it was defined
in `blobs_consistency` and re-exported, which is backwards once the
DB-only tenant has no entry to scope. The create-grace window goes
with the probe — it existed to avoid flagging a blob whose bytes had
landed before its row, and the refcount comparison reads one
consistent snapshot.

Known cost: `backend_consistency` returns `backend_unenumerable` on
Azure and mid-migration, so on those configs missing bytes now go
unreported where the per-row probe caught them. That argues for the
Azure enumeration impl, not for keeping the probe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Edouard Vanbelle
2026-08-29 01:44:08 +02:00
parent 1ea3826660
commit f1f327a6c4
11 changed files with 361 additions and 474 deletions
+25 -14
View File
@@ -234,27 +234,38 @@ can be safely dropped.
---
## 4. Blob consistency (`blobs_consistency`)
## 4. Blob consistency — two jobs, split by what they read
Read-only recoverable job that walks `storage.blobs` and reports
divergence between the DB registry and the physical backend.
The registry side and the physical side are separate tenants. They
used to be one, with `blobs_consistency` probing the backend once per
row; that probe found strictly less than the merge-join below, at N
round-trips instead of one enumeration, so it was removed.
### Shallow mode (default)
### `blobs_consistency` — database only
Per row:
Walks `storage.blobs` and compares `ref_count` against the reference
count computed from `storage.files.blob_hash` +
`chunk_manifests.chunk_hashes[]`. On mismatch: `refcount_mismatch`
(severity `inconsistent`), repairable under `?repair=true`.
- `blob_exists(hash)` on the active backend → if false, record
`blob_missing_from_backend` (severity `data_loss`)
- Compare `ref_count` against the actual reference count computed
from `SUM` over `storage.files.blob_hash` + `chunk_manifests.chunk_hashes[]`
→ if mismatch, record `refcount_mismatch` (severity `inconsistent`)
It opens no backend and makes no network call. `?storage=<name>` and
`?deep=true` are inert. Cost is one aggregate SQL per row.
Cost: one existence probe + one aggregate SQL per row. Fast on
S3/Azure (single HEAD).
### `backend_consistency` — everything physical
### Deep mode (`?deep=true`)
Merge-joins the backend's enumeration against `storage.blobs`, both
ordered by hash, yielding both deltas in one pass:
Adds a full read of every blob:
- bytes with no registry row → `orphan_blob` (severity `inconsistent`)
- a registry row with no bytes → `blob_missing_from_backend`
(severity `data_loss`)
`?storage=<name>` scopes it to any declared entry rather than the live
backend.
### Deep mode (`?deep=true`, on `backend_consistency`)
For every hash present on both sides, adds a full read:
- Stream the blob through `EncryptedBlobBackend::get_blob_stream`
(strips header, decrypts if needed, applies BLAKE3 rescue for
+6 -2
View File
@@ -100,12 +100,16 @@ This one-shot repair command re-runs the same env-parse the server does at boot,
### Auditing entries other than the active one
`blobs_consistency` and `backend_consistency` (recoverable jobs on the Jobs tab) accept `?storage=<name>` to probe any declared entry — not just the live one. Use this to verify a migration target before cutover, or to audit an old backend after cutover but before decommissioning:
`backend_consistency` (a recoverable job on the Jobs tab) accepts `?storage=<name>` to audit any declared entry — not just the live one. Use this to verify a migration target before cutover, or to audit an old backend after cutover but before decommissioning:
```
POST /api/admin/jobs/blobs_consistency/trigger?storage=<name>
POST /api/admin/jobs/backend_consistency/trigger?storage=<name>
```
Add `?deep=true` to also read every blob back and re-hash it, which catches silent bit-rot. That is a full read of the entry and can take hours.
`blobs_consistency` does *not* accept `?storage=<name>`: it only reads the database, so there is no entry for it to scope.
Unknown names 400 at the HTTP layer.
## Data Storage
+2 -2
View File
@@ -949,8 +949,8 @@ Findings each job reports today, and where the new tables land:
| # | Edge | Direction | Mechanism | Status |
|---|---|---|---|---|
| 1 | backend → `storage.blobs` | orphan bytes | `orphan_blob` (backend_consistency) | ✓ |
| 2 | `storage.blobs` → backend | missing bytes | `blob_missing_from_backend` | ✓ |
| 3 | chunk bytes | corruption | `blob_corrupted`, `blob_unreadable` | ✓ |
| 2 | `storage.blobs` → backend | missing bytes | `blob_missing_from_backend` (backend_consistency) | ✓ |
| 3 | chunk bytes | corruption | `blob_corrupted`, `blob_unreadable` (backend_consistency, `?deep=true`) | ✓ |
| 4 | manifest → chunks | chunk reaped | `chunk_missing` (files_consistency) | ✓ |
| 5 | `files` → Blob | dangling | `missing_blob` (files_consistency) | ✓ |
| 6 | `storage.blobs.ref_count` | recompute | `refcount_mismatch` | ✓ chunk level only |
+11 -1
View File
@@ -179,7 +179,17 @@ Guardrail:
head pair's fingerprint differs from the second entry's — signals "you
added new keys but haven't rotated legacy blobs yet".
* A **legacy-blob counter** is surfaced in the admin panel per storage entry.
The counter is maintained by `blobs_consistency`: during its normal walk it
> **Retarget (post-split).** This plan names `blobs_consistency` as the
> host for the magic-byte branch throughout. That is no longer the right
> tenant: `blobs_consistency` became database-only and opens no backend,
> while `backend_consistency` owns every physical check and already holds
> the enumeration. Read `backend_consistency` wherever the sections below
> say `blobs_consistency`. Nothing else about the design changes — the
> magic-byte check still rides along on an existing walk, and still lands
> in the run's `stats` bag.
The counter is maintained by that scan: during its normal walk it
branches on the magic-byte check and records the legacy count as a run
statistic on `jobs.recoverable_runs` (existing surface, no schema hit). The
admin panel reads the most recent count and displays it. Refresh cadence is
+4 -2
View File
@@ -407,10 +407,12 @@ Per slice, plus these end-to-end scenarios in Hurl:
6. **In-place encryption rotation refused**: two entries, same S3 bucket,
different encryption keys. Trigger migration → refuses with the specific
error message pointing at the encryption case and the two-step workaround.
7. **`?storage=<name>` on blobs_consistency**: run against `s3_prod` before
7. **`?storage=<name>` on backend_consistency**: run against `s3_prod` before
cutover. Full walk, `probed_storage` in run row. Then cutover, then rerun
against `local_main` — verifies old backend still has everything.
8. **Unknown storage name**: `POST /admin/jobs/blobs_consistency/trigger?storage=nope`
(This scenario named `blobs_consistency` until that tenant became
database-only; entry scoping belongs to whichever job opens a backend.)
8. **Unknown storage name**: `POST /admin/jobs/backend_consistency/trigger?storage=nope`
→ 400 with known-names list. No run row created.
9. **Missing entry at boot**: `active_backend_name = "gone"` but `_ENTRIES`
doesn't include it → boot aborts with the specific message pointing at