Re-chunk pre-CDC legacy blobs into CDC manifests at startup

Files uploaded before chunk_manifests landed (20260414000000) are stored
as ONE whole-file blob with no manifest. Every legacy fallback in
DedupService exists to serve them, and the cost concentrates on Range
reads: with encryption enabled, seeking inside a legacy video decrypts
the ENTIRE blob (AES-GCM is all-or-nothing), where a CDC file decrypts
only the overlapping <=1 MiB chunks.

This adds a one-time, idempotent background migration (spawned from the
composition root after dedup init, maintenance pool) that converts each
legacy blob into a regular CDC file, indistinguishable from a native
upload:

  1. Spool the blob through the normal read path (decrypts when
     encryption is on) to a per-attempt-unique temp file, verifying
     BLAKE3 == hash; sizes come from the verified spool, never from the
     legacy storage.blobs.size column (the manifest's total_size drives
     Range arithmetic).
  2. CDC-chunk + store chunks via the existing store_chunks (one
     manifest reference per distinct chunk).
  3. One short accounting TX with the blob row locked: manifest INSERT
     with ref_count = N current file references, blob ref_count -= N,
     row deleted only at exactly 0 - so single-chunk files (chunk hash
     == file hash) keep the physical blob, which IS the chunk; only
     bookkeeping moves, no bytes are rewritten.
  4. Physical whole-file blob deleted only when its row dropped.

Races lean on the row lock: a concurrent identical upload landing a
legacy reference after commit keeps the blob row alive and that file
readable via the fallback (bounded space leak, never data loss); a
crash between chunk store and the TX over-counts one file's chunk refs
(also a bounded leak). Corrupt blobs (content != hash) are logged,
counted, excluded from the sweep and left untouched, with a hard cap
before aborting.

Per-hash failures never block the sweep; manifests are the resumability
marker, so a restart continues where it left off. The legacy read/write
fallbacks stay in place as the safety net while a deployment converges;
they can be deleted once fleets report "legacy re-chunk: nothing to do".

Opt-out via OXICLOUD_LEGACY_RECHUNK=false (documented in example.env)
for metered remote backends where the one-time re-read should be
scheduled deliberately.

Covered by five integration tests against real PostgreSQL (multi-chunk
accounting + Range across a chunk boundary, single-chunk physical-blob
preservation, corrupt-blob isolation, empty blob, and the full
encrypted-backend roundtrip); they run concurrently, which also
exercises the cross-sweep race handling.

https://claude.ai/code/session_0193Hff42gaA962wThxMGSd1
This commit is contained in:
Claude
2026-06-11 10:43:45 +00:00
parent ecbdaee19e
commit 54c494419c
4 changed files with 783 additions and 0 deletions
+9
View File
@@ -87,6 +87,15 @@ OXICLOUD_SERVER_HOST=127.0.0.1
# fresher sync detection, higher = fewer background UPDATEs. Minimum: 100.
#OXICLOUD_TREE_ETAG_FLUSH_MS=500
# One-time startup migration that converts pre-CDC whole-file blobs (files
# uploaded before chunked dedup landed) into CDC chunk manifests, in the
# background on the maintenance pool. Fixes the legacy penalty where a Range
# read (video seek) reads — and with encryption, DECRYPTS — the entire blob.
# Idempotent; a no-op once no legacy blobs remain. Disable only on metered
# remote backends (S3/Azure egress) where the one-time re-read of every
# legacy blob should be scheduled deliberately, e.g. off-peak.
#OXICLOUD_LEGACY_RECHUNK=true
# Allow multiple processes to bind to the same port (SO_REUSEPORT).
# DISABLED by default — leaving this off means a second accidental instance
# will fail immediately with "address already in use", which is the safe behaviour.