refactor(storage): make blob enumeration ordered and hash-cursored
Precondition for the merge-join in backend_consistency (step 6 / option A of docs/plan/derived-blobs.md), landed separately because it is independently useful and carries the risk. Two contract changes on BlobStorageBackend::list_blob_hashes: 1. Entries MUST be in ascending hash order. Every shipped backend already did this — local sorts within each shard and walks 00..ff, and since the shard IS the hash prefix that is globally sorted; S3 and Azure list lexicographically by key and blobs/<xx>/<hash> sorts identically to <hash>. It was accidental, and a future backend enumerating in any other order would have silently made the merge-join emit bogus blob_missing_from_backend findings at data_loss severity. 2. The cursor is the last hash returned, not an opaque backend token. This is what lets a caller resume from a checkpoint it already holds — the merge-join keeps one cursor for both the DB walk and the backend walk instead of a compound one, which in turn means blobs_consistency's existing cursor format survives and no paused run is stranded. Local already derived its position from a hash; it now emits the bare hash instead of "<shard>/<hash>", and still accepts both legacy forms so a run paused across this deploy resumes. The bare-shard form works through the same path unchanged, since "3f" sorts before every 64-char hash beginning "3f". S3 moves from continuation_token to StartAfter, which supports this natively. One non-obvious case handled: a page can contain only non-canonical keys (.tmp spool files, .corrupt sidecars), which are filtered into `unknowns`, leaving `blobs` empty — a naive blobs.last() would return no cursor and silently end enumeration while is_truncated said otherwise, making an audit job under-report. It now falls back to the last key seen; StartAfter is a string comparison, so a non-hash resume point is fine. "Cursor is a hash" constrains what callers may synthesise, not what backends may return. Azure is unaffected — it does not implement list_blob_hashes (TODO, inherits the NotSupported default). Adds the first test for enumeration at all: ordering across shards with deliberately out-of-order inserts, complete paged traversal, and resume from a caller-synthesised cursor. NOT verified against real S3 — no bucket available here. The local path is covered by the new test; the StartAfter change is reasoned from the API contract and needs exercising against a real bucket before it is relied on. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -244,11 +244,28 @@ pub trait BlobStorageBackend: Send + Sync + 'static {
|
||||
///
|
||||
/// * `cursor` — opaque continuation token from a prior call, or
|
||||
/// `None` to start from the beginning. Format is per-backend
|
||||
/// (local = last path visited; S3 = continuation token; Azure
|
||||
/// = list marker); callers treat it as opaque.
|
||||
/// — **the last blob hash returned by the previous page**.
|
||||
/// Enumeration resumes strictly AFTER that hash.
|
||||
///
|
||||
/// This is deliberately NOT an opaque backend token. Callers may
|
||||
/// synthesise a cursor from any hash they hold, which is what lets a
|
||||
/// consistency sweep merge-join this stream against a
|
||||
/// `storage.blobs` walk and resume both sides from one checkpoint.
|
||||
/// An opaque token would force the backend side to re-enumerate from
|
||||
/// the beginning on every resume.
|
||||
/// * `limit` — soft cap on batch size; backends may return
|
||||
/// fewer (e.g. end of a shard directory).
|
||||
///
|
||||
/// **Entries MUST be returned in ascending hash order**, and pages must
|
||||
/// be contiguous in that order. Every shipped backend already satisfies
|
||||
/// this — local sorts within each shard and walks shards `00`..`ff`
|
||||
/// (the shard IS the hash prefix, so that is globally sorted); S3 and
|
||||
/// Azure list lexicographically by key, and `blobs/<xx>/<hash>` sorts
|
||||
/// identically to `<hash>`. It is stated here because the merge-join in
|
||||
/// `backend_consistency` depends on it: an unordered backend would
|
||||
/// silently emit bogus `blob_missing_from_backend` findings at
|
||||
/// `data_loss` severity.
|
||||
///
|
||||
/// Returns `(entries, next_cursor)`. `next_cursor = None` means
|
||||
/// enumeration is complete. Each `BackendBlobEntry` carries the
|
||||
/// hash + optional mtime for grace-window filtering.
|
||||
|
||||
Reference in New Issue
Block a user