refactor(storage): make blob enumeration ordered and hash-cursored

Precondition for the merge-join in backend_consistency (step 6 /
option A of docs/plan/derived-blobs.md), landed separately because it
is independently useful and carries the risk.

Two contract changes on BlobStorageBackend::list_blob_hashes:

1. Entries MUST be in ascending hash order. Every shipped backend
   already did this — local sorts within each shard and walks 00..ff,
   and since the shard IS the hash prefix that is globally sorted; S3
   and Azure list lexicographically by key and blobs/<xx>/<hash> sorts
   identically to <hash>. It was accidental, and a future backend
   enumerating in any other order would have silently made the
   merge-join emit bogus blob_missing_from_backend findings at
   data_loss severity.

2. The cursor is the last hash returned, not an opaque backend token.
   This is what lets a caller resume from a checkpoint it already
   holds — the merge-join keeps one cursor for both the DB walk and
   the backend walk instead of a compound one, which in turn means
   blobs_consistency's existing cursor format survives and no paused
   run is stranded.

Local already derived its position from a hash; it now emits the bare
hash instead of "<shard>/<hash>", and still accepts both legacy forms
so a run paused across this deploy resumes. The bare-shard form works
through the same path unchanged, since "3f" sorts before every 64-char
hash beginning "3f".

S3 moves from continuation_token to StartAfter, which supports this
natively. One non-obvious case handled: a page can contain only
non-canonical keys (.tmp spool files, .corrupt sidecars), which are
filtered into `unknowns`, leaving `blobs` empty — a naive
blobs.last() would return no cursor and silently end enumeration while
is_truncated said otherwise, making an audit job under-report. It now
falls back to the last key seen; StartAfter is a string comparison, so
a non-hash resume point is fine. "Cursor is a hash" constrains what
callers may synthesise, not what backends may return.

Azure is unaffected — it does not implement list_blob_hashes (TODO,
inherits the NotSupported default).

Adds the first test for enumeration at all: ordering across shards with
deliberately out-of-order inserts, complete paged traversal, and
resume from a caller-synthesised cursor.

NOT verified against real S3 — no bucket available here. The local path
is covered by the new test; the StartAfter change is reasoned from the
API contract and needs exercising against a real bucket before it is
relied on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Edouard Vanbelle
2026-08-24 21:25:08 +02:00
parent a8223cab65
commit 7f5ee7401f
3 changed files with 145 additions and 15 deletions
+40 -4
View File
@@ -497,8 +497,13 @@ impl BlobStorageBackend for S3BlobBackend {
.list_objects_v2()
.bucket(&self.bucket)
.max_keys(limit.min(1000) as i32);
// Resume after a HASH, not a continuation token (port contract).
// ListObjectsV2 supports this natively via StartAfter, and it is
// what lets a caller resume the backend side of a merge-join from
// a checkpoint it holds — a continuation token would force a
// re-enumeration from the start on every resume.
if let Some(c) = cursor {
req = req.continuation_token(c);
req = req.start_after(Self::object_key(&c));
}
let resp = req.send().await.map_err(|e| {
@@ -512,6 +517,9 @@ impl BlobStorageBackend for S3BlobBackend {
let objects = resp.contents.unwrap_or_default();
let mut blobs: Vec<BackendBlobEntry> = Vec::with_capacity(objects.len());
let mut unknowns: Vec<BackendUnknownEntry> = Vec::new();
// Last key of the page regardless of classification — the resume
// fallback for an all-unknowns page (see next_cursor below).
let mut last_key: Option<String> = None;
for obj in objects {
let Some(key) = obj.key else { continue };
@@ -539,13 +547,41 @@ impl BlobStorageBackend for S3BlobBackend {
});
match is_canonical {
Some(hash) => blobs.push(BackendBlobEntry { hash, mtime }),
None => unknowns.push(BackendUnknownEntry { path: key, mtime }),
Some(hash) => {
last_key = Some(hash.clone());
blobs.push(BackendBlobEntry { hash, mtime })
}
None => {
last_key = Some(key.clone());
unknowns.push(BackendUnknownEntry { path: key, mtime })
}
}
}
// Resume point: the last hash of this page, not the continuation
// token — see the StartAfter note above.
//
// `blobs.last()` alone is NOT sufficient. A page can legitimately
// contain only non-canonical keys (`.tmp` spool files, `.corrupt`
// sidecars), which are filtered into `unknowns`; `blobs` is then
// empty and a naive `blobs.last()` yields None, silently ending
// enumeration while `is_truncated` says otherwise. A consistency
// sweep would under-report rather than fail — the worst shape of
// bug for an audit job.
//
// So fall back to the last KEY seen. StartAfter is a plain string
// comparison, so any key works as a resume point; it need not be
// a hash. The port contract's "cursor is a hash" is what CALLERS
// may synthesise, not a restriction on what backends may return.
let next_cursor = if resp.is_truncated.unwrap_or(false) {
resp.next_continuation_token
match blobs.last() {
Some(entry) => Some(entry.hash.clone()),
None => last_key.map(|k| {
// Strip the `blobs/<xx>/` prefix: object_key() re-adds
// it when this comes back as a cursor.
k.rsplit('/').next().unwrap_or(&k).to_string()
}),
}
} else {
None
};