fix(storage): a read failure is not proof the blob is gone

Ed's point, and the most dangerous bug in the batch: NotFound is a
conclusion callers ACT on. Every read path in all three backends
returned it unconditionally.

    // s3, azure, local — all of them
    .map_err(|e| DomainError::new(ErrorKind::NotFound, …))

So a refused connection, a 503, an expired credential, a stale NFS
handle and an unmounted iSCSI target all reported "blob missing". Nine
sites: get / get-range / stat on each backend.

## Why it is disastrous rather than untidy

`backend_migration` probes its source before copying. A transient probe
error used to `continue` — skip the row, record NOTHING, and let the
cursor advance past it at the end of the batch. With `failed` still 0
the run reached `finish_completed` and FLIPPED THE POINTER to a target
missing every blob the outage covered. A migration reporting success
having silently dropped whatever was unreachable at the time.

That path now pauses when the probe error is transient, and records a
finding when it is permanent, so a run can no longer report clean while
having skipped rows.

## Local storage is not exempt

Ed again: a local backend is a PATH, and that path may be an iSCSI or
NVMe-oF LUN, an NFS mount, or a disk with a failing sector. It matters
MORE there than for a remote backend, because `RetryBlobBackend` is only
applied when the active backend is not Local — nothing below retries, so
the classification is the only thing between a flaky mount and a run
concluding the data is gone.

`local_io_error` maps the network-mount family (TimedOut,
HostUnreachable, NetworkDown, ConnectionReset, StaleNetworkFileHandle)
plus Interrupted and ResourceBusy to transient. PermissionDenied,
ReadOnlyFilesystem and StorageFull stay permanent because retrying
changes nothing without an operator, and InvalidData stays permanent
because corruption is a finding worth keeping. A bad sector arrives as
an uncategorised EIO and lands there too, which is right: the useful
outcome is a finding naming the blob, not a run that waits for a disk to
heal.

## Shape of the fix

Only a genuine absence is NotFound — `NoSuchKey` on S3 GET,
`is_not_found` on S3 HEAD, HTTP 404 on Azure, `ErrorKind::NotFound` on
local. Everything else goes through the classifier, so a 403 stays
permanent rather than being retried forever.

Tested at the local layer, which is where the mapping table is dense
enough to get wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Edouard Vanbelle
2026-09-07 22:15:40 +02:00
parent 0cdb2bb0a9
commit 34a2607658
4 changed files with 262 additions and 59 deletions
+52 -15
View File
@@ -286,11 +286,30 @@ impl BlobStorageBackend for S3BlobBackend {
.send()
.await
.map_err(|e| {
DomainError::new(
ErrorKind::NotFound,
"S3",
format!("Failed to get blob {}: {}", hash, e),
)
// Only a real NoSuchKey is NotFound. This used to
// label EVERY read failure that way — a refused
// connection, a 503, an expired credential all
// reported as "blob missing".
//
// That is the most dangerous wrong answer available
// here, because callers ACT on NotFound by concluding
// the bytes are gone. A migration reading its source
// through this would treat an outage as "the source
// does not have this blob" and move on.
//
// Everything else goes through the normal classifier,
// so a 403 stays permanent rather than being retried
// forever.
if let aws_sdk_s3::error::SdkError::ServiceError(svc) = &e
&& svc.err().is_no_such_key()
{
return DomainError::new(
ErrorKind::NotFound,
"S3",
format!("Failed to get blob {hash}: no such key"),
);
}
s3_domain_error("S3", format!("Failed to get blob {hash}"), &e)
})?;
// Convert S3 ByteStream into a Stream<Item = Result<Bytes, io::Error>>
@@ -325,11 +344,21 @@ impl BlobStorageBackend for S3BlobBackend {
.send()
.await
.map_err(|e| {
DomainError::new(
ErrorKind::NotFound,
"S3",
format!("Failed to get blob range {}: {}", hash, e),
)
// Same rule as the full read: only a real NoSuchKey
// is NotFound. Ranged reads feed CDC reassembly and
// deep verification, so mislabelling an outage here
// reads as "this chunk is gone" — a data-loss
// conclusion drawn from a network problem.
if let aws_sdk_s3::error::SdkError::ServiceError(svc) = &e
&& svc.err().is_no_such_key()
{
return DomainError::new(
ErrorKind::NotFound,
"S3",
format!("Failed to get blob range {hash}: no such key"),
);
}
s3_domain_error("S3", format!("Failed to get blob range {hash}"), &e)
})?;
let reader = output.body.into_async_read();
@@ -407,11 +436,19 @@ impl BlobStorageBackend for S3BlobBackend {
.send()
.await
.map_err(|e| {
DomainError::new(
ErrorKind::NotFound,
"S3",
format!("Failed to stat blob {}: {}", hash, e),
)
// `head_object` reports a missing key as NotFound
// rather than NoSuchKey, so match on the typed
// variant the SDK actually returns here.
if let aws_sdk_s3::error::SdkError::ServiceError(svc) = &e
&& svc.err().is_not_found()
{
return DomainError::new(
ErrorKind::NotFound,
"S3",
format!("Failed to stat blob {hash}: not found"),
);
}
s3_domain_error("S3", format!("Failed to stat blob {hash}"), &e)
})?;
Ok(output.content_length().unwrap_or(0) as u64)