fix(storage): a read failure is not proof the blob is gone
Ed's point, and the most dangerous bug in the batch: NotFound is a
conclusion callers ACT on. Every read path in all three backends
returned it unconditionally.
// s3, azure, local — all of them
.map_err(|e| DomainError::new(ErrorKind::NotFound, …))
So a refused connection, a 503, an expired credential, a stale NFS
handle and an unmounted iSCSI target all reported "blob missing". Nine
sites: get / get-range / stat on each backend.
## Why it is disastrous rather than untidy
`backend_migration` probes its source before copying. A transient probe
error used to `continue` — skip the row, record NOTHING, and let the
cursor advance past it at the end of the batch. With `failed` still 0
the run reached `finish_completed` and FLIPPED THE POINTER to a target
missing every blob the outage covered. A migration reporting success
having silently dropped whatever was unreachable at the time.
That path now pauses when the probe error is transient, and records a
finding when it is permanent, so a run can no longer report clean while
having skipped rows.
## Local storage is not exempt
Ed again: a local backend is a PATH, and that path may be an iSCSI or
NVMe-oF LUN, an NFS mount, or a disk with a failing sector. It matters
MORE there than for a remote backend, because `RetryBlobBackend` is only
applied when the active backend is not Local — nothing below retries, so
the classification is the only thing between a flaky mount and a run
concluding the data is gone.
`local_io_error` maps the network-mount family (TimedOut,
HostUnreachable, NetworkDown, ConnectionReset, StaleNetworkFileHandle)
plus Interrupted and ResourceBusy to transient. PermissionDenied,
ReadOnlyFilesystem and StorageFull stay permanent because retrying
changes nothing without an operator, and InvalidData stays permanent
because corruption is a finding worth keeping. A bad sector arrives as
an uncategorised EIO and lands there too, which is right: the useful
outcome is a finding naming the blob, not a run that waits for a disk to
heal.
## Shape of the fix
Only a genuine absence is NotFound — `NoSuchKey` on S3 GET,
`is_not_found` on S3 HEAD, HTTP 404 on Azure, `ErrorKind::NotFound` on
local. Everything else goes through the classifier, so a 403 stays
permanent rather than being retried forever.
Tested at the local layer, which is where the mapping table is dense
enough to get wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -239,11 +239,7 @@ impl BlobStorageBackend for AzureBlobBackend {
|
||||
let first = match pages.next().await {
|
||||
Some(Ok(response)) => response,
|
||||
Some(Err(e)) => {
|
||||
return Err(DomainError::new(
|
||||
ErrorKind::NotFound,
|
||||
"Azure",
|
||||
format!("Failed to get blob {hash}: {e}"),
|
||||
));
|
||||
return Err(azure_read_error(format!("Failed to get blob {hash}"), &e));
|
||||
}
|
||||
None => {
|
||||
let empty: BlobStream =
|
||||
@@ -324,10 +320,9 @@ impl BlobStorageBackend for AzureBlobBackend {
|
||||
let first = match pages.next().await {
|
||||
Some(Ok(response)) => response,
|
||||
Some(Err(e)) => {
|
||||
return Err(DomainError::new(
|
||||
ErrorKind::NotFound,
|
||||
"Azure",
|
||||
format!("Failed to get blob range {hash}: {e}"),
|
||||
return Err(azure_read_error(
|
||||
format!("Failed to get blob range {hash}"),
|
||||
&e,
|
||||
));
|
||||
}
|
||||
None => {
|
||||
@@ -415,13 +410,10 @@ impl BlobStorageBackend for AzureBlobBackend {
|
||||
let hash = hash.to_owned();
|
||||
Box::pin(async move {
|
||||
let client = self.blob_client(&hash);
|
||||
let props = client.get_properties().await.map_err(|e| {
|
||||
DomainError::new(
|
||||
ErrorKind::NotFound,
|
||||
"Azure",
|
||||
format!("Failed to stat blob {hash}: {e}"),
|
||||
)
|
||||
})?;
|
||||
let props = client
|
||||
.get_properties()
|
||||
.await
|
||||
.map_err(|e| azure_read_error(format!("Failed to stat blob {hash}"), &e))?;
|
||||
Ok(props.blob.properties.content_length)
|
||||
})
|
||||
}
|
||||
@@ -683,6 +675,33 @@ impl BlobStorageBackend for AzureBlobBackend {
|
||||
/// second wearing the clothes of the first. So the policy is to retry
|
||||
/// as if transient and let the bounded attempt cap turn the difference
|
||||
/// into a Paused run an operator can act on.
|
||||
/// Read-path variant of [`azure_domain_error`]: only a real 404 is
|
||||
/// `NotFound`.
|
||||
///
|
||||
/// Every Azure read used to label EVERY failure `NotFound` — a refused
|
||||
/// connection, a 503, an expired SAS token all reported as "blob
|
||||
/// missing". That is the most dangerous wrong answer available on a read
|
||||
/// path, because callers ACT on NotFound by concluding the bytes are
|
||||
/// gone: a migration reading its source would treat an outage as "the
|
||||
/// source does not have this blob" and move past it.
|
||||
///
|
||||
/// Everything that is not a 404 goes through the normal classifier, so a
|
||||
/// 403 stays permanent instead of being retried.
|
||||
pub(crate) fn azure_read_error(context: String, err: &azure_core::Error) -> DomainError {
|
||||
use azure_core::error::ErrorKind as AzKind;
|
||||
|
||||
if let AzKind::HttpResponse { status, .. } = err.kind()
|
||||
&& u16::from(*status) == 404
|
||||
{
|
||||
return DomainError::new(
|
||||
ErrorKind::NotFound,
|
||||
"Azure",
|
||||
format!("{context}: not found"),
|
||||
);
|
||||
}
|
||||
azure_domain_error(context, err)
|
||||
}
|
||||
|
||||
pub(crate) fn azure_domain_error(context: String, err: &azure_core::Error) -> DomainError {
|
||||
use azure_core::error::ErrorKind as AzKind;
|
||||
|
||||
|
||||
Reference in New Issue
Block a user