feat(storage): a backend that never answers is now a transient failure
Ed pulled the network mid-migration and got nothing: no log, no pause,
after more than two minutes. The cause is not the classification work
that preceded this — it is that there was no error to classify.
Pull a network on an ESTABLISHED TCP connection and there is no RST and
no ICMP. The peer simply stops answering and the socket read blocks
until the OS abandons retransmission, on the order of fifteen minutes.
For that whole window the job is neither running nor failed. Nothing
retries, because nothing failed. It looks exactly like a slow migration.
A refused connection is instant and does surface, which is what made
the earlier `127.0.0.1` test look reassuring. It exercised the one
network failure that cannot hang.
## Two layers, because one does not fit
`TimeoutBlobBackend` is innermost, below retry — a hang has to become an
error before any layer above can react to it. Bounds are per operation
class, because one number cannot fit both a HEAD and a 5 GB upload:
metadata 30s exists / size / delete / init / health / list
open 60s time to FIRST BYTE, not transfer duration
write off the whole transfer is inside the future, so any
bound here is also a maximum upload duration
Write is unbounded by default deliberately: guessing it wrong truncates
legitimate uploads, which is worse than the hang it would prevent. All
three are configurable (`OXICLOUD_STORAGE_TIMEOUT_*_MS`, 0 = unbounded).
The S3 client also gets what it could always have had. It was built from
a bare `config::Builder::new()`, which carries NO `TimeoutConfig` at
all — so `SdkError::TimeoutError`, an arm `s3_domain_error` already
handles, was unreachable. It now sets connect/read timeouts plus
stalled-stream protection, which measures throughput rather than
elapsed time and is therefore the correct instrument for a stream: it
bounds a stalled upload without capping how long a large one may take.
## Local is not the justification
Ed's correction, and it is right: a local path is reached through the
kernel, and the kernel owns that timeout. iSCSI gives up after
`replacement_timeout` (120s default) and returns an I/O error; NVMe-oF
and soft-mounted NFS behave the same. Those arrive as `io::Error` and
`local_io_error` already classifies them. Local passes through the
decorator only because a uniform chain beats a conditional one, and a
bound that never fires costs nothing.
The real asymmetry is Azure: its 0.21 client has no timeout knob short
of a custom transport, and the SDK migration is deferred. That is why
this lives in the chain rather than being configured per SDK.
Also fixes the log gap: the timeout warns with the wrapper, backend,
operation and bound, so a stalled layer is visible before the pause
rather than only afterwards.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -7,6 +7,7 @@ use aws_sdk_s3::primitives::ByteStream;
|
||||
use bytes::Bytes;
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::pin::Pin;
|
||||
use std::time::Duration;
|
||||
use tokio::fs;
|
||||
use tokio_util::io::ReaderStream;
|
||||
|
||||
@@ -39,9 +40,36 @@ impl S3BlobBackend {
|
||||
"oxicloud",
|
||||
);
|
||||
|
||||
// `Builder::new()` starts from nothing — in particular with no
|
||||
// `TimeoutConfig` at all, which meant a lost network on an
|
||||
// established connection produced no error until the OS gave up
|
||||
// on TCP retransmission (~15 minutes). For that whole window a
|
||||
// migration looked merely slow: no error, so no retry, no log
|
||||
// and no pause. It also made the `SdkError::TimeoutError` arm of
|
||||
// `s3_domain_error` unreachable.
|
||||
//
|
||||
// These bounds are deliberately not the ones in `TimeoutPolicy`:
|
||||
// that decorator provides the configurable outer bound for every
|
||||
// backend, while these are the SDK's finer, per-attempt
|
||||
// instruments underneath it.
|
||||
let timeouts = aws_sdk_s3::config::timeout::TimeoutConfig::builder()
|
||||
.connect_timeout(Duration::from_secs(10))
|
||||
// Time to first byte, not transfer duration — a large object
|
||||
// is never punished for being large.
|
||||
.read_timeout(Duration::from_secs(30))
|
||||
.build();
|
||||
|
||||
let mut builder = aws_sdk_s3::config::Builder::new()
|
||||
.region(aws_sdk_s3::config::Region::new(config.region.clone()))
|
||||
.credentials_provider(credentials)
|
||||
.timeout_config(timeouts)
|
||||
// The right tool for a network pulled mid-transfer: it
|
||||
// measures throughput rather than elapsed time, so it can
|
||||
// bound a streaming upload without capping how long a
|
||||
// legitimately large one may take.
|
||||
.stalled_stream_protection(
|
||||
aws_sdk_s3::config::StalledStreamProtectionConfig::enabled().build(),
|
||||
)
|
||||
.behavior_version_latest();
|
||||
|
||||
if let Some(ref endpoint) = config.endpoint_url {
|
||||
|
||||
Reference in New Issue
Block a user