perf(upload): one durability barrier per file instead of two fsyncs per chunk
The CDC chunk path paid `sync_all` + parent-dir fsync + one ref-count upsert round-trip PER ~256 KB chunk: a new 1 GB file ≈ 8 000 fsyncs + 4 000 sequential INSERTs on the upload critical path (seconds on SSD, tens of seconds on HDD). - New `put_blob_from_bytes_unsynced` + `sync_blobs` on BlobStorageBackend. Defaults delegate to the durable variants, so remote backends (S3/Azure — durable on PUT) are untouched. LocalBlobBackend writes chunks without fsync and `sync_blobs` then fsyncs all files concurrently (kernel coalesces the writeback) plus one fsync per DISTINCT shard directory — at most 256 — instead of one per chunk. EncryptedBlobBackend delegates both so encryption over the local backend keeps the optimization. - store_chunks: uploads use the unsynced variant; one `sync_blobs` barrier runs before anything references the chunks, and the per-chunk ref-count upserts collapse into a single `INSERT ... SELECT unnest(...) ON CONFLICT` statement. Durability contract is unchanged: every chunk is on stable storage before the manifest row that references it commits. A crash mid-upload leaves chunk files without DB rows — the same orphan class the per-chunk scheme already produced, just a wider window. https://claude.ai/code/session_01Dp3oWon5GBMVn4j3QXZdgx
This commit is contained in:
@@ -60,6 +60,36 @@ pub trait BlobStorageBackend: Send + Sync + 'static {
|
||||
/// without overwriting. Returns the number of bytes stored.
|
||||
fn put_blob_from_bytes(&self, hash: &str, data: Bytes) -> BoxFut<'_, Result<u64, DomainError>>;
|
||||
|
||||
/// Store a blob from in-memory bytes **without a durability barrier**.
|
||||
///
|
||||
/// The CDC chunk path writes thousands of small chunks per file; paying
|
||||
/// two fsyncs per chunk (file + parent dir) put ~8 000 fsyncs on the
|
||||
/// critical path of a 1 GB upload. Callers using this MUST issue one
|
||||
/// [`Self::sync_blobs`] barrier over the written hashes before
|
||||
/// persisting any record that references them (the chunk manifest).
|
||||
///
|
||||
/// Default: delegates to [`Self::put_blob_from_bytes`] — correct for
|
||||
/// remote backends (S3, Azure) where a successful PUT is already
|
||||
/// durable and the "unsynced" notion does not exist.
|
||||
fn put_blob_from_bytes_unsynced(
|
||||
&self,
|
||||
hash: &str,
|
||||
data: Bytes,
|
||||
) -> BoxFut<'_, Result<u64, DomainError>> {
|
||||
self.put_blob_from_bytes(hash, data)
|
||||
}
|
||||
|
||||
/// Durability barrier for blobs previously written with
|
||||
/// [`Self::put_blob_from_bytes_unsynced`]: when this resolves, the
|
||||
/// listed blobs and their directory entries survive a power loss.
|
||||
///
|
||||
/// Default: no-op — matches the default `put_blob_from_bytes_unsynced`,
|
||||
/// which is already durable on completion.
|
||||
fn sync_blobs(&self, hashes: &[String]) -> BoxFut<'_, Result<(), DomainError>> {
|
||||
let _ = hashes;
|
||||
Box::pin(async { Ok(()) })
|
||||
}
|
||||
|
||||
/// Stream the full blob content in chunks.
|
||||
fn get_blob_stream(&self, hash: &str) -> BoxFut<'_, Result<BlobStream, DomainError>>;
|
||||
|
||||
|
||||
Reference in New Issue
Block a user