perf(upload): one durability barrier per file instead of two fsyncs per chunk

The CDC chunk path paid `sync_all` + parent-dir fsync + one ref-count
upsert round-trip PER ~256 KB chunk: a new 1 GB file ≈ 8 000 fsyncs +
4 000 sequential INSERTs on the upload critical path (seconds on SSD,
tens of seconds on HDD).

- New `put_blob_from_bytes_unsynced` + `sync_blobs` on
  BlobStorageBackend. Defaults delegate to the durable variants, so
  remote backends (S3/Azure — durable on PUT) are untouched.
  LocalBlobBackend writes chunks without fsync and `sync_blobs` then
  fsyncs all files concurrently (kernel coalesces the writeback) plus
  one fsync per DISTINCT shard directory — at most 256 — instead of one
  per chunk. EncryptedBlobBackend delegates both so encryption over the
  local backend keeps the optimization.
- store_chunks: uploads use the unsynced variant; one `sync_blobs`
  barrier runs before anything references the chunks, and the per-chunk
  ref-count upserts collapse into a single
  `INSERT ... SELECT unnest(...) ON CONFLICT` statement.

Durability contract is unchanged: every chunk is on stable storage
before the manifest row that references it commits. A crash mid-upload
leaves chunk files without DB rows — the same orphan class the
per-chunk scheme already produced, just a wider window.

https://claude.ai/code/session_01Dp3oWon5GBMVn4j3QXZdgx
This commit is contained in:
Claude
2026-06-10 09:52:47 +00:00
parent 7687766bf7
commit 908b8f4d4b
4 changed files with 204 additions and 53 deletions
@@ -60,6 +60,36 @@ pub trait BlobStorageBackend: Send + Sync + 'static {
/// without overwriting. Returns the number of bytes stored.
fn put_blob_from_bytes(&self, hash: &str, data: Bytes) -> BoxFut<'_, Result<u64, DomainError>>;
/// Store a blob from in-memory bytes **without a durability barrier**.
///
/// The CDC chunk path writes thousands of small chunks per file; paying
/// two fsyncs per chunk (file + parent dir) put ~8 000 fsyncs on the
/// critical path of a 1 GB upload. Callers using this MUST issue one
/// [`Self::sync_blobs`] barrier over the written hashes before
/// persisting any record that references them (the chunk manifest).
///
/// Default: delegates to [`Self::put_blob_from_bytes`] — correct for
/// remote backends (S3, Azure) where a successful PUT is already
/// durable and the "unsynced" notion does not exist.
fn put_blob_from_bytes_unsynced(
&self,
hash: &str,
data: Bytes,
) -> BoxFut<'_, Result<u64, DomainError>> {
self.put_blob_from_bytes(hash, data)
}
/// Durability barrier for blobs previously written with
/// [`Self::put_blob_from_bytes_unsynced`]: when this resolves, the
/// listed blobs and their directory entries survive a power loss.
///
/// Default: no-op — matches the default `put_blob_from_bytes_unsynced`,
/// which is already durable on completion.
fn sync_blobs(&self, hashes: &[String]) -> BoxFut<'_, Result<(), DomainError>> {
let _ = hashes;
Box::pin(async { Ok(()) })
}
/// Stream the full blob content in chunks.
fn get_blob_stream(&self, hash: &str) -> BoxFut<'_, Result<BlobStream, DomainError>>;