Stream uploads directly into the CDC chunk store (no spool, single write)

Every upload surface previously wrote each byte to disk twice: the HTTP
body was spooled to a temp file (or assembled from chunk parts), then
mmap-re-read for FastCDC analysis, and finally the new chunks were
written to the blob backend. CDC could not start until the last byte
arrived, so large uploads paid receive + reread + rewrite latency.

The dedup engine now chunks, hashes and settles the stream WHILE it
arrives (fastcdc AsyncStreamCDC + incremental BLAKE3):

- Each batch of distinct chunks is pinned-or-classified by ONE
  `UPDATE … RETURNING` (no check-then-bump TOCTOU; pinned chunks can't
  be reclaimed mid-upload), and only chunks the store doesn't have are
  written — a full dedup hit performs zero content writes.
- Durability before visibility is preserved: one batched fsync sweep,
  then one batched INSERT, then the manifest. Identical concurrent
  uploads are resolved at the manifest INSERT via ON CONFLICT (the
  loser releases its references and becomes a dedup hit).
- A drop guard rolls back pins and surfaces written-but-unregistered
  chunks to GC if the request future is cancelled mid-stream.
- MIME sniffing now peeks the first bytes in-flight; client-requested
  MD5/SHA-256 checksums are computed by a stream tee — the post-upload
  re-read of the assembled file is gone.

All surfaces converge on the new interfaces::upload_ingest helper:
REST multipart, WebDAV PUT, NextCloud PUT, WOPI PutFile, the dedup
endpoint, and both chunked-upload completions (which now stream their
ordered parts straight into the store instead of writing an assembled
file — chunk parts persist until finalize, so completion is genuinely
retryable). The legacy blob re-chunk migration streams from the
backend with no spool file either.

Legacy removed: store_from_file + mmap CDC analysers + temp-path
plumbing through every port (pre_computed_hash, save_file_from_temp,
update_file_content_from_temp), upload_spool + assembled-file
assembly in both chunked services, create_file/update_file byte-slice
variants (no callers), common::temp, the OXICLOUD_UPLOAD_TMPDIR
config, and the memmap2 dependency.

Verified end-to-end against PostgreSQL 16: 8 MB upload (26 chunks),
identical re-upload (dedup hit, zero writes), 3-byte edit re-upload
(26 chunks, 1 written), byte-identical downloads, Range across chunk
boundaries, concurrent identical-upload race (manifest ref 2), and
trash-empty reclaiming exactly the unshared chunk while the shared 25
survive for the edited file. The empty/sub-8KB multipart path found a
post-EOF re-poll panic in the MIME peek (fixed with fuse + regression
test).

https://claude.ai/code/session_01WdNenpnujNR2sc32XVvwfS
This commit is contained in:
Claude
2026-06-11 13:06:33 +00:00
parent 7157454afd
commit e3f04d58aa
34 changed files with 1864 additions and 2440 deletions
+19 -5
View File
@@ -98,6 +98,21 @@ pub struct UploadStatusResponseDto {
pub is_complete: bool,
}
/// A completed upload session, ready to be streamed into the blob store.
///
/// `chunk_paths` lists every chunk file in assembly order — the caller
/// concatenates them as one byte stream (CDC chunking + hashing happen in
/// that single pass). The files stay on disk until `finalize_upload`, so a
/// failed completion (e.g. checksum mismatch) remains retryable.
#[derive(Debug, Clone)]
pub struct CompletedUploadParts {
pub chunk_paths: Vec<PathBuf>,
pub filename: String,
pub folder_id: Option<String>,
pub content_type: String,
pub total_size: u64,
}
/// Port for chunked/resumable upload operations.
///
/// Implementations manage upload sessions, chunk storage, reassembly,
@@ -137,16 +152,15 @@ pub trait ChunkedUploadPort: Send + Sync + 'static {
user_id: Uuid,
) -> Result<UploadStatusResponseDto, DomainError>;
/// Assemble all chunks into the final file.
/// Validate completion and hand back the ordered chunk parts.
///
/// Returns `(assembled_file_path, filename, folder_id, content_type, total_size, blake3_hash)`.
/// The hash is computed during assembly (hash-on-write), eliminating a
/// second sequential read of the assembled file.
/// No assembled file is written — the caller streams the parts straight
/// into the content-addressable store.
async fn complete_upload(
&self,
upload_id: &str,
user_id: Uuid,
) -> Result<(PathBuf, String, Option<String>, String, u64, String), DomainError>;
) -> Result<CompletedUploadParts, DomainError>;
/// Finalize upload: clean up the session and temporary files.
async fn finalize_upload(&self, upload_id: &str, user_id: Uuid) -> Result<(), DomainError>;