Stream uploads directly into the CDC chunk store (no spool, single write)
Every upload surface previously wrote each byte to disk twice: the HTTP body was spooled to a temp file (or assembled from chunk parts), then mmap-re-read for FastCDC analysis, and finally the new chunks were written to the blob backend. CDC could not start until the last byte arrived, so large uploads paid receive + reread + rewrite latency. The dedup engine now chunks, hashes and settles the stream WHILE it arrives (fastcdc AsyncStreamCDC + incremental BLAKE3): - Each batch of distinct chunks is pinned-or-classified by ONE `UPDATE … RETURNING` (no check-then-bump TOCTOU; pinned chunks can't be reclaimed mid-upload), and only chunks the store doesn't have are written — a full dedup hit performs zero content writes. - Durability before visibility is preserved: one batched fsync sweep, then one batched INSERT, then the manifest. Identical concurrent uploads are resolved at the manifest INSERT via ON CONFLICT (the loser releases its references and becomes a dedup hit). - A drop guard rolls back pins and surfaces written-but-unregistered chunks to GC if the request future is cancelled mid-stream. - MIME sniffing now peeks the first bytes in-flight; client-requested MD5/SHA-256 checksums are computed by a stream tee — the post-upload re-read of the assembled file is gone. All surfaces converge on the new interfaces::upload_ingest helper: REST multipart, WebDAV PUT, NextCloud PUT, WOPI PutFile, the dedup endpoint, and both chunked-upload completions (which now stream their ordered parts straight into the store instead of writing an assembled file — chunk parts persist until finalize, so completion is genuinely retryable). The legacy blob re-chunk migration streams from the backend with no spool file either. Legacy removed: store_from_file + mmap CDC analysers + temp-path plumbing through every port (pre_computed_hash, save_file_from_temp, update_file_content_from_temp), upload_spool + assembled-file assembly in both chunked services, create_file/update_file byte-slice variants (no callers), common::temp, the OXICLOUD_UPLOAD_TMPDIR config, and the memmap2 dependency. Verified end-to-end against PostgreSQL 16: 8 MB upload (26 chunks), identical re-upload (dedup hit, zero writes), 3-byte edit re-upload (26 chunks, 1 written), byte-identical downloads, Range across chunk boundaries, concurrent identical-upload race (manifest ref 2), and trash-empty reclaiming exactly the unshared chunk while the shared 25 survive for the edited file. The empty/sub-8KB multipart path found a post-EOF re-poll panic in the MIME peek (fixed with fuse + regression test). https://claude.ai/code/session_01WdNenpnujNR2sc32XVvwfS
This commit is contained in:
@@ -98,6 +98,21 @@ pub struct UploadStatusResponseDto {
|
||||
pub is_complete: bool,
|
||||
}
|
||||
|
||||
/// A completed upload session, ready to be streamed into the blob store.
|
||||
///
|
||||
/// `chunk_paths` lists every chunk file in assembly order — the caller
|
||||
/// concatenates them as one byte stream (CDC chunking + hashing happen in
|
||||
/// that single pass). The files stay on disk until `finalize_upload`, so a
|
||||
/// failed completion (e.g. checksum mismatch) remains retryable.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct CompletedUploadParts {
|
||||
pub chunk_paths: Vec<PathBuf>,
|
||||
pub filename: String,
|
||||
pub folder_id: Option<String>,
|
||||
pub content_type: String,
|
||||
pub total_size: u64,
|
||||
}
|
||||
|
||||
/// Port for chunked/resumable upload operations.
|
||||
///
|
||||
/// Implementations manage upload sessions, chunk storage, reassembly,
|
||||
@@ -137,16 +152,15 @@ pub trait ChunkedUploadPort: Send + Sync + 'static {
|
||||
user_id: Uuid,
|
||||
) -> Result<UploadStatusResponseDto, DomainError>;
|
||||
|
||||
/// Assemble all chunks into the final file.
|
||||
/// Validate completion and hand back the ordered chunk parts.
|
||||
///
|
||||
/// Returns `(assembled_file_path, filename, folder_id, content_type, total_size, blake3_hash)`.
|
||||
/// The hash is computed during assembly (hash-on-write), eliminating a
|
||||
/// second sequential read of the assembled file.
|
||||
/// No assembled file is written — the caller streams the parts straight
|
||||
/// into the content-addressable store.
|
||||
async fn complete_upload(
|
||||
&self,
|
||||
upload_id: &str,
|
||||
user_id: Uuid,
|
||||
) -> Result<(PathBuf, String, Option<String>, String, u64, String), DomainError>;
|
||||
) -> Result<CompletedUploadParts, DomainError>;
|
||||
|
||||
/// Finalize upload: clean up the session and temporary files.
|
||||
async fn finalize_upload(&self, upload_id: &str, user_id: Uuid) -> Result<(), DomainError>;
|
||||
|
||||
Reference in New Issue
Block a user