Stream uploads directly into the CDC chunk store (no spool, single write)
Every upload surface previously wrote each byte to disk twice: the HTTP body was spooled to a temp file (or assembled from chunk parts), then mmap-re-read for FastCDC analysis, and finally the new chunks were written to the blob backend. CDC could not start until the last byte arrived, so large uploads paid receive + reread + rewrite latency. The dedup engine now chunks, hashes and settles the stream WHILE it arrives (fastcdc AsyncStreamCDC + incremental BLAKE3): - Each batch of distinct chunks is pinned-or-classified by ONE `UPDATE … RETURNING` (no check-then-bump TOCTOU; pinned chunks can't be reclaimed mid-upload), and only chunks the store doesn't have are written — a full dedup hit performs zero content writes. - Durability before visibility is preserved: one batched fsync sweep, then one batched INSERT, then the manifest. Identical concurrent uploads are resolved at the manifest INSERT via ON CONFLICT (the loser releases its references and becomes a dedup hit). - A drop guard rolls back pins and surfaces written-but-unregistered chunks to GC if the request future is cancelled mid-stream. - MIME sniffing now peeks the first bytes in-flight; client-requested MD5/SHA-256 checksums are computed by a stream tee — the post-upload re-read of the assembled file is gone. All surfaces converge on the new interfaces::upload_ingest helper: REST multipart, WebDAV PUT, NextCloud PUT, WOPI PutFile, the dedup endpoint, and both chunked-upload completions (which now stream their ordered parts straight into the store instead of writing an assembled file — chunk parts persist until finalize, so completion is genuinely retryable). The legacy blob re-chunk migration streams from the backend with no spool file either. Legacy removed: store_from_file + mmap CDC analysers + temp-path plumbing through every port (pre_computed_hash, save_file_from_temp, update_file_content_from_temp), upload_spool + assembled-file assembly in both chunked services, create_file/update_file byte-slice variants (no callers), common::temp, the OXICLOUD_UPLOAD_TMPDIR config, and the memmap2 dependency. Verified end-to-end against PostgreSQL 16: 8 MB upload (26 chunks), identical re-upload (dedup hit, zero writes), 3-byte edit re-upload (26 chunks, 1 written), byte-identical downloads, Range across chunk boundaries, concurrent identical-upload race (manifest ref 2), and trash-empty reclaiming exactly the unshared chunk while the shared 25 survive for the edited file. The empty/sub-8KB multipart path found a post-EOF re-poll panic in the MIME peek (fixed with fuse + regression test). https://claude.ai/code/session_01WdNenpnujNR2sc32XVvwfS
This commit is contained in:
@@ -251,21 +251,21 @@ pub struct CopyFolderTreeResult {
|
||||
|
||||
/// Secondary port for file **writing**.
|
||||
///
|
||||
/// Covers: upload (buffered + streaming), move, delete, update,
|
||||
/// and deferred registration for the write-behind cache.
|
||||
/// Covers: upload registration, move, delete, update, and deferred
|
||||
/// registration for the write-behind cache.
|
||||
pub trait FileWritePort: Send + Sync + 'static {
|
||||
/// Streaming upload — saves a file from a temp file already on disk.
|
||||
/// Register a file row pointing at a blob already stored in the
|
||||
/// content-addressable chunk store.
|
||||
///
|
||||
/// When `pre_computed_hash` is provided, the dedup service skips the
|
||||
/// hash re-read — zero extra I/O beyond the initial spool.
|
||||
async fn save_file_from_temp(
|
||||
/// Takes ownership of one blob reference: on any failure the reference
|
||||
/// is released before the error is returned.
|
||||
async fn save_file_with_blob(
|
||||
&self,
|
||||
name: String,
|
||||
folder_id: Option<String>,
|
||||
content_type: String,
|
||||
temp_path: &std::path::Path,
|
||||
blob_hash: &str,
|
||||
size: u64,
|
||||
pre_computed_hash: Option<String>,
|
||||
) -> Result<File, DomainError>;
|
||||
|
||||
/// Moves a file to another folder.
|
||||
@@ -281,22 +281,20 @@ pub trait FileWritePort: Send + Sync + 'static {
|
||||
/// Deletes a file.
|
||||
async fn delete_file(&self, id: &str) -> Result<(), DomainError>;
|
||||
|
||||
/// Streaming update — replaces file content from a temp file on disk.
|
||||
/// Atomically swap a file's content to a blob already stored in the
|
||||
/// content-addressable chunk store.
|
||||
///
|
||||
/// When `pre_computed_hash` is provided, the dedup service skips the
|
||||
/// hash re-read — zero extra I/O beyond the initial spool.
|
||||
/// Peak RAM: ~256 KB regardless of file size.
|
||||
/// Takes ownership of one blob reference (released on failure); the
|
||||
/// previous content's reference is dropped after the swap.
|
||||
///
|
||||
/// Returns `(new_blob_hash, updated_at_epoch)` — everything a caller
|
||||
/// needs to rebuild the fresh entity/ETag from a `File` it already
|
||||
/// holds, without re-reading the row it just updated.
|
||||
async fn update_file_content_from_temp(
|
||||
async fn update_file_content_with_blob(
|
||||
&self,
|
||||
file_id: &str,
|
||||
temp_path: &std::path::Path,
|
||||
blob_hash: &str,
|
||||
size: u64,
|
||||
content_type: Option<String>,
|
||||
pre_computed_hash: Option<String>,
|
||||
modified_at: Option<i64>,
|
||||
) -> Result<(String, i64), DomainError>;
|
||||
|
||||
|
||||
Reference in New Issue
Block a user