e3f04d58aa
Every upload surface previously wrote each byte to disk twice: the HTTP body was spooled to a temp file (or assembled from chunk parts), then mmap-re-read for FastCDC analysis, and finally the new chunks were written to the blob backend. CDC could not start until the last byte arrived, so large uploads paid receive + reread + rewrite latency. The dedup engine now chunks, hashes and settles the stream WHILE it arrives (fastcdc AsyncStreamCDC + incremental BLAKE3): - Each batch of distinct chunks is pinned-or-classified by ONE `UPDATE … RETURNING` (no check-then-bump TOCTOU; pinned chunks can't be reclaimed mid-upload), and only chunks the store doesn't have are written — a full dedup hit performs zero content writes. - Durability before visibility is preserved: one batched fsync sweep, then one batched INSERT, then the manifest. Identical concurrent uploads are resolved at the manifest INSERT via ON CONFLICT (the loser releases its references and becomes a dedup hit). - A drop guard rolls back pins and surfaces written-but-unregistered chunks to GC if the request future is cancelled mid-stream. - MIME sniffing now peeks the first bytes in-flight; client-requested MD5/SHA-256 checksums are computed by a stream tee — the post-upload re-read of the assembled file is gone. All surfaces converge on the new interfaces::upload_ingest helper: REST multipart, WebDAV PUT, NextCloud PUT, WOPI PutFile, the dedup endpoint, and both chunked-upload completions (which now stream their ordered parts straight into the store instead of writing an assembled file — chunk parts persist until finalize, so completion is genuinely retryable). The legacy blob re-chunk migration streams from the backend with no spool file either. Legacy removed: store_from_file + mmap CDC analysers + temp-path plumbing through every port (pre_computed_hash, save_file_from_temp, update_file_content_from_temp), upload_spool + assembled-file assembly in both chunked services, create_file/update_file byte-slice variants (no callers), common::temp, the OXICLOUD_UPLOAD_TMPDIR config, and the memmap2 dependency. Verified end-to-end against PostgreSQL 16: 8 MB upload (26 chunks), identical re-upload (dedup hit, zero writes), 3-byte edit re-upload (26 chunks, 1 written), byte-identical downloads, Range across chunk boundaries, concurrent identical-upload race (manifest ref 2), and trash-empty reclaiming exactly the unshared chunk while the shared 25 survive for the edited file. The empty/sub-8KB multipart path found a post-EOF re-poll panic in the MIME peek (fixed with fuse + regression test). https://claude.ai/code/session_01WdNenpnujNR2sc32XVvwfS
174 lines
6.0 KiB
Rust
174 lines
6.0 KiB
Rust
//! Chunked Upload Port - Application layer abstraction for resumable chunked uploads.
|
|
//!
|
|
//! This module defines the port (trait) and DTOs for chunked/resumable upload
|
|
//! operations, keeping the application and interface layers independent of
|
|
//! the specific upload implementation (TUS-like protocol, S3 multipart, etc.).
|
|
|
|
use crate::common::errors::DomainError;
|
|
use bytes::Bytes;
|
|
use serde::Serialize;
|
|
use std::path::PathBuf;
|
|
use utoipa::ToSchema;
|
|
use uuid::Uuid;
|
|
|
|
/// Default chunk size (5 MB) — optimised for parallel transfers.
|
|
pub const DEFAULT_CHUNK_SIZE: usize = 5 * 1024 * 1024;
|
|
|
|
/// Minimum file size to use chunked upload (10 MB).
|
|
pub const CHUNKED_UPLOAD_THRESHOLD: usize = 10 * 1024 * 1024;
|
|
|
|
/// Algorithm used by the client-side chunk checksum.
|
|
///
|
|
/// The wire format is `?checksum=<hex>&checksumalg=<name>` (or the
|
|
/// equivalent header pair for older clients that send only `Content-MD5`).
|
|
/// Clients that omit `checksumalg` are assumed to mean MD5 — that's the
|
|
/// algorithm baked into the legacy `Content-MD5` header (RFC 1864), TUS-
|
|
/// like upload protocols, and S3 multipart ETags.
|
|
///
|
|
/// Three supported variants, all from already-declared dependencies:
|
|
/// - `Md5` — legacy default; weak cryptographically but fine for
|
|
/// transport-integrity checks under TLS.
|
|
/// - `Sha256` — industry-standard, FIPS-compliant, widely supported by
|
|
/// sync clients (AWS S3 also accepts SHA-256 trailers).
|
|
/// - `Blake3` — fastest of the three; already used by the blob-storage
|
|
/// layer, so the chunk-level integrity check and the assembled-file
|
|
/// dedup hash use the same algorithm when clients opt in.
|
|
///
|
|
/// Skipped intentionally: SHA-1 (deprecated, broken), CRC32 (too weak for
|
|
/// integrity claims). Both can be added if a real client need appears.
|
|
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
|
pub enum ChecksumAlg {
|
|
Md5,
|
|
Sha256,
|
|
Blake3,
|
|
}
|
|
|
|
impl ChecksumAlg {
|
|
/// Parse a client-supplied algorithm name. Case-insensitive. Accepts
|
|
/// `sha-256` as a synonym for `sha256` since both forms are common
|
|
/// in HTTP headers. Unknown names return `None` so the handler can
|
|
/// 400 with the offending value.
|
|
pub fn parse(s: &str) -> Option<Self> {
|
|
match s.trim().to_ascii_lowercase().as_str() {
|
|
"md5" => Some(Self::Md5),
|
|
"sha256" | "sha-256" => Some(Self::Sha256),
|
|
"blake3" => Some(Self::Blake3),
|
|
_ => None,
|
|
}
|
|
}
|
|
|
|
pub fn as_str(self) -> &'static str {
|
|
match self {
|
|
Self::Md5 => "md5",
|
|
Self::Sha256 => "sha256",
|
|
Self::Blake3 => "blake3",
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Response returned when a new upload session is created.
|
|
#[derive(Debug, Clone, Serialize, ToSchema)]
|
|
pub struct CreateUploadResponseDto {
|
|
pub upload_id: String,
|
|
pub chunk_size: usize,
|
|
pub total_chunks: usize,
|
|
pub expires_at: u64,
|
|
}
|
|
|
|
/// Response returned after a single chunk is uploaded.
|
|
#[derive(Debug, Clone, Serialize, ToSchema)]
|
|
pub struct ChunkUploadResponseDto {
|
|
pub chunk_index: usize,
|
|
pub bytes_received: u64,
|
|
pub progress: f64,
|
|
pub is_complete: bool,
|
|
}
|
|
|
|
/// Response for querying upload session status.
|
|
#[derive(Debug, Clone, Serialize, ToSchema)]
|
|
pub struct UploadStatusResponseDto {
|
|
pub upload_id: String,
|
|
pub filename: String,
|
|
pub total_size: u64,
|
|
pub bytes_received: u64,
|
|
pub progress: f64,
|
|
pub total_chunks: usize,
|
|
pub completed_chunks: usize,
|
|
pub pending_chunks: Vec<usize>,
|
|
pub is_complete: bool,
|
|
}
|
|
|
|
/// A completed upload session, ready to be streamed into the blob store.
|
|
///
|
|
/// `chunk_paths` lists every chunk file in assembly order — the caller
|
|
/// concatenates them as one byte stream (CDC chunking + hashing happen in
|
|
/// that single pass). The files stay on disk until `finalize_upload`, so a
|
|
/// failed completion (e.g. checksum mismatch) remains retryable.
|
|
#[derive(Debug, Clone)]
|
|
pub struct CompletedUploadParts {
|
|
pub chunk_paths: Vec<PathBuf>,
|
|
pub filename: String,
|
|
pub folder_id: Option<String>,
|
|
pub content_type: String,
|
|
pub total_size: u64,
|
|
}
|
|
|
|
/// Port for chunked/resumable upload operations.
|
|
///
|
|
/// Implementations manage upload sessions, chunk storage, reassembly,
|
|
/// and cleanup, while the application layer only interacts through
|
|
/// this abstraction.
|
|
pub trait ChunkedUploadPort: Send + Sync + 'static {
|
|
/// Create a new upload session.
|
|
///
|
|
/// Returns session metadata including the upload ID, chunk size,
|
|
/// total number of chunks, and expiration timestamp.
|
|
async fn create_session(
|
|
&self,
|
|
user_id: Uuid,
|
|
filename: String,
|
|
folder_id: Option<String>,
|
|
content_type: String,
|
|
total_size: u64,
|
|
chunk_size: Option<usize>,
|
|
) -> Result<CreateUploadResponseDto, DomainError>;
|
|
|
|
/// Upload a single chunk.
|
|
///
|
|
/// `checksum` is an optional MD5 hex string for integrity verification.
|
|
async fn upload_chunk(
|
|
&self,
|
|
upload_id: &str,
|
|
user_id: Uuid,
|
|
chunk_index: usize,
|
|
data: Bytes,
|
|
checksum: Option<String>,
|
|
) -> Result<ChunkUploadResponseDto, DomainError>;
|
|
|
|
/// Get the current status of an upload session.
|
|
async fn get_status(
|
|
&self,
|
|
upload_id: &str,
|
|
user_id: Uuid,
|
|
) -> Result<UploadStatusResponseDto, DomainError>;
|
|
|
|
/// Validate completion and hand back the ordered chunk parts.
|
|
///
|
|
/// No assembled file is written — the caller streams the parts straight
|
|
/// into the content-addressable store.
|
|
async fn complete_upload(
|
|
&self,
|
|
upload_id: &str,
|
|
user_id: Uuid,
|
|
) -> Result<CompletedUploadParts, DomainError>;
|
|
|
|
/// Finalize upload: clean up the session and temporary files.
|
|
async fn finalize_upload(&self, upload_id: &str, user_id: Uuid) -> Result<(), DomainError>;
|
|
|
|
/// Cancel an upload and clean up all temporary data.
|
|
async fn cancel_upload(&self, upload_id: &str, user_id: Uuid) -> Result<(), DomainError>;
|
|
|
|
/// Check if a file size qualifies for chunked upload.
|
|
fn should_use_chunked(&self, size: u64) -> bool;
|
|
}
|