feat(dedup): CDC sub-file deduplication with FastCDC + parallel chunk storage + dedup skip
- Replace whole-file SHA-256 dedup with FastCDC 2020 content-defined chunking (min 64KB, avg 256KB, max 1MB) + BLAKE3 hashing - Add chunk_manifests table (file_hash → chunk_hashes[] + chunk_sizes[]) - Add put_blob_from_bytes to BlobStorageBackend trait (all 7 backends) - 3-phase store_chunks pipeline: Phase 0: batch-check existing chunks (single PG query) Phase 1: selective disk read (skip existing chunks entirely) Phase 2: parallel upload with buffer_unordered(8) - CDC-aware read_blob_stream and read_blob_range_stream with legacy fallback - Transactional manifest + chunk ref-count cascade on remove_reference - 12 CDC tests (determinism, reassembly, contiguity, sub-file dedup, etc.) - Update deduplication.md to reflect new architecture
This commit is contained in:
@@ -28,16 +28,11 @@ pub struct BlobMetadataDto {
|
||||
#[derive(Debug, Clone)]
|
||||
pub enum DedupResultDto {
|
||||
/// New content was stored (first occurrence).
|
||||
NewBlob {
|
||||
hash: String,
|
||||
size: u64,
|
||||
blob_path: PathBuf,
|
||||
},
|
||||
NewBlob { hash: String, size: u64 },
|
||||
/// Content already existed; a reference was added instead.
|
||||
ExistingBlob {
|
||||
hash: String,
|
||||
size: u64,
|
||||
blob_path: PathBuf,
|
||||
saved_bytes: u64,
|
||||
},
|
||||
}
|
||||
@@ -57,13 +52,6 @@ impl DedupResultDto {
|
||||
}
|
||||
}
|
||||
|
||||
pub fn blob_path(&self) -> &Path {
|
||||
match self {
|
||||
DedupResultDto::NewBlob { blob_path, .. } => blob_path,
|
||||
DedupResultDto::ExistingBlob { blob_path, .. } => blob_path,
|
||||
}
|
||||
}
|
||||
|
||||
pub fn was_deduplicated(&self) -> bool {
|
||||
matches!(self, DedupResultDto::ExistingBlob { .. })
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user