Stream uploads directly into the CDC chunk store (no spool, single write)

Every upload surface previously wrote each byte to disk twice: the HTTP
body was spooled to a temp file (or assembled from chunk parts), then
mmap-re-read for FastCDC analysis, and finally the new chunks were
written to the blob backend. CDC could not start until the last byte
arrived, so large uploads paid receive + reread + rewrite latency.

The dedup engine now chunks, hashes and settles the stream WHILE it
arrives (fastcdc AsyncStreamCDC + incremental BLAKE3):

- Each batch of distinct chunks is pinned-or-classified by ONE
  `UPDATE … RETURNING` (no check-then-bump TOCTOU; pinned chunks can't
  be reclaimed mid-upload), and only chunks the store doesn't have are
  written — a full dedup hit performs zero content writes.
- Durability before visibility is preserved: one batched fsync sweep,
  then one batched INSERT, then the manifest. Identical concurrent
  uploads are resolved at the manifest INSERT via ON CONFLICT (the
  loser releases its references and becomes a dedup hit).
- A drop guard rolls back pins and surfaces written-but-unregistered
  chunks to GC if the request future is cancelled mid-stream.
- MIME sniffing now peeks the first bytes in-flight; client-requested
  MD5/SHA-256 checksums are computed by a stream tee — the post-upload
  re-read of the assembled file is gone.

All surfaces converge on the new interfaces::upload_ingest helper:
REST multipart, WebDAV PUT, NextCloud PUT, WOPI PutFile, the dedup
endpoint, and both chunked-upload completions (which now stream their
ordered parts straight into the store instead of writing an assembled
file — chunk parts persist until finalize, so completion is genuinely
retryable). The legacy blob re-chunk migration streams from the
backend with no spool file either.

Legacy removed: store_from_file + mmap CDC analysers + temp-path
plumbing through every port (pre_computed_hash, save_file_from_temp,
update_file_content_from_temp), upload_spool + assembled-file
assembly in both chunked services, create_file/update_file byte-slice
variants (no callers), common::temp, the OXICLOUD_UPLOAD_TMPDIR
config, and the memmap2 dependency.

Verified end-to-end against PostgreSQL 16: 8 MB upload (26 chunks),
identical re-upload (dedup hit, zero writes), 3-byte edit re-upload
(26 chunks, 1 written), byte-identical downloads, Range across chunk
boundaries, concurrent identical-upload race (manifest ref 2), and
trash-empty reclaiming exactly the unshared chunk while the shared 25
survive for the edited file. The empty/sub-8KB multipart path found a
post-EOF re-poll panic in the MIME peek (fixed with fuse + regression
test).

https://claude.ai/code/session_01WdNenpnujNR2sc32XVvwfS
This commit is contained in:
Claude
2026-06-11 13:06:33 +00:00
parent 7157454afd
commit e3f04d58aa
34 changed files with 1864 additions and 2440 deletions
+4 -19
View File
@@ -228,12 +228,6 @@ pub struct StorageConfig {
/// returns 413 with a "use chunked upload" hint when a direct PUT
/// exceeds this cap. Env: `OXICLOUD_DIRECT_PUT_MAX_BYTES`.
pub direct_put_max_bytes: usize,
/// Directory for upload spool temp files. When `Some`, large uploads are
/// spooled here instead of the OS default temp dir (often tmpfs/RAM in
/// containers, where the spool's page-cache counts against the cgroup
/// memory limit and can trigger OOMKill on large files). Env:
/// `OXICLOUD_UPLOAD_TMPDIR`.
pub upload_temp_dir: Option<PathBuf>,
/// Root directory for chunked-upload sessions. When `Some`, chunks land
/// under `{chunk_dir}/{upload_id}/` (REST) and
/// `{chunk_dir}/nextcloud/{user}/{upload_id}/` (NC). When `None`, falls
@@ -401,7 +395,6 @@ impl Default for StorageConfig {
max_upload_size: MAX_UPLOAD_SIZE,
chunk_max_bytes: 100 * 1024 * 1024, // 100 MB — sane upper bound for a single chunked-upload PUT
direct_put_max_bytes: 1024 * 1024 * 1024, // 1 GiB — pushes larger uploads onto the chunked protocol
upload_temp_dir: None,
chunk_dir: None,
usage_reconcile_secs: 600, // 10 minutes
tree_etag_flush_ms: 500,
@@ -1281,18 +1274,10 @@ impl AppConfig {
config.storage.direct_put_max_bytes = val;
}
// Upload spool directory — keep large upload temp files off tmpfs/RAM
// (otherwise their page-cache counts against the cgroup memory limit).
if let Ok(dir) = env::var("OXICLOUD_UPLOAD_TMPDIR")
&& !dir.trim().is_empty()
{
config.storage.upload_temp_dir = Some(PathBuf::from(dir.trim()));
}
// Chunked-upload session root — separate from the PUT spool because
// chunked sessions accumulate disk on long uploads (multi-chunk
// resumable transfers) while PUT spool is short-lived. Sysadmins
// commonly want one of them on fast/local storage (NVMe) and the
// other on bulk storage; this knob lets that be expressed.
// Chunked-upload session root — chunked sessions accumulate disk on
// long uploads (multi-chunk resumable transfers); sysadmins commonly
// want them on fast/local storage (NVMe). This knob lets that be
// expressed.
if let Ok(dir) = env::var("OXICLOUD_CHUNK_DIR")
&& !dir.trim().is_empty()
{
+1 -2
View File
@@ -452,8 +452,7 @@ impl AppServiceFactory {
repos.file_read_repository.clone(),
)
.with_content_cache(core.file_content_cache.clone())
.with_file_lifecycle_hook(core.file_lifecycle.clone())
.with_upload_temp_dir(self.config.storage.upload_temp_dir.clone()),
.with_file_lifecycle_hook(core.file_lifecycle.clone()),
);
let file_retrieval_service = Arc::new(FileRetrievalService::new_with_cache(
+19 -90
View File
@@ -8,21 +8,25 @@
//!
//! Performance: < 1µs for the `infer` check (reads only header bytes, no allocation).
use std::path::Path;
use tokio::io::AsyncReadExt;
/// Maximum bytes to read for magic-byte detection.
const MAGIC_BYTES_LEN: usize = 8192;
/// Maximum bytes needed for magic-byte detection. Upload ingestion peeks
/// this many bytes off the stream before forwarding them unchanged.
pub const MAGIC_BYTES_LEN: usize = 8192;
/// Extract the filename component from a `/`-separated path.
pub fn filename_from_path(path: &str) -> &str {
path.rsplit('/').next().unwrap_or(path)
}
/// Whether a claimed Content-Type is too generic to trust — these trigger
/// magic-byte detection on the upload path.
pub fn is_generic_mime(claimed: &str) -> bool {
claimed.is_empty() || claimed == "application/octet-stream" || claimed == "binary/octet-stream"
}
/// Refine a claimed MIME type using magic bytes and filename extension.
///
/// This is a synchronous function — the caller should already have the first
/// bytes of the file available (or call the async wrapper below).
/// bytes of the content available (upload ingestion peeks them in-flight).
///
/// # Arguments
/// * `buf` — first bytes of the file (at least 8192 for best results)
@@ -30,10 +34,7 @@ pub fn filename_from_path(path: &str) -> &str {
/// * `claimed` — the Content-Type sent by the client
pub fn refine_content_type(buf: &[u8], filename: &str, claimed: &str) -> String {
// If the client sent a specific type (not generic), trust it
if !claimed.is_empty()
&& claimed != "application/octet-stream"
&& claimed != "binary/octet-stream"
{
if !is_generic_mime(claimed) {
return claimed.to_string();
}
@@ -52,49 +53,9 @@ pub fn refine_content_type(buf: &[u8], filename: &str, claimed: &str) -> String
claimed.to_string()
}
/// Async helper: reads the first bytes of a file on disk and refines the MIME type.
///
/// Designed for the upload path where the file has been spooled to a temp path.
pub async fn refine_content_type_from_file(
temp_path: &Path,
filename: &str,
claimed: &str,
) -> String {
// Fast path: if the client gave us a specific type, trust it
if !claimed.is_empty()
&& claimed != "application/octet-stream"
&& claimed != "binary/octet-stream"
{
return claimed.to_string();
}
// Read only the first bytes needed for magic detection (not the whole file).
match tokio::fs::File::open(temp_path).await {
Ok(mut file) => {
let mut buf = vec![0u8; MAGIC_BYTES_LEN];
let n = file.read(&mut buf).await.unwrap_or(0);
refine_content_type(&buf[..n], filename, claimed)
}
Err(e) => {
tracing::warn!(
"MIME detection: failed to read {} for magic bytes: {}",
temp_path.display(),
e
);
// Fall back to extension
let guess = mime_guess::from_path(filename);
if let Some(mime) = guess.first() {
return mime.to_string();
}
claimed.to_string()
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::io::Write;
// ── refine_content_type (sync) ──────────────────────────────
@@ -145,47 +106,15 @@ mod tests {
assert_eq!(result, "image/png");
}
// ── refine_content_type_from_file (async) ───────────────────
// ── is_generic_mime ─────────────────────────────────────────
#[tokio::test]
async fn from_file_detects_png() {
let mut tmp = tempfile::NamedTempFile::new().unwrap();
let png = b"\x89PNG\r\n\x1a\n\x00\x00\x00\rIHDR";
tmp.write_all(png).unwrap();
tmp.flush().unwrap();
let result =
refine_content_type_from_file(tmp.path(), "photo", "application/octet-stream").await;
assert_eq!(result, "image/png");
}
#[tokio::test]
async fn from_file_falls_back_to_extension() {
let mut tmp = tempfile::NamedTempFile::new().unwrap();
tmp.write_all(b"not magic").unwrap();
tmp.flush().unwrap();
let result =
refine_content_type_from_file(tmp.path(), "doc.css", "application/octet-stream").await;
assert_eq!(result, "text/css");
}
#[tokio::test]
async fn from_file_trusts_specific_claimed() {
let result =
refine_content_type_from_file(Path::new("/nonexistent"), "file", "image/webp").await;
assert_eq!(result, "image/webp");
}
#[tokio::test]
async fn from_file_missing_file_falls_back_to_extension() {
let result = refine_content_type_from_file(
Path::new("/nonexistent/file"),
"photo.jpg",
"application/octet-stream",
)
.await;
assert_eq!(result, "image/jpeg");
#[test]
fn generic_mime_detection() {
assert!(is_generic_mime(""));
assert!(is_generic_mime("application/octet-stream"));
assert!(is_generic_mime("binary/octet-stream"));
assert!(!is_generic_mime("image/png"));
assert!(!is_generic_mime("text/plain"));
}
// ── filename_from_path ──────────────────────────────────────
-1
View File
@@ -4,4 +4,3 @@ pub mod errors;
pub mod locale;
pub mod mime_detect;
pub mod stubs;
pub mod temp;
+8 -59
View File
@@ -24,6 +24,7 @@ use crate::application::dtos::search_dto::{
};
use crate::application::ports::file_ports::{
FileManagementUseCase, FileRetrievalUseCase, FileUploadUseCase, OptimizedFileContent,
StoredBlob,
};
use crate::application::ports::folder_ports::FolderUseCase;
@@ -149,14 +150,13 @@ impl FileReadPort for StubFileReadPort {
pub struct StubFileWritePort;
impl FileWritePort for StubFileWritePort {
async fn save_file_from_temp(
async fn save_file_with_blob(
&self,
_name: String,
_folder_id: Option<String>,
_content_type: String,
_temp_path: &Path,
_blob_hash: &str,
_size: u64,
_pre_computed_hash: Option<String>,
) -> Result<File, DomainError> {
Ok(File::default())
}
@@ -185,13 +185,11 @@ impl FileWritePort for StubFileWritePort {
Ok(())
}
async fn update_file_content_from_temp(
async fn update_file_content_with_blob(
&self,
_file_id: &str,
_temp_path: &Path,
_blob_hash: &str,
_size: u64,
_content_type: Option<String>,
_pre_computed_hash: Option<String>,
_modified_at: Option<i64>,
) -> Result<(String, i64), DomainError> {
Ok((String::new(), 0))
@@ -478,40 +476,7 @@ impl FileUploadUseCase for StubFileUploadUseCase {
_name: String,
_folder_id: Option<String>,
_content_type: String,
_temp_path: &Path,
_size: u64,
_pre_computed_hash: Option<String>,
) -> Result<FileDto, DomainError> {
Ok(FileDto::default())
}
async fn upload_file_from_path(
&self,
_name: String,
_folder_id: Option<String>,
_content_type: String,
_file_path: &Path,
_pre_computed_hash: Option<String>,
) -> Result<FileDto, DomainError> {
Ok(FileDto::default())
}
async fn create_file(
&self,
_parent_path: &str,
_filename: &str,
_content: &[u8],
_content_type: &str,
) -> Result<FileDto, DomainError> {
Ok(FileDto::default())
}
async fn update_file(
&self,
_path: &str,
_content: &[u8],
_content_type: &str,
_modified_at: Option<i64>,
_blob: StoredBlob,
) -> Result<FileDto, DomainError> {
Ok(FileDto::default())
}
@@ -519,10 +484,8 @@ impl FileUploadUseCase for StubFileUploadUseCase {
async fn update_file_streaming(
&self,
_path: &str,
_temp_path: &Path,
_size: u64,
_blob: StoredBlob,
_content_type: &str,
_pre_computed_hash: Option<String>,
_modified_at: Option<i64>,
) -> Result<FileDto, DomainError> {
Ok(FileDto::default())
@@ -756,25 +719,11 @@ impl SearchUseCase for StubSearchUseCase {
// DedupPort
// ---------------------------------------------------------------------------
use crate::application::ports::dedup_ports::{
BlobMetadataDto, DedupPort, DedupResultDto, DedupStatsDto,
};
use crate::application::ports::dedup_ports::{BlobMetadataDto, DedupPort, DedupStatsDto};
pub struct StubDedupPort;
impl DedupPort for StubDedupPort {
async fn store_from_file(
&self,
_source_path: &Path,
_content_type: Option<String>,
_pre_computed_hash: Option<String>,
) -> Result<DedupResultDto, DomainError> {
Err(DomainError::internal_error(
"DedupService",
"DedupService not initialized",
))
}
async fn blob_exists(&self, _hash: &str) -> bool {
false
}
-26
View File
@@ -1,26 +0,0 @@
//! Shared helper for creating upload spool temp files.
//!
//! Upload paths spool the request body to a temp file before deduplication.
//! By default `tempfile` uses the OS temp dir (`std::env::temp_dir()`, i.e.
//! `$TMPDIR` / `/tmp`), which in many container setups is **tmpfs (RAM)**.
//! Writing a multi-hundred-MB upload there fills page-cache that counts
//! against the cgroup memory limit and can OOMKill the process. Pointing the
//! spool at a real-disk directory (`OXICLOUD_UPLOAD_TMPDIR`) keeps the upload
//! footprint proportional to the streaming buffer, not the file size.
use std::path::Path;
use tempfile::NamedTempFile;
/// Create a [`NamedTempFile`], honoring an optional configured spool directory.
///
/// When `dir` is `Some`, the temp file is created there (the directory is
/// created if missing); otherwise the OS default temp dir is used.
pub fn new_spool_temp_file(dir: Option<&Path>) -> std::io::Result<NamedTempFile> {
match dir {
Some(d) => {
std::fs::create_dir_all(d)?;
NamedTempFile::new_in(d)
}
None => NamedTempFile::new(),
}
}