Stream uploads directly into the CDC chunk store (no spool, single write)
Every upload surface previously wrote each byte to disk twice: the HTTP body was spooled to a temp file (or assembled from chunk parts), then mmap-re-read for FastCDC analysis, and finally the new chunks were written to the blob backend. CDC could not start until the last byte arrived, so large uploads paid receive + reread + rewrite latency. The dedup engine now chunks, hashes and settles the stream WHILE it arrives (fastcdc AsyncStreamCDC + incremental BLAKE3): - Each batch of distinct chunks is pinned-or-classified by ONE `UPDATE … RETURNING` (no check-then-bump TOCTOU; pinned chunks can't be reclaimed mid-upload), and only chunks the store doesn't have are written — a full dedup hit performs zero content writes. - Durability before visibility is preserved: one batched fsync sweep, then one batched INSERT, then the manifest. Identical concurrent uploads are resolved at the manifest INSERT via ON CONFLICT (the loser releases its references and becomes a dedup hit). - A drop guard rolls back pins and surfaces written-but-unregistered chunks to GC if the request future is cancelled mid-stream. - MIME sniffing now peeks the first bytes in-flight; client-requested MD5/SHA-256 checksums are computed by a stream tee — the post-upload re-read of the assembled file is gone. All surfaces converge on the new interfaces::upload_ingest helper: REST multipart, WebDAV PUT, NextCloud PUT, WOPI PutFile, the dedup endpoint, and both chunked-upload completions (which now stream their ordered parts straight into the store instead of writing an assembled file — chunk parts persist until finalize, so completion is genuinely retryable). The legacy blob re-chunk migration streams from the backend with no spool file either. Legacy removed: store_from_file + mmap CDC analysers + temp-path plumbing through every port (pre_computed_hash, save_file_from_temp, update_file_content_from_temp), upload_spool + assembled-file assembly in both chunked services, create_file/update_file byte-slice variants (no callers), common::temp, the OXICLOUD_UPLOAD_TMPDIR config, and the memmap2 dependency. Verified end-to-end against PostgreSQL 16: 8 MB upload (26 chunks), identical re-upload (dedup hit, zero writes), 3-byte edit re-upload (26 chunks, 1 written), byte-identical downloads, Range across chunk boundaries, concurrent identical-upload race (manifest ref 2), and trash-empty reclaiming exactly the unshared chunk while the shared 25 survive for the edited file. The empty/sub-8KB multipart path found a post-EOF re-poll panic in the MIME peek (fixed with fuse + regression test). https://claude.ai/code/session_01WdNenpnujNR2sc32XVvwfS
This commit is contained in:
@@ -165,10 +165,6 @@ async fn put_file(
|
||||
State(state): State<WopiState>,
|
||||
req: Request<Body>,
|
||||
) -> Response {
|
||||
use http_body_util::BodyStream;
|
||||
use tokio::io::AsyncWriteExt;
|
||||
use tokio_stream::StreamExt;
|
||||
|
||||
let claims = match state
|
||||
.token_service
|
||||
.validate_token(&token_query.access_token)
|
||||
@@ -218,75 +214,32 @@ async fn put_file(
|
||||
Err(_) => return StatusCode::NOT_FOUND.into_response(),
|
||||
};
|
||||
|
||||
// ── Streaming spool: body → temp file + incremental BLAKE3 ──
|
||||
let temp_file = match tempfile::NamedTempFile::new() {
|
||||
Ok(f) => f,
|
||||
Err(e) => {
|
||||
tracing::error!("WOPI PutFile: failed to create temp file: {}", e);
|
||||
return StatusCode::INTERNAL_SERVER_ERROR.into_response();
|
||||
}
|
||||
};
|
||||
let temp_path = temp_file.path().to_path_buf();
|
||||
|
||||
let mut file_out = match tokio::fs::File::create(&temp_path).await {
|
||||
Ok(f) => f,
|
||||
Err(e) => {
|
||||
tracing::error!("WOPI PutFile: failed to open temp file: {}", e);
|
||||
return StatusCode::INTERNAL_SERVER_ERROR.into_response();
|
||||
}
|
||||
};
|
||||
|
||||
// ── Streaming ingest: body → CDC chunk store (no temp file) ──
|
||||
let content_type = file.mime_type.clone();
|
||||
let mut hasher = blake3::Hasher::new();
|
||||
let mut total_bytes: u64 = 0;
|
||||
let mut stream = BodyStream::new(req.into_body());
|
||||
|
||||
while let Some(frame_result) = stream.next().await {
|
||||
let frame = match frame_result {
|
||||
Ok(f) => f,
|
||||
Err(e) => {
|
||||
let _ = tokio::fs::remove_file(&temp_path).await;
|
||||
tracing::error!("WOPI PutFile: body read error: {}", e);
|
||||
return StatusCode::INTERNAL_SERVER_ERROR.into_response();
|
||||
}
|
||||
};
|
||||
if let Some(chunk) = frame.data_ref() {
|
||||
total_bytes += chunk.len() as u64;
|
||||
hasher.update(chunk);
|
||||
if let Err(e) = file_out.write_all(chunk).await {
|
||||
let _ = tokio::fs::remove_file(&temp_path).await;
|
||||
tracing::error!("WOPI PutFile: temp write error: {}", e);
|
||||
return StatusCode::INTERNAL_SERVER_ERROR.into_response();
|
||||
}
|
||||
let ingested = match crate::interfaces::upload_ingest::ingest_body_to_cas(
|
||||
req.into_body(),
|
||||
&state.app_state.core.dedup_service,
|
||||
&file.name,
|
||||
&content_type,
|
||||
usize::MAX,
|
||||
)
|
||||
.await
|
||||
{
|
||||
Ok(ingested) => ingested,
|
||||
Err(e) => {
|
||||
tracing::error!("WOPI PutFile: ingest failed: {}", e.message);
|
||||
return StatusCode::INTERNAL_SERVER_ERROR.into_response();
|
||||
}
|
||||
}
|
||||
if let Err(e) = file_out.flush().await {
|
||||
let _ = tokio::fs::remove_file(&temp_path).await;
|
||||
tracing::error!("WOPI PutFile: flush error: {}", e);
|
||||
return StatusCode::INTERNAL_SERVER_ERROR.into_response();
|
||||
}
|
||||
drop(file_out);
|
||||
};
|
||||
|
||||
let hash = hasher.finalize().to_hex().to_string();
|
||||
|
||||
// ── Atomic store: temp file → dedup blob + DB metadata update ──
|
||||
// ── Atomic store: swap the file row onto the ingested blob ──
|
||||
let result = state
|
||||
.app_state
|
||||
.applications
|
||||
.file_upload_service
|
||||
.update_file_streaming(
|
||||
&file.path,
|
||||
&temp_path,
|
||||
total_bytes,
|
||||
&content_type,
|
||||
Some(hash),
|
||||
None,
|
||||
)
|
||||
.update_file_streaming(&file.path, ingested.stored(), &content_type, None)
|
||||
.await;
|
||||
|
||||
// Clean up temp file (may already be moved by dedup, ignore error)
|
||||
let _ = tokio::fs::remove_file(&temp_path).await;
|
||||
|
||||
match result {
|
||||
Ok(_file_dto) => StatusCode::OK.into_response(),
|
||||
Err(e) => {
|
||||
|
||||
Reference in New Issue
Block a user