Sweep aborted-upload orphans from the periodic trash job; pipeline ZIP reads
Two follow-ups to the streaming-upload work: The dedup garbage_collect() pass only ran when a user manually emptied their trash. Zero-reference rows — chunks orphaned by aborted streaming uploads (registered at ref_count 0 by the ingest rollback) and blobs dereferenced by trash expiry itself — could linger indefinitely on instances where nobody empties trash. The periodic TrashCleanupService sweep now ends every run with garbage_collect() (maintenance pool, batched), bounding orphan lifetime to the cleanup interval. Folder-ZIP creation was strictly sequential: open blob stream, deflate, close, repeat — every per-file blob-store round-trip (PG lookup + backend open; a full HTTP round-trip on S3/Azure) added to the wall clock. It now runs as a 2-stage pipeline: a prefetch task streams the planned files' content ahead of the writer through a bounded channel (~4 MiB), so the next file's read latency overlaps the current file's compression. ZIP entries are still written strictly in order, peak RAM stays flat, and a writer error hangs up the channel so the prefetcher stops on its own. Verified end-to-end against PostgreSQL 16: an upload aborted at ~14 MB left exactly 30 ref_count=0 chunk rows which the GC then reclaimed (8.5 MB, rows + physical files); a chunked upload completed with a wrong MD5 returned 400 with the tee-computed digest and the SAME session then completed successfully with the right checksum (parts persist — the old assembly deleted them, so the documented retry never actually worked); a 4-file folder ZIP downloaded and extracted byte-identical. https://claude.ai/code/session_01WdNenpnujNR2sc32XVvwfS
This commit is contained in:
@@ -6,21 +6,35 @@ use tracing::{debug, error, info, instrument};
|
||||
use crate::common::errors::Result;
|
||||
use crate::domain::repositories::trash_repository::TrashRepository;
|
||||
use crate::infrastructure::repositories::pg::trash_db_repository::TrashDbRepository;
|
||||
use crate::infrastructure::services::dedup_service::DedupService;
|
||||
|
||||
/// Service for automatic cleanup of expired items in the trash.
|
||||
///
|
||||
/// Uses `TrashRepository::delete_expired_bulk` to purge all expired items
|
||||
/// in **2 SQL statements inside a single transaction**, instead of the
|
||||
/// previous N+1 pattern that issued 3 queries per expired item.
|
||||
///
|
||||
/// Each sweep ends with a dedup `garbage_collect()` pass: it reclaims the
|
||||
/// blobs the expiry just dereferenced AND any other zero-reference rows —
|
||||
/// notably chunks left behind by aborted streaming uploads, whose rollback
|
||||
/// registers them at ref_count 0 precisely so this sweep can find them.
|
||||
/// Without it, orphans would only be collected when a user happens to
|
||||
/// empty their trash by hand.
|
||||
pub struct TrashCleanupService {
|
||||
trash_repository: Arc<TrashDbRepository>,
|
||||
dedup_service: Arc<DedupService>,
|
||||
cleanup_interval_hours: u64,
|
||||
}
|
||||
|
||||
impl TrashCleanupService {
|
||||
pub fn new(trash_repository: Arc<TrashDbRepository>, cleanup_interval_hours: u64) -> Self {
|
||||
pub fn new(
|
||||
trash_repository: Arc<TrashDbRepository>,
|
||||
dedup_service: Arc<DedupService>,
|
||||
cleanup_interval_hours: u64,
|
||||
) -> Self {
|
||||
Self {
|
||||
trash_repository,
|
||||
dedup_service,
|
||||
cleanup_interval_hours: cleanup_interval_hours.max(1), // Minimum 1 hour
|
||||
}
|
||||
}
|
||||
@@ -29,6 +43,7 @@ impl TrashCleanupService {
|
||||
#[instrument(skip(self))]
|
||||
pub async fn start_cleanup_job(&self) {
|
||||
let trash_repository = self.trash_repository.clone();
|
||||
let dedup_service = self.dedup_service.clone();
|
||||
let interval_hours = self.cleanup_interval_hours;
|
||||
|
||||
info!(
|
||||
@@ -41,7 +56,7 @@ impl TrashCleanupService {
|
||||
let mut interval = time::interval(interval_duration);
|
||||
|
||||
// First immediate execution
|
||||
Self::cleanup_expired_items(trash_repository.clone())
|
||||
Self::cleanup_expired_items(trash_repository.clone(), dedup_service.clone())
|
||||
.await
|
||||
.unwrap_or_else(|e| error!("Error in initial trash cleanup: {:?}", e));
|
||||
|
||||
@@ -49,16 +64,24 @@ impl TrashCleanupService {
|
||||
interval.tick().await;
|
||||
debug!("Running scheduled trash cleanup task");
|
||||
|
||||
if let Err(e) = Self::cleanup_expired_items(trash_repository.clone()).await {
|
||||
if let Err(e) =
|
||||
Self::cleanup_expired_items(trash_repository.clone(), dedup_service.clone())
|
||||
.await
|
||||
{
|
||||
error!("Error in scheduled trash cleanup: {:?}", e);
|
||||
}
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
/// Bulk-delete all expired trash items in a single transaction.
|
||||
#[instrument(skip(trash_repository))]
|
||||
async fn cleanup_expired_items(trash_repository: Arc<TrashDbRepository>) -> Result<()> {
|
||||
/// Bulk-delete all expired trash items in a single transaction, then
|
||||
/// garbage-collect every zero-reference manifest/blob (expired content
|
||||
/// plus aborted-upload orphans).
|
||||
#[instrument(skip(trash_repository, dedup_service))]
|
||||
async fn cleanup_expired_items(
|
||||
trash_repository: Arc<TrashDbRepository>,
|
||||
dedup_service: Arc<DedupService>,
|
||||
) -> Result<()> {
|
||||
debug!("Starting bulk cleanup of expired trash items");
|
||||
|
||||
let (files, folders) = trash_repository.delete_expired_bulk().await?;
|
||||
@@ -72,6 +95,16 @@ impl TrashCleanupService {
|
||||
);
|
||||
}
|
||||
|
||||
// Runs on the maintenance pool; batched (500 rows/iteration) with
|
||||
// yield points, so it never starves request-path queries.
|
||||
match dedup_service.garbage_collect().await {
|
||||
Ok((0, _)) => debug!("Trash cleanup GC: nothing to collect"),
|
||||
Ok((items, bytes)) => {
|
||||
info!("Trash cleanup GC: reclaimed {items} orphaned blobs ({bytes} bytes)");
|
||||
}
|
||||
Err(e) => error!("Trash cleanup GC failed: {:?}", e),
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user