perf: stream_files_in_subtree — replace Vec<File> with async Stream

Replace list_files_in_subtree (fetch_all → Vec) with stream_files_in_subtree
that returns a Pin<Box<dyn Stream<Item = Result<File/FileDto>>>> backed by a
PostgreSQL cursor via sqlx::fetch().

Changes:
- FileReadPort::stream_files_in_subtree() returns streaming cursor (no default)
- FileRetrievalUseCase::stream_files_in_subtree() maps File→FileDto on the fly
- FileBlobReadRepository: async_stream::try_stream! + sqlx::fetch() cursor
- batch_operations: consume stream into HashMap incrementally
- zip_service: consume stream into HashMap incrementally
- All stubs/mocks updated (return empty stream)

Eliminates:
- Double allocation: Vec<(9-tuple)> + Vec<File> materialized simultaneously
- Unbounded RAM proportional to subtree size (was ~500 bytes × N files)
- Latency: callers blocked until last row fetched from PG

RAM is now O(folders) for the HashMap, not O(files).
This commit is contained in:
Dionisio
2026-02-26 00:07:10 +01:00
parent 1ad7a32a61
commit 9f8a6f5177
9 changed files with 127 additions and 69 deletions
+8 -5
View File
@@ -154,12 +154,15 @@ pub trait FileRetrievalUseCase: Send + Sync + 'static {
end: Option<u64>,
) -> Result<Box<dyn Stream<Item = Result<Bytes, std::io::Error>> + Send>, DomainError>;
/// Lists every file in the subtree rooted at `folder_id`.
/// Streams every file in the subtree rooted at `folder_id`.
///
/// Default: falls back to `list_files(Some(folder_id))` (one level).
async fn list_files_in_subtree(&self, folder_id: &str) -> Result<Vec<FileDto>, DomainError> {
self.list_files(Some(folder_id)).await
}
/// Returns a streaming cursor — RAM stays O(1) per row. Callers
/// consume incrementally (e.g. group into a HashMap by folder_id)
/// without materializing the full result set.
async fn stream_files_in_subtree(
&self,
folder_id: &str,
) -> Result<Pin<Box<dyn Stream<Item = Result<FileDto, DomainError>> + Send>>, DomainError>;
/// Lists files in a folder with LIMIT/OFFSET pagination.
///
+10 -6
View File
@@ -3,6 +3,7 @@ use bytes::Bytes;
use futures::Stream;
use serde_json::Value;
use std::path::PathBuf;
use std::pin::Pin;
use crate::application::dtos::search_dto::SearchCriteriaDto;
use crate::common::errors::DomainError;
@@ -94,15 +95,18 @@ pub trait FileReadPort: Send + Sync + 'static {
Ok(all.into_iter().skip(start).take(end - start).collect())
}
/// Lists every file in the subtree rooted at `folder_id`.
/// Streams every file in the subtree rooted at `folder_id`.
///
/// Uses an ltree `<@` join against `storage.folders` so the entire
/// subtree is fetched in a single GiST-indexed query.
/// subtree is resolved in a single GiST-indexed query, but rows are
/// delivered via a PostgreSQL cursor — RAM stays O(1) per row.
///
/// Default: falls back to `list_files(Some(folder_id))` (one level).
async fn list_files_in_subtree(&self, folder_id: &str) -> Result<Vec<File>, DomainError> {
self.list_files(Some(folder_id)).await
}
/// Callers consume the stream incrementally (e.g. build a HashMap
/// keyed by folder_id) without ever materializing the full Vec.
async fn stream_files_in_subtree(
&self,
folder_id: &str,
) -> Result<Pin<Box<dyn Stream<Item = Result<File, DomainError>> + Send>>, DomainError>;
/// Search files with pagination and filtering at database level.
///