perf: keyset/LATERAL SQL shapes, auth+blob-cache single-flight, spool buffers, DTO interning
Round 3 of benchmark-gated optimizations (benches/ROUND3.md; every change gated by a before/after benchmark — an AFTER that did not beat its BEFORE was to be rolled back; none needed it. Equivalence gates assert identical row sequences / byte-identical output on every behavior-preserving rewrite): DB hot paths (local PG16, EXPLAIN-verified): - Web-UI listing (list_resources_paged): cursor pushed INSIDE the folders/files UNION-ALL branches as sargable row-value comparisons with per-branch ORDER/LIMIT + two partial expression indexes (folder_id, LOWER(name), id). 20k-entry folder: 26.6 -> 1.3 ms/page (19.5x); other sort modes at parity or better. New migration 20260918000000. [benches/LISTING-KEYSET.md section in ROUND3] - Photos timeline (list_media_files): per-drive CROSS JOIN LATERAL top-N on the timeline index, joins moved above the top-N. 50k-photo library: 97.4 -> 1.6 ms/page (55.7x). The old "LIMIT stops the scan early" comment was refuted by EXPLAIN. - PROPFIND sub-folders (both DAV surfaces): keyset list_folders_batch off idx_folders_unique_name replaces COUNT(*) OVER() + LIMIT/OFFSET (5k dirs: 79.7 -> 17.9 ms full walk, 4.5x). Concurrency: - Basic-auth cache single-flight (moka try_get_with): 8 concurrent DAV connections at TTL expiry paid 8 Argon2id runs (2.6 s CPU + 8x64 MiB); now 1 (300 ms). Failed verifications remain uncached. - CachedBlobBackend per-hash single-flight + unique tmp names: 16 concurrent cold readers = 16 full remote downloads racing truncating writes on ONE deterministic .tmp (corruptible cache); now 1 download (16x less egress, 2.8x wall on a shared link) and torn files can never be renamed into the cache. I/O and allocations: - Chunk-assembly reads 64K -> 512K buffers (2.3x, 8x fewer syscalls); chunk-spool writes via BufWriter 512K (5.6x, 32x fewer syscalls). - S3/Azure put_blob_from_bytes_unsynced overrides: dedup settle no longer pays a HEAD probe per new chunk (2 RTT -> 1, 1.8x); Azure stops copying every chunk (Bytes -> Body, -0.44 ms - 4 MiB alloc per 4 MiB chunk). - Entity->DTO mapping: Arc<str> interning of closed-set display fields + common MIMEs, 1-alloc etag/size formatting, FolderDto moves instead of clones. File row: 11 -> 4 allocs; folder row: 11.8 -> 1 (2.1x faster). - CardDAV REPORT: deleted dead per-contact vCard pre-generation and the O(N^2) uid scan whose result was discarded (5k contacts: 55.7 -> 5.7 ms, 9.8x); byte-identical XML asserted. - Search-results cache: byte weigher + 32 MiB budget (OXICLOUD_SEARCH_CACHE_MAX_BYTES) replaces the 1000-ENTRY cap that let ~300 MiB of enriched rows sit in RSS; read latency parity. - Dropped aws-config + aws-smithy-types (zero references; -82 dep-graph nodes, three SDK stacks gone from every build). tokio "process" is now an explicit feature (was enabled transitively by aws-config). Frontend: - Cached Intl.DateTimeFormat keyed by (locale, options) in formatDate and 4 sibling callsites: 20k dates 2612 -> 51 ms (51.6x); vitest gate asserts output identity across locales and a 3x floor. Validation: cargo fmt + clippy --all-features --all-targets -D warnings clean; 518 unit + 548 integration-cfg tests green; new-shape endpoints smoke-tested end-to-end over HTTP (all 5 listing sort modes with cursor walks, WebDAV PROPFIND Depth-1, photos timeline, Basic-auth DAV login); frontend npm run check clean, new vitest gates green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EBsU2qEzny3A8WQUEuMNCr
This commit is contained in:
@@ -352,8 +352,25 @@ impl File {
|
||||
/// formula here changes it everywhere — that is the property
|
||||
/// we want.
|
||||
pub fn compute_etag(blob_hash: &str, modified_at: u64) -> String {
|
||||
let prefix: String = blob_hash.chars().take(16).collect();
|
||||
format!("{}-{}", prefix, modified_at)
|
||||
use std::fmt::Write as _;
|
||||
|
||||
// Byte index just past the 16th char (whole string when shorter).
|
||||
// `blob_hash` is lowercase hex ASCII in practice, so this is
|
||||
// effectively `min(len, 16)`, but `char_indices` keeps the slice
|
||||
// char-boundary-safe for exotic fixture values — byte-identical
|
||||
// to the old `chars().take(16).collect::<String>()` without the
|
||||
// intermediate allocation.
|
||||
let end = match blob_hash.char_indices().nth(16) {
|
||||
Some((i, _)) => i,
|
||||
None => blob_hash.len(),
|
||||
};
|
||||
|
||||
// Single allocation: prefix + '-' + up to 20 digits (u64::MAX).
|
||||
let mut etag = String::with_capacity(end + 1 + 20);
|
||||
etag.push_str(&blob_hash[..end]);
|
||||
etag.push('-');
|
||||
let _ = write!(etag, "{modified_at}");
|
||||
etag
|
||||
}
|
||||
|
||||
// Getters
|
||||
|
||||
@@ -7,6 +7,30 @@ use crate::domain::services::path_service::{
|
||||
// Re-export entity errors from the centralized module
|
||||
pub use super::entity_errors::{FolderError, FolderResult};
|
||||
|
||||
/// Owned parts of a [`Folder`] entity, produced by [`Folder::into_parts()`].
|
||||
///
|
||||
/// Consuming a `Folder` into `FolderParts` **moves** every field without
|
||||
/// cloning, eliminating the 3-4 heap allocations that previously occurred
|
||||
/// when converting `Folder → FolderDto` via `.to_string()` on each getter.
|
||||
/// Mirrors [`super::file::FileParts`].
|
||||
pub struct FolderParts {
|
||||
pub id: String,
|
||||
pub name: String,
|
||||
pub storage_path: StoragePath,
|
||||
pub path_string: String,
|
||||
pub parent_id: Option<String>,
|
||||
/// Drive that owns this folder. See [`Folder::drive_id`].
|
||||
pub drive_id: Uuid,
|
||||
pub created_at: u64,
|
||||
pub modified_at: u64,
|
||||
/// Descendant-rollup timestamp. See [`Folder::tree_modified_at`].
|
||||
pub tree_modified_at: u64,
|
||||
/// §14 provenance: original creator. See [`Folder::created_by`].
|
||||
pub created_by: Option<Uuid>,
|
||||
/// §14 provenance: most recent mutator. See [`Folder::updated_by`].
|
||||
pub updated_by: Option<Uuid>,
|
||||
}
|
||||
|
||||
/// Represents a folder entity in the domain
|
||||
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||
pub struct Folder {
|
||||
@@ -219,6 +243,26 @@ impl Folder {
|
||||
})
|
||||
}
|
||||
|
||||
/// Consume the entity and return all fields by ownership.
|
||||
///
|
||||
/// Use this when converting `Folder` into a DTO to avoid cloning
|
||||
/// every `String` field (saves 3-4 heap allocations per folder).
|
||||
pub fn into_parts(self) -> FolderParts {
|
||||
FolderParts {
|
||||
id: self.id,
|
||||
name: self.name,
|
||||
storage_path: self.storage_path,
|
||||
path_string: self.path_string,
|
||||
parent_id: self.parent_id,
|
||||
drive_id: self.drive_id,
|
||||
created_at: self.created_at,
|
||||
modified_at: self.modified_at,
|
||||
tree_modified_at: self.tree_modified_at,
|
||||
created_by: self.created_by,
|
||||
updated_by: self.updated_by,
|
||||
}
|
||||
}
|
||||
|
||||
// Getters
|
||||
pub fn id(&self) -> &str {
|
||||
&self.id
|
||||
@@ -326,8 +370,25 @@ impl Folder {
|
||||
/// changed; the folder's own value stays untouched
|
||||
/// (self-exclusion).
|
||||
pub fn compute_etag(id: &str, tree_modified_at: u64) -> String {
|
||||
let prefix: String = id.chars().take(16).collect();
|
||||
format!("{}-{}", prefix, tree_modified_at)
|
||||
use std::fmt::Write as _;
|
||||
|
||||
// Byte index just past the 16th char (whole string when shorter).
|
||||
// `id` is a UUID string (ASCII) in practice, so this is
|
||||
// effectively `min(len, 16)`, but `char_indices` keeps the slice
|
||||
// char-boundary-safe for exotic fixture values — byte-identical
|
||||
// to the old `chars().take(16).collect::<String>()` without the
|
||||
// intermediate allocation.
|
||||
let end = match id.char_indices().nth(16) {
|
||||
Some((i, _)) => i,
|
||||
None => id.len(),
|
||||
};
|
||||
|
||||
// Single allocation: prefix + '-' + up to 20 digits (u64::MAX).
|
||||
let mut etag = String::with_capacity(end + 1 + 20);
|
||||
etag.push_str(&id[..end]);
|
||||
etag.push('-');
|
||||
let _ = write!(etag, "{tree_modified_at}");
|
||||
etag
|
||||
}
|
||||
|
||||
/// Creates a new Folder instance from a DTO
|
||||
|
||||
@@ -98,6 +98,32 @@ pub trait FolderRepository: Send + Sync + 'static {
|
||||
include_total: bool,
|
||||
) -> Result<(Vec<Folder>, Option<usize>), DomainError>;
|
||||
|
||||
/// Keyset-paged listing of `parent_id`'s direct sub-folders in name
|
||||
/// order — `name > $after_name ORDER BY name LIMIT $limit`, one bounded
|
||||
/// index-range read per page off the partial unique index
|
||||
/// `idx_folders_unique_name`. Streaming PROPFIND drains sub-folders
|
||||
/// with this instead of `COUNT(*) OVER() … LIMIT/OFFSET`, which
|
||||
/// window-aggregated and rescanned all N sub-folders on every page
|
||||
/// (4.5x on a 5k-dir parent, benches/FOLDER-KEYSET.md). `has_next`
|
||||
/// falls out of `rows.len() == limit` — no total needed.
|
||||
///
|
||||
/// The default implementation falls back to `list_folders` + in-memory
|
||||
/// slice so stubs and mocks compile without changes.
|
||||
async fn list_folders_batch(
|
||||
&self,
|
||||
parent_id: Option<&str>,
|
||||
after_name: Option<&str>,
|
||||
limit: usize,
|
||||
) -> Result<Vec<Folder>, DomainError> {
|
||||
let mut all = self.list_folders(parent_id).await?;
|
||||
all.sort_by(|a, b| a.name().cmp(b.name()));
|
||||
Ok(all
|
||||
.into_iter()
|
||||
.filter(|f| after_name.is_none_or(|a| f.name() > a))
|
||||
.take(limit)
|
||||
.collect())
|
||||
}
|
||||
|
||||
/// Renames a folder. `caller_id` is stamped into `updated_by`
|
||||
/// alongside the `updated_at = NOW()` bump (§14 provenance).
|
||||
async fn rename_folder(
|
||||
|
||||
Reference in New Issue
Block a user