perf: keyset/LATERAL SQL shapes, auth+blob-cache single-flight, spool buffers, DTO interning

Round 3 of benchmark-gated optimizations (benches/ROUND3.md; every change
gated by a before/after benchmark — an AFTER that did not beat its BEFORE
was to be rolled back; none needed it. Equivalence gates assert identical
row sequences / byte-identical output on every behavior-preserving rewrite):

DB hot paths (local PG16, EXPLAIN-verified):
- Web-UI listing (list_resources_paged): cursor pushed INSIDE the
  folders/files UNION-ALL branches as sargable row-value comparisons with
  per-branch ORDER/LIMIT + two partial expression indexes
  (folder_id, LOWER(name), id). 20k-entry folder: 26.6 -> 1.3 ms/page
  (19.5x); other sort modes at parity or better. New migration
  20260918000000. [benches/LISTING-KEYSET.md section in ROUND3]
- Photos timeline (list_media_files): per-drive CROSS JOIN LATERAL top-N
  on the timeline index, joins moved above the top-N. 50k-photo library:
  97.4 -> 1.6 ms/page (55.7x). The old "LIMIT stops the scan early"
  comment was refuted by EXPLAIN.
- PROPFIND sub-folders (both DAV surfaces): keyset list_folders_batch off
  idx_folders_unique_name replaces COUNT(*) OVER() + LIMIT/OFFSET
  (5k dirs: 79.7 -> 17.9 ms full walk, 4.5x).

Concurrency:
- Basic-auth cache single-flight (moka try_get_with): 8 concurrent DAV
  connections at TTL expiry paid 8 Argon2id runs (2.6 s CPU + 8x64 MiB);
  now 1 (300 ms). Failed verifications remain uncached.
- CachedBlobBackend per-hash single-flight + unique tmp names: 16
  concurrent cold readers = 16 full remote downloads racing truncating
  writes on ONE deterministic .tmp (corruptible cache); now 1 download
  (16x less egress, 2.8x wall on a shared link) and torn files can never
  be renamed into the cache.

I/O and allocations:
- Chunk-assembly reads 64K -> 512K buffers (2.3x, 8x fewer syscalls);
  chunk-spool writes via BufWriter 512K (5.6x, 32x fewer syscalls).
- S3/Azure put_blob_from_bytes_unsynced overrides: dedup settle no longer
  pays a HEAD probe per new chunk (2 RTT -> 1, 1.8x); Azure stops copying
  every chunk (Bytes -> Body, -0.44 ms - 4 MiB alloc per 4 MiB chunk).
- Entity->DTO mapping: Arc<str> interning of closed-set display fields +
  common MIMEs, 1-alloc etag/size formatting, FolderDto moves instead of
  clones. File row: 11 -> 4 allocs; folder row: 11.8 -> 1 (2.1x faster).
- CardDAV REPORT: deleted dead per-contact vCard pre-generation and the
  O(N^2) uid scan whose result was discarded (5k contacts: 55.7 -> 5.7 ms,
  9.8x); byte-identical XML asserted.
- Search-results cache: byte weigher + 32 MiB budget
  (OXICLOUD_SEARCH_CACHE_MAX_BYTES) replaces the 1000-ENTRY cap that let
  ~300 MiB of enriched rows sit in RSS; read latency parity.
- Dropped aws-config + aws-smithy-types (zero references; -82 dep-graph
  nodes, three SDK stacks gone from every build). tokio "process" is now
  an explicit feature (was enabled transitively by aws-config).

Frontend:
- Cached Intl.DateTimeFormat keyed by (locale, options) in formatDate and
  4 sibling callsites: 20k dates 2612 -> 51 ms (51.6x); vitest gate
  asserts output identity across locales and a 3x floor.

Validation: cargo fmt + clippy --all-features --all-targets -D warnings
clean; 518 unit + 548 integration-cfg tests green; new-shape endpoints
smoke-tested end-to-end over HTTP (all 5 listing sort modes with cursor
walks, WebDAV PROPFIND Depth-1, photos timeline, Basic-auth DAV login);
frontend npm run check clean, new vitest gates green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBsU2qEzny3A8WQUEuMNCr
This commit is contained in:
Claude
2026-07-17 11:10:27 +00:00
parent 7d95a19907
commit cd4c62042a
43 changed files with 5290 additions and 439 deletions
+32 -36
View File
@@ -651,7 +651,6 @@ impl CardDavAdapter {
pub fn generate_contacts_response<W: Write>(
writer: W,
contacts: &[ContactDto],
vcards: &[(String, String)], // (uid, vcard_data)
report: &CardDavReportType,
base_href: &str,
) -> Result<()> {
@@ -672,17 +671,9 @@ impl CardDavAdapter {
for contact in contacts {
let href = format!("{}{}.vcf", base_href, contact.uid);
let vcard = vcards
.iter()
.find(|(uid, _)| *uid == contact.uid)
.map(|(_, data)| data.as_str())
.unwrap_or("");
// `write_contact_response` generates the vCard on demand when (and
// only when) address-data is actually requested.
Self::write_contact_response(&mut xml_writer, contact, &props, &href)?;
// If address-data is requested, include vcard
if props.iter().any(|p| p.name == "address-data") || props.is_empty() {
// Already handled in write_contact_response
}
let _ = vcard; // suppress warning - used via contact_to_vcard fallback
}
xml_writer.write_event(Event::End(BytesEnd::new("D:multistatus")))?;
@@ -868,20 +859,25 @@ impl CardDavAdapter {
/// Convert a ContactDto to vCard 3.0 format
pub fn contact_to_vcard(contact: &ContactDto) -> String {
// `write!` into a String is infallible; `let _ =` discards the Ok(()).
// Formatting straight into the buffer avoids one temporary String per
// vCard line compared to `push_str(&format!(…))`.
use std::fmt::Write as _;
let mut vcard = String::from("BEGIN:VCARD\r\nVERSION:3.0\r\n");
vcard.push_str(&format!("UID:{}\r\n", contact.uid));
let _ = write!(vcard, "UID:{}\r\n", contact.uid);
if let (Some(last), Some(first)) = (&contact.last_name, &contact.first_name) {
vcard.push_str(&format!("N:{};{};;;\r\n", last, first));
let _ = write!(vcard, "N:{};{};;;\r\n", last, first);
} else if let Some(last) = &contact.last_name {
vcard.push_str(&format!("N:{};;;;\r\n", last));
let _ = write!(vcard, "N:{};;;;\r\n", last);
} else if let Some(first) = &contact.first_name {
vcard.push_str(&format!("N:;{};;;\r\n", first));
let _ = write!(vcard, "N:;{};;;\r\n", first);
}
if let Some(fn_name) = &contact.full_name {
vcard.push_str(&format!("FN:{}\r\n", fn_name));
let _ = write!(vcard, "FN:{}\r\n", fn_name);
} else {
// FN is mandatory in vCard 3.0
let fn_name = format!(
@@ -892,68 +888,68 @@ pub fn contact_to_vcard(contact: &ContactDto) -> String {
.trim()
.to_string();
if !fn_name.is_empty() {
vcard.push_str(&format!("FN:{}\r\n", fn_name));
let _ = write!(vcard, "FN:{}\r\n", fn_name);
} else {
vcard.push_str("FN:Unknown\r\n");
}
}
if let Some(nickname) = &contact.nickname {
vcard.push_str(&format!("NICKNAME:{}\r\n", nickname));
let _ = write!(vcard, "NICKNAME:{}\r\n", nickname);
}
for email in &contact.email {
vcard.push_str(&format!(
let _ = write!(
vcard,
"EMAIL;TYPE={}:{}\r\n",
email.r#type.to_uppercase(),
email.email
));
);
}
for phone in &contact.phone {
vcard.push_str(&format!(
let _ = write!(
vcard,
"TEL;TYPE={}:{}\r\n",
phone.r#type.to_uppercase(),
phone.number
));
);
}
for addr in &contact.address {
let adr = format!(
";;{};{};{};{};{}",
let _ = write!(
vcard,
"ADR;TYPE={}:;;{};{};{};{};{}\r\n",
addr.r#type.to_uppercase(),
addr.street.as_deref().unwrap_or(""),
addr.city.as_deref().unwrap_or(""),
addr.state.as_deref().unwrap_or(""),
addr.postal_code.as_deref().unwrap_or(""),
addr.country.as_deref().unwrap_or(""),
);
vcard.push_str(&format!(
"ADR;TYPE={}:{}\r\n",
addr.r#type.to_uppercase(),
adr
));
}
if let Some(org) = &contact.organization {
vcard.push_str(&format!("ORG:{}\r\n", org));
let _ = write!(vcard, "ORG:{}\r\n", org);
}
if let Some(title) = &contact.title {
vcard.push_str(&format!("TITLE:{}\r\n", title));
let _ = write!(vcard, "TITLE:{}\r\n", title);
}
if let Some(notes) = &contact.notes {
vcard.push_str(&format!("NOTE:{}\r\n", notes.replace('\n', "\\n")));
let _ = write!(vcard, "NOTE:{}\r\n", notes.replace('\n', "\\n"));
}
if let Some(bday) = &contact.birthday {
vcard.push_str(&format!("BDAY:{}\r\n", bday.format("%Y-%m-%d")));
let _ = write!(vcard, "BDAY:{}\r\n", bday.format("%Y-%m-%d"));
}
if let Some(photo) = &contact.photo_url {
vcard.push_str(&format!("PHOTO;VALUE=URI:{}\r\n", photo));
let _ = write!(vcard, "PHOTO;VALUE=URI:{}\r\n", photo);
}
vcard.push_str(&format!(
let _ = write!(
vcard,
"REV:{}\r\n",
contact.updated_at.format("%Y%m%dT%H%M%SZ")
));
);
vcard.push_str("END:VCARD\r\n");
vcard
@@ -484,10 +484,6 @@ mod tests {
#[test]
fn test_generate_contacts_response() {
let contacts = vec![sample_contact()];
let vcards = vec![(
"contact-001".to_string(),
contact_to_vcard(&sample_contact()),
)];
let report = CardDavReportType::AddressbookQuery {
props: vec![
QualifiedName {
@@ -505,7 +501,6 @@ mod tests {
let result = CardDavAdapter::generate_contacts_response(
&mut output,
&contacts,
&vcards,
&report,
"/carddav/ab-001",
);
@@ -528,14 +523,12 @@ mod tests {
#[test]
fn test_generate_empty_contacts_response() {
let contacts: Vec<ContactDto> = vec![];
let vcards: Vec<(String, String)> = vec![];
let report = CardDavReportType::AddressbookQuery { props: vec![] };
let mut output = Vec::new();
let result = CardDavAdapter::generate_contacts_response(
&mut output,
&contacts,
&vcards,
&report,
"/carddav/ab-001",
);
+228 -5
View File
@@ -8,6 +8,175 @@
//! then fall back to the file extension when the MIME is generic
//! (`application/octet-stream` or empty).
use std::collections::HashMap;
use std::fmt::Write as _;
use std::sync::{Arc, LazyLock};
// ─── Arc<str> interning for closed-set display values ────────────────
//
// `FileDto` / `FolderDto` store their display fields as `Arc<str>` so DTO
// clones are O(1). But `Arc::<str>::from(&str)` always allocates + copies,
// so building the DTO paid 3-4 heap allocations per row even though the
// value space is a small closed set. Interning turns each conversion into
// a HashMap lookup + refcount bump.
/// Every `&'static str` that [`icon_class_for`], [`icon_special_class_for`]
/// and [`category_for`] can return, plus the folder-DTO constants.
///
/// Keep this table in sync when adding a value to those functions — a
/// missing entry is not a bug (callers fall back to `Arc::from`, same
/// bytes, one extra allocation), just a lost optimization.
static DISPLAY_INTERN: LazyLock<HashMap<&'static str, Arc<str>>> = LazyLock::new(|| {
const CLOSED_SET: &[&str] = &[
// icon_class_for
"fas fa-file-pdf",
"fas fa-file-word",
"fas fa-file-excel",
"fas fa-file-powerpoint",
"fas fa-file-archive",
"fas fa-file-code",
"fas fa-hdd",
"fas fa-file-image",
"fas fa-file-video",
"fas fa-file-audio",
"fas fa-file-alt",
"fas fa-terminal",
"fas fa-file",
// icon_special_class_for
"pdf-icon",
"doc-icon",
"spreadsheet-icon",
"presentation-icon",
"archive-icon",
"code-icon json-icon",
"code-icon js-icon",
"code-icon ts-icon",
"code-icon html-icon",
"code-icon sql-icon",
"code-icon config-icon",
"code-icon php-icon",
"script-icon",
"installer-icon",
"image-icon",
"video-icon",
"audio-icon",
"code-icon py-icon",
"code-icon rust-icon",
"code-icon",
"code-icon go-icon",
"code-icon ruby-icon",
"code-icon md-icon",
"code-icon css-icon",
"code-icon java-icon",
"code-icon c-icon",
"code-icon cs-icon",
"code-icon swift-icon",
"",
// category_for
"PDF",
"Document",
"Spreadsheet",
"Presentation",
"Archive",
"Code",
"Installer",
"Image",
"Video",
"Audio",
"Markdown",
"Text",
// FolderDto constants
"fas fa-folder",
"folder-icon",
"Folder",
];
CLOSED_SET.iter().map(|s| (*s, Arc::from(*s))).collect()
});
/// Returns a shared `Arc<str>` for a display value from the closed sets
/// above (icon class, icon special class, category). Lookup + refcount
/// bump instead of alloc + copy; unknown values (future additions not
/// yet in the table) fall back to `Arc::from` with identical bytes.
pub fn intern_display(s: &'static str) -> Arc<str> {
DISPLAY_INTERN
.get(s)
.cloned()
.unwrap_or_else(|| Arc::from(s))
}
/// The MIME types that dominate real storage rows. Exotic types fall back
/// to a per-row `Arc::from` — correctness is unaffected, only the alloc is.
static MIME_INTERN: LazyLock<HashMap<&'static str, Arc<str>>> = LazyLock::new(|| {
const COMMON_MIMES: &[&str] = &[
"",
"directory",
"application/octet-stream",
// Images
"image/jpeg",
"image/png",
"image/gif",
"image/webp",
"image/svg+xml",
"image/heic",
"image/heif",
"image/avif",
"image/bmp",
"image/tiff",
"image/x-icon",
// Video
"video/mp4",
"video/quicktime",
"video/webm",
"video/x-matroska",
"video/x-msvideo",
// Audio
"audio/mpeg",
"audio/mp4",
"audio/ogg",
"audio/flac",
"audio/wav",
"audio/x-wav",
"audio/aac",
// Documents
"application/pdf",
"application/msword",
"application/vnd.openxmlformats-officedocument.wordprocessingml.document",
"application/vnd.ms-excel",
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
"application/vnd.ms-powerpoint",
"application/vnd.openxmlformats-officedocument.presentationml.presentation",
"application/vnd.oasis.opendocument.text",
"application/vnd.oasis.opendocument.spreadsheet",
// Text / code
"text/plain",
"text/csv",
"text/html",
"text/css",
"text/markdown",
"text/xml",
"application/json",
"application/javascript",
"application/xml",
"application/x-yaml",
// Archives
"application/zip",
"application/gzip",
"application/x-tar",
"application/x-7z-compressed",
"application/x-rar-compressed",
];
COMMON_MIMES.iter().map(|s| (*s, Arc::from(*s))).collect()
});
/// Returns a shared `Arc<str>` for the given MIME type. Common types hit
/// the intern table (refcount bump); exotic ones allocate as before.
pub fn intern_mime(mime: &str) -> Arc<str> {
MIME_INTERN
.get(mime)
.cloned()
.unwrap_or_else(|| Arc::from(mime))
}
// ─── Private: extract lowercase extension from a filename ────────────
fn ext_of(name: &str) -> Option<&str> {
let name = name.rsplit('/').next().unwrap_or(name); // strip path
@@ -388,11 +557,21 @@ pub fn format_file_size(bytes: u64) -> String {
let value = bytes as f64 / K.powi(i as i32);
// Two decimal places, then strip trailing zeros (matches JS parseFloat behaviour)
let formatted = format!("{:.2}", value);
let formatted = formatted.trim_end_matches('0').trim_end_matches('.');
format!("{} {}", formatted, SIZES[i])
// Single buffer: write the 2-decimal value, strip trailing zeros in
// place (matches JS parseFloat behaviour), then append the unit.
// 16 chars covers the worst case ("16777216 TB" for u64::MAX,
// "1023.99 Bytes" for the longest unit), so no realloc occurs.
let mut out = String::with_capacity(16);
let _ = write!(out, "{:.2}", value);
while out.ends_with('0') {
out.pop();
}
if out.ends_with('.') {
out.pop();
}
out.push(' ');
out.push_str(SIZES[i]);
out
}
#[cfg(test)]
@@ -506,6 +685,50 @@ mod tests {
);
}
/// Every value the closed-set display functions can return must hit
/// the intern table (same bytes, shared allocation) — a miss is only
/// a lost optimization, but this test keeps the table in sync.
#[test]
fn test_intern_display_covers_closed_sets_and_shares_storage() {
for s in [
"fas fa-file-pdf",
"fas fa-file",
"fas fa-terminal",
"fas fa-folder",
"code-icon rust-icon",
"folder-icon",
"",
"PDF",
"Folder",
"Document",
"Markdown",
] {
let a = intern_display(s);
let b = intern_display(s);
assert_eq!(&*a, s, "interned bytes must be identical");
assert!(
Arc::ptr_eq(&a, &b),
"closed-set value {s:?} must come from the intern table"
);
}
}
#[test]
fn test_intern_mime_common_hits_table_exotic_falls_back() {
let a = intern_mime("image/jpeg");
let b = intern_mime("image/jpeg");
assert_eq!(&*a, "image/jpeg");
assert!(Arc::ptr_eq(&a, &b), "common MIME must be interned");
let exotic = intern_mime("chemical/x-pdb");
assert_eq!(&*exotic, "chemical/x-pdb");
let exotic2 = intern_mime("chemical/x-pdb");
assert!(
!Arc::ptr_eq(&exotic, &exotic2),
"exotic MIME falls back to a fresh Arc"
);
}
#[test]
fn test_ext_of() {
assert_eq!(ext_of("file.txt"), Some("txt"));
+14 -9
View File
@@ -6,7 +6,8 @@ use utoipa::ToSchema;
use uuid::Uuid;
use super::display_helpers::{
category_for, format_file_size, icon_class_for, icon_special_class_for,
category_for, format_file_size, icon_class_for, icon_special_class_for, intern_display,
intern_mime,
};
/// DTO for file responses
@@ -101,11 +102,15 @@ impl From<File> for FileDto {
// for id, name, path, folder_id (previously 4× .to_string()).
let parts = file.into_parts();
let icon_class = Arc::from(icon_class_for(&parts.name, &parts.mime_type));
let icon_special_class = Arc::from(icon_special_class_for(&parts.name, &parts.mime_type));
let category = Arc::from(category_for(&parts.name, &parts.mime_type));
// Display fields come from closed static tables and MIME values
// repeat massively across rows — intern instead of allocating a
// fresh Arc<str> per row (`Arc::from(&str)` always allocs+copies).
let icon_class = intern_display(icon_class_for(&parts.name, &parts.mime_type));
let icon_special_class =
intern_display(icon_special_class_for(&parts.name, &parts.mime_type));
let category = intern_display(category_for(&parts.name, &parts.mime_type));
let size_formatted = format_file_size(parts.size);
let mime_type = Arc::from(parts.mime_type.as_str());
let mime_type = intern_mime(&parts.mime_type);
Self {
id: parts.id,
@@ -169,13 +174,13 @@ impl FileDto {
name: "stub-file".to_string(),
path: "/stub/path".to_string(),
size: 0,
mime_type: Arc::from("application/octet-stream"),
mime_type: intern_mime("application/octet-stream"),
folder_id: None,
created_at: 0,
modified_at: 0,
icon_class: Arc::from("fas fa-file"),
icon_special_class: Arc::from(""),
category: Arc::from("Document"),
icon_class: intern_display("fas fa-file"),
icon_special_class: intern_display(""),
category: intern_display("Document"),
size_formatted: "0 Bytes".to_string(),
content_hash: String::new(),
etag: String::new(),
+27 -17
View File
@@ -1,6 +1,7 @@
use std::sync::Arc;
use crate::application::dtos::cursor::{CursorListResponse, CursorQuery, PageCursor};
use crate::application::dtos::display_helpers::intern_display;
use crate::application::dtos::grant_dto::{ResourceContentDto, ResourceTypeDto};
use crate::domain::entities::folder::Folder;
use crate::domain::services::authorization::ResourceKind;
@@ -99,24 +100,33 @@ pub struct FolderDto {
impl From<Folder> for FolderDto {
fn from(folder: Folder) -> Self {
let is_root = folder.parent_id().is_none();
let etag = folder.etag().to_string();
// Consume the entity by moving all fields — zero heap allocations
// for id, name, path, parent_id (previously 3-4× .to_string()).
let parts = folder.into_parts();
let is_root = parts.parent_id.is_none();
// Single-allocation ETag straight from the owned parts. The old
// shape (`folder.etag().to_string()`) built the String and then
// cloned it — a pure double-alloc.
let etag = Folder::compute_etag(&parts.id, parts.tree_modified_at);
Self {
id: folder.id().to_string(),
name: folder.name().to_string(),
path: folder.path_string().to_string(),
parent_id: folder.parent_id().map(String::from),
drive_id: folder.drive_id(),
created_at: folder.created_at(),
modified_at: folder.modified_at(),
id: parts.id,
name: parts.name,
path: parts.path_string,
parent_id: parts.parent_id,
drive_id: parts.drive_id,
created_at: parts.created_at,
modified_at: parts.modified_at,
is_root,
icon_class: Arc::from("fas fa-folder"),
icon_special_class: Arc::from("folder-icon"),
category: Arc::from("Folder"),
// Constant display fields: refcount bump on interned statics
// instead of 3 fresh Arc allocations per row.
icon_class: intern_display("fas fa-folder"),
icon_special_class: intern_display("folder-icon"),
category: intern_display("Folder"),
etag,
created_by: folder.created_by(),
updated_by: folder.updated_by(),
created_by: parts.created_by,
updated_by: parts.updated_by,
}
}
}
@@ -163,9 +173,9 @@ impl FolderDto {
created_at: 0,
modified_at: 0,
is_root: true,
icon_class: Arc::from("fas fa-folder"),
icon_special_class: Arc::from("folder-icon"),
category: Arc::from("Folder"),
icon_class: intern_display("fas fa-folder"),
icon_special_class: intern_display("folder-icon"),
category: intern_display("Folder"),
etag: String::new(),
created_by: None,
updated_by: None,
+25
View File
@@ -77,6 +77,31 @@ pub trait FolderUseCase: Send + Sync + 'static {
pagination: &crate::application::dtos::pagination::PaginationRequestDto,
) -> Result<crate::application::dtos::pagination::PaginatedResponseDto<FolderDto>, DomainError>;
/// Keyset-paged sub-folder listing in name order, scoped to a caller —
/// `name > after_name LIMIT limit`, `has_next = len() == limit`.
///
/// Used by streaming WebDAV/NC PROPFIND: O(page) per page off the
/// `idx_folders_unique_name` index instead of the quadratic
/// `COUNT(*) OVER() … LIMIT/OFFSET` walk (benches/FOLDER-KEYSET.md).
///
/// The default implementation falls back to `list_folders_with_perms`
/// + in-memory slice so stubs and mocks compile without changes.
async fn list_folders_batch_with_perms(
&self,
parent_id: Option<&str>,
caller_id: Uuid,
after_name: Option<&str>,
limit: usize,
) -> Result<Vec<FolderDto>, DomainError> {
let mut all = self.list_folders_with_perms(parent_id, caller_id).await?;
all.sort_by(|a, b| a.name.cmp(&b.name));
Ok(all
.into_iter()
.filter(|f| after_name.is_none_or(|a| f.name.as_str() > a))
.take(limit)
.collect())
}
/// Renames a folder (ownership verified against caller_id)
async fn rename_folder_with_perms(
&self,
@@ -304,12 +304,45 @@ impl AppPasswordService {
let cache_key: [u8; 32] =
blake3::hash(format!("{}:{}", username, password).as_bytes()).into();
// ── 2. Cache hit → return immediately ────────────────────────
if let Some(cached) = self.auth_cache.get(&cache_key).await {
return Ok((cached.user_id, cached.username, cached.email, cached.role));
}
// ── 2. Single-flight cache lookup ─────────────────────────────
// Concurrent misses on the same credential coalesce into ONE
// full verification: DAV sync clients hold 4-8 parallel
// connections, so an expiring cache entry used to fan out into
// K simultaneous Argon2id runs (~100-300 ms CPU + 64 MiB RAM
// apiece) every TTL — a recurring p99 spike on every DAV
// surface (8 -> 1 verifications, benches/AUTH-HERD.md).
// `try_get_with` caches only `Ok` results, so failed
// verifications are still never cached, preserving the full
// Argon2id cost as a brute-force deterrent.
let result = self
.auth_cache
.try_get_with(
cache_key,
self.verify_basic_auth_uncached(username, password),
)
.await
.map_err(
|e: std::sync::Arc<DomainError>| match std::sync::Arc::try_unwrap(e) {
Ok(err) => err,
// Another coalesced waiter still holds the Arc — rebuild
// an equivalent error (the source chain isn't clonable).
Err(shared) => {
DomainError::new(shared.kind, shared.entity_type, shared.message.clone())
}
},
)?;
Ok((result.user_id, result.username, result.email, result.role))
}
// ── 3. Cache miss → full verification ────────────────────────
/// The uncached Basic Auth slow path: user lookup, prefix-scoped
/// candidate fetch, Argon2id verification. Runs at most once per
/// credential per TTL — `verify_basic_auth` coalesces concurrent
/// callers onto a single in-flight instance of this future.
async fn verify_basic_auth_uncached(
&self,
username: &str,
password: &str,
) -> Result<CachedBasicAuthResult, DomainError> {
let user = self
.user_repo
.get_user_by_username(username)
@@ -363,15 +396,14 @@ impl AppPasswordService {
{
let _ = self.repo.touch_last_used(ap.id).await;
let result = CachedBasicAuthResult {
// Caching happens in `verify_basic_auth`: `try_get_with`
// stores this value under the blake3 key on return.
return Ok(CachedBasicAuthResult {
user_id: user.id(),
username: user.username().unwrap_or("").to_string(),
email: user.email().to_string(),
role: user.role().to_string(),
};
self.auth_cache.insert(cache_key, result.clone()).await;
return Ok((result.user_id, result.username, result.email, result.role));
});
}
}
@@ -429,6 +429,62 @@ impl FolderUseCase for FolderService {
Ok(response)
}
/// Keyset-paged sub-folder listing (name order), caller-scoped.
///
/// AuthZ mirrors `list_folders_paginated_with_perms`: one
/// `authz.require(Read)` on the parent per batch; root scope goes
/// through the caller's drive-membership listing.
async fn list_folders_batch_with_perms(
&self,
parent_id: Option<&str>,
caller_id: Uuid,
after_name: Option<&str>,
limit: usize,
) -> Result<Vec<FolderDto>, DomainError> {
match parent_id {
Some(pid) => {
self.authz
.require(
Subject::User(caller_id),
Permission::Read,
Self::folder_resource(pid)?,
)
.await?;
let folders = self
.folder_storage
.list_folders_batch(parent_id, after_name, limit)
.await
.map_err(|e| {
DomainError::internal_error(
"FolderStorage",
format!("Failed to batch-list folders in parent {pid}: {e}"),
)
})?;
Ok(folders.into_iter().map(FolderDto::from).collect())
}
None => {
// Root scope: one row per readable drive — a handful.
let mut all = self
.folder_storage
.list_root_folders_for_caller(caller_id)
.await
.map_err(|e| {
DomainError::internal_error(
"FolderStorage",
format!("Failed to batch-list root folders for '{caller_id}': {e}"),
)
})?;
all.sort_by(|a, b| a.name().cmp(b.name()));
Ok(all
.into_iter()
.filter(|f| after_name.is_none_or(|a| f.name() > a))
.take(limit)
.map(FolderDto::from)
.collect())
}
}
}
/// Lists folders with pagination, scoped to a specific owner.
async fn list_folders_paginated_with_perms(
&self,
+159 -5
View File
@@ -67,9 +67,80 @@ pub struct SearchService {
/// Lock-free concurrent cache with automatic TTL and LRU eviction (moka).
/// Values are `Arc<SearchResultsDto>` so cache insert/hit is a single
/// atomic ref-count increment (~1 ns) instead of cloning thousands of Strings.
///
/// **Byte-bounded**, not entry-bounded: entries are weighed by
/// [`search_results_entry_weight`] and `max_capacity` is a byte budget.
/// Keys span user × query × offset × limit, and each page holds up to 500
/// enriched rows (~500–900 B of owned Strings each) — an entry-count bound
/// let hundreds of MB of result pages accumulate invisibly.
search_cache: moka::future::Cache<u64, Arc<SearchResultsDto>>,
}
// ─── Search-results cache (byte-bounded) ─────────────────────────────────
/// Approximate heap bytes retained by one cached search page.
///
/// With a `weigher` installed, moka's `max_capacity` is the sum of entry
/// *weights*, so this converts the cache bound from "number of entries" to
/// real bytes: the length of every owned `String` in each file/folder row,
/// plus a fixed per-row and per-entry overhead for struct fields, the 24-B
/// `String` headers, `Vec` slots and allocator slop. Same pattern as the
/// file-content cache and the dedup manifest cache.
///
/// `pub` so `examples/bench_search_cache_mem.rs` can recompute retained
/// bytes with the exact production formula.
pub fn search_results_entry_weight(_key: &u64, value: &Arc<SearchResultsDto>) -> u32 {
/// Fixed per-row overhead: struct scalars + one 24-B header per `String`
/// field (12 on a file row, 4 on a folder row) + `Vec` slot + allocator
/// slop. Deliberately a round upper-ish estimate — under-weighing is the
/// failure mode that re-opens the memory hole.
const ROW_OVERHEAD: usize = 200;
/// Fixed per-entry overhead: `Arc` + `SearchResultsDto` scalars + `Vec`
/// headers + moka's own bookkeeping per entry.
const ENTRY_OVERHEAD: usize = 256;
fn opt_len(s: &Option<String>) -> usize {
s.as_deref().map_or(0, str::len)
}
let mut bytes = ENTRY_OVERHEAD + value.sort_by.len();
for f in &value.files {
bytes += ROW_OVERHEAD
+ f.id.len()
+ f.name.len()
+ f.path.len()
+ f.mime_type.len()
+ opt_len(&f.folder_id)
+ f.size_formatted.len()
+ f.icon_class.len()
+ f.icon_special_class.len()
+ f.category.len()
+ f.blob_hash.len()
+ opt_len(&f.snippet)
+ opt_len(&f.match_source);
}
for d in &value.folders {
bytes += ROW_OVERHEAD + d.id.len() + d.name.len() + d.path.len() + opt_len(&d.parent_id);
}
bytes.min(u32::MAX as usize) as u32
}
/// Build the search-results cache exactly as production wires it: a byte
/// budget enforced through [`search_results_entry_weight`], plus TTL.
///
/// Shared with `examples/bench_search_cache_mem.rs` so the benchmark
/// measures the identical cache configuration that serves requests.
pub fn build_search_results_cache(
cache_ttl_secs: u64,
max_bytes: u64,
) -> moka::future::Cache<u64, Arc<SearchResultsDto>> {
moka::future::Cache::builder()
.max_capacity(max_bytes)
.weigher(search_results_entry_weight)
.time_to_live(Duration::from_secs(cache_ttl_secs))
.build()
}
// ─── Utility functions (pure, no self — computed on the server) ─────────
/// Compute relevance score (0–100) for a name against a query.
@@ -160,6 +231,10 @@ fn get_category(name: &str, mime: &str) -> String {
impl SearchService {
/**
* Creates a new instance of the search service.
*
* `max_cache_bytes` is the byte budget for the results cache (weigher-
* bounded, see [`search_results_entry_weight`]) — it replaced the old
* entry-count capacity, which was blind to how big each cached page is.
*/
pub fn new(
file_repository: Arc<FileBlobReadRepository>,
@@ -168,12 +243,9 @@ impl SearchService {
authorization: Option<Arc<crate::infrastructure::services::pg_acl_engine::PgAclEngine>>,
drive_repo: Option<Arc<dyn crate::domain::repositories::drive_repository::DriveRepository>>,
cache_ttl: u64,
max_cache_size: usize,
max_cache_bytes: u64,
) -> Self {
let search_cache = moka::future::Cache::builder()
.max_capacity(max_cache_size as u64)
.time_to_live(Duration::from_secs(cache_ttl))
.build();
let search_cache = build_search_results_cache(cache_ttl, max_cache_bytes);
Self {
file_repository,
@@ -815,6 +887,88 @@ mod tests {
}
}
#[test]
fn entry_weight_counts_every_owned_string_plus_overheads() {
// Empty page: entry overhead + sort_by ("relevance" = 9 bytes).
let empty = Arc::new(SearchResultsDto::empty());
let base = search_results_entry_weight(&0, &empty) as usize;
assert_eq!(base, 256 + 9);
// One file row: base + row overhead + its owned string bytes
// (id 7 + name 7 + path 8 + mime 10; the rest are empty/None).
let one_file = Arc::new(SearchResultsDto::new(
vec![dto("abc.txt", 50, 10, 1)],
Vec::new(),
100,
0,
Some(1),
0,
"relevance".to_string(),
));
let w = search_results_entry_weight(&0, &one_file) as usize;
assert_eq!(w, base + 200 + 7 + 7 + 8 + 10);
// Folder rows weigh too (id 2 + name 4 + path 5 + parent 6 = 17).
let one_folder = Arc::new(SearchResultsDto::new(
Vec::new(),
vec![SearchFolderResultDto {
id: "f1".to_string(),
name: "docs".to_string(),
path: "/docs".to_string(),
parent_id: Some("parent".to_string()),
drive_id: Uuid::nil(),
created_at: 0,
modified_at: 0,
is_root: false,
relevance_score: 50,
}],
100,
0,
Some(1),
0,
"relevance".to_string(),
));
let w = search_results_entry_weight(&0, &one_folder) as usize;
assert_eq!(w, base + 200 + 2 + 4 + 5 + 6);
}
#[tokio::test]
async fn cache_evicts_down_to_the_byte_budget() {
// Budget fits ~2 of these entries; inserting 20 must never let the
// weighted size settle above the budget.
let entry = |i: usize| {
Arc::new(SearchResultsDto::new(
(0..50)
.map(|r| dto(&format!("file_{i}_{r}_{}", "x".repeat(100)), 50, 1, 1))
.collect(),
Vec::new(),
50,
0,
Some(50),
0,
"relevance".to_string(),
))
};
let per_entry = search_results_entry_weight(&0, &entry(0)) as u64;
let budget = per_entry * 2 + per_entry / 2;
let cache = build_search_results_cache(300, budget);
for i in 0..20u64 {
cache.insert(i, entry(i as usize)).await;
}
cache.run_pending_tasks().await;
let retained: u64 = cache
.iter()
.map(|(k, v)| search_results_entry_weight(&k, &v) as u64)
.sum();
assert!(
retained <= budget,
"retained {retained} B exceeds budget {budget} B"
);
assert!(cache.entry_count() <= 2);
}
#[test]
fn merged_files_resort_by_relevance_and_by_column() {
let mut files = vec![