Files
Oxicloud/benches/ROUND3.md
T
Claude cd4c62042a perf: keyset/LATERAL SQL shapes, auth+blob-cache single-flight, spool buffers, DTO interning
Round 3 of benchmark-gated optimizations (benches/ROUND3.md; every change
gated by a before/after benchmark — an AFTER that did not beat its BEFORE
was to be rolled back; none needed it. Equivalence gates assert identical
row sequences / byte-identical output on every behavior-preserving rewrite):

DB hot paths (local PG16, EXPLAIN-verified):
- Web-UI listing (list_resources_paged): cursor pushed INSIDE the
  folders/files UNION-ALL branches as sargable row-value comparisons with
  per-branch ORDER/LIMIT + two partial expression indexes
  (folder_id, LOWER(name), id). 20k-entry folder: 26.6 -> 1.3 ms/page
  (19.5x); other sort modes at parity or better. New migration
  20260918000000. [benches/LISTING-KEYSET.md section in ROUND3]
- Photos timeline (list_media_files): per-drive CROSS JOIN LATERAL top-N
  on the timeline index, joins moved above the top-N. 50k-photo library:
  97.4 -> 1.6 ms/page (55.7x). The old "LIMIT stops the scan early"
  comment was refuted by EXPLAIN.
- PROPFIND sub-folders (both DAV surfaces): keyset list_folders_batch off
  idx_folders_unique_name replaces COUNT(*) OVER() + LIMIT/OFFSET
  (5k dirs: 79.7 -> 17.9 ms full walk, 4.5x).

Concurrency:
- Basic-auth cache single-flight (moka try_get_with): 8 concurrent DAV
  connections at TTL expiry paid 8 Argon2id runs (2.6 s CPU + 8x64 MiB);
  now 1 (300 ms). Failed verifications remain uncached.
- CachedBlobBackend per-hash single-flight + unique tmp names: 16
  concurrent cold readers = 16 full remote downloads racing truncating
  writes on ONE deterministic .tmp (corruptible cache); now 1 download
  (16x less egress, 2.8x wall on a shared link) and torn files can never
  be renamed into the cache.

I/O and allocations:
- Chunk-assembly reads 64K -> 512K buffers (2.3x, 8x fewer syscalls);
  chunk-spool writes via BufWriter 512K (5.6x, 32x fewer syscalls).
- S3/Azure put_blob_from_bytes_unsynced overrides: dedup settle no longer
  pays a HEAD probe per new chunk (2 RTT -> 1, 1.8x); Azure stops copying
  every chunk (Bytes -> Body, -0.44 ms - 4 MiB alloc per 4 MiB chunk).
- Entity->DTO mapping: Arc<str> interning of closed-set display fields +
  common MIMEs, 1-alloc etag/size formatting, FolderDto moves instead of
  clones. File row: 11 -> 4 allocs; folder row: 11.8 -> 1 (2.1x faster).
- CardDAV REPORT: deleted dead per-contact vCard pre-generation and the
  O(N^2) uid scan whose result was discarded (5k contacts: 55.7 -> 5.7 ms,
  9.8x); byte-identical XML asserted.
- Search-results cache: byte weigher + 32 MiB budget
  (OXICLOUD_SEARCH_CACHE_MAX_BYTES) replaces the 1000-ENTRY cap that let
  ~300 MiB of enriched rows sit in RSS; read latency parity.
- Dropped aws-config + aws-smithy-types (zero references; -82 dep-graph
  nodes, three SDK stacks gone from every build). tokio "process" is now
  an explicit feature (was enabled transitively by aws-config).

Frontend:
- Cached Intl.DateTimeFormat keyed by (locale, options) in formatDate and
  4 sibling callsites: 20k dates 2612 -> 51 ms (51.6x); vitest gate
  asserts output identity across locales and a 3x floor.

Validation: cargo fmt + clippy --all-features --all-targets -D warnings
clean; 518 unit + 548 integration-cfg tests green; new-shape endpoints
smoke-tested end-to-end over HTTP (all 5 listing sort modes with cursor
walks, WebDAV PROPFIND Depth-1, photos timeline, Basic-auth DAV login);
frontend npm run check clean, new vitest gates green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EBsU2qEzny3A8WQUEuMNCr
2026-07-17 11:10:27 +00:00

13 KiB

Round 3 — listing/timeline SQL shapes, auth herd, blob-cache stampede, spool I/O, DTO allocs

Twelve benchmark-gated changes. Rule of the round (same as ROUND2): every change ships with a BEFORE/AFTER benchmark; an AFTER that doesn't beat its BEFORE gets rolled back — none did. Equivalence gates (byte-identical output / identical row sequences) guard every behavior-preserving rewrite.

Measured on 4 cores / 15 GiB, local PostgreSQL 16 (fsync off), release profile. Reproduce any row with the command in its section.

Summary

# change key metric before → after
1 Web-UI listing keyset pushdown ms/page p50, 20k-entry folder 26.6 → 1.30 (19.5x)
2 Photos timeline LATERAL top-N ms/page p50, 50k-photo library 97.4 → 1.61 (55.7x)
3 PROPFIND subfolder keyset full walk, 5k dirs 79.7 → 17.9 ms (4.5x)
4 Basic-auth single-flight herd CPU, 8 conns 2620 → 300 ms (8.7x)
5 Blob-cache miss single-flight remote fetches / wall 16 → 1, 519 → 188 ms (2.8x)
6 Chunk-assembly read buffer 512K wall / read syscalls 251 → 109 ms (2.3x), 2580 → 340
7 Chunk-spool BufWriter 512K wall / write syscalls 877 → 158 ms (5.6x), 12800 → 400
8 S3/Azure unsynced PUT (no HEAD) wall / requests, 500 chunks 1604 → 868 ms (1.8x), 1000 → 500
9 DTO mapping interning allocs/row file / folder 11.0 → 4.0, 11.8 → 1.0
10 CardDAV REPORT dead work 5k contacts, getetag 55.7 → 5.7 ms (9.8x)
11 Search-cache byte weigher retained RSS worst case ~298 MiB → 31.9 MiB (bounded)
12 Drop aws-config/aws-smithy-types dep-graph nodes 1728 → 1646

Frontend (gated by vitest, frontend/src/lib/utils/formatDate.bench.test.ts): cached Intl.DateTimeFormat — 20k dates 2612 → 50.6 ms (51.6x), output identity asserted across locales.


[1] Web-UI folder listing — whole-folder rescan → per-branch keyset — 19.5x

list_resources_paged (SPA files view) applied its keyset cursor OUTSIDE the folders/files UNION-ALL on computed columns (sort_str = LOWER(name), folder_first), so Postgres re-scanned and top-N-sorted every remaining row of the folder on every page (EXPLAIN: Seq Scan, 17,999 rows removed by filter, 29 ms / 565 buffers per 200-row page on a 20k-file folder).

Now the cursor is pushed into each branch as a sargable row-value comparison on base columns ((LOWER(name), id) > ($str, $id)), constants folded per branch in Rust (a cursor in the file group drops the folder branch outright), each branch pre-sorts + pre-limits, and the outer query merges ≤ 2·limit rows. Two new expression indexes (migration 20260918000000): idx_files_folder_lname (folder_id, LOWER(name), id) and idx_folders_parent_lname (parent_id, LOWER(name), id), both partial on NOT is_trashed.

cargo run --release --features bench --example bench_listing_keyset
# full drain, 20k files + 300 dirs, 200/page       total ms   p50/pg   p99/pg
# name        OLD/no-idx                             2717.2    26.57    33.55
# name        OLD/idx (indexes alone don't help)     2786.8    27.83    35.62
# name        NEW/idx                                 139.6     1.30     1.81   19.5x
# modified_at OLD → NEW (no dedicated index)         1653.4 → 1367.5             1.2x

Equivalence: the drained (type, id) sequence is asserted identical across all modes and both sort orders; the example exits 1 on mismatch.

[2] Photos timeline — full-library scan → per-drive LATERAL top-N — 55.7x

list_media_files claimed idx_files_media_timeline_by_drive let LIMIT stop the scan early; EXPLAIN refuted it — the folders/file_metadata joins and the global sort sat ABOVE the drive_id IN (grants) nested loop, so every page fed the ENTIRE media library through the join into a top-N heapsort. Now the accessible drive ids materialise once, a CROSS JOIN LATERAL (… ORDER BY media_sort_date DESC LIMIT k) per drive does one bounded index scan each, and the joins run on the k emitted rows only.

cargo run --release --features bench --example bench_photos_timeline
# 10 pages of 100, 50k photos, 3 drives      total ms   p50 ms/page
# OLD                                          1032.1        97.41
# NEW                                            18.5         1.61   55.7x

Equivalence: page-by-page id sequences asserted identical (seed uses strictly distinct capture dates so ties can't mask reordering).

[3] PROPFIND subfolder paging — LIMIT/OFFSET + COUNT(*) OVER() → keyset — 4.5x

The exact quadratic shape PROPFIND-PAGING fixed for files still applied to sub-folders on both DAV surfaces: every page window-aggregated and re-scanned all N sub-folders, and the total was only used for has_next. New FolderRepository::list_folders_batch (keyset name > $last, served by the existing idx_folders_unique_name, no migration) wired into both streaming PROPFIND walkers via list_folders_batch_with_perms (same per-batch authz as before).

cargo run --release --features bench --example bench_folder_keyset
# full walk, 5k dirs, 500/page     total ms   p50 ms/page
# OFFSET                               79.7          6.54
# KEYSET                               17.9          1.64   4.5x

[4] Basic-auth cache — thundering herd → single-flight — 8.7x CPU

Every DAV/NC request authenticates via verify_basic_auth. On a cache miss each concurrent caller independently ran the full slow path — an Argon2id verification (m=64 MiB, t=3, p=2 ≈ 290 ms CPU here) apiece. DAV sync clients hold 4-8 parallel connections, so every TTL expiry (300 s) fanned out K verifications: a recurring p99 spike + CPU/RAM burst. try_get_with now coalesces concurrent misses; errors are never cached (brute-force cost preserved), revocation via invalidate_entries_if unchanged.

cargo run --release --features bench --example bench_auth_herd
# herd of 8, cold cache          wall ms   CPU ms   verifications
# BEFORE (per-caller)                764     2620             9.0
# AFTER  (single-flight)             311      300             1.0
# warm hit p50: 0.6 us

[5] CachedBlobBackend — miss stampede → per-hash single-flight — 16 fetches → 1

K concurrent cold readers of one blob (video player's parallel Range probes; N clients pulling the same new file) each downloaded the FULL blob from S3/Azure — and raced truncating writes on ONE deterministic .tmp path (a torn interleaving could be renamed into the cache). Fixes: a per-hash DashMap gate (leader fetches, waiters re-check and serve locally), plus unique .{uuid}.tmp names + error-path cleanup so a corrupt file can never land at the final path.

cargo run --release --features bench --example bench_blob_cache
# 16 cold readers, 32 MiB blob, shared 1 GiB/s link   wall ms   fetches   remote MiB
# BEFORE (per-caller)                                     519        16          512
# AFTER  (single-flight)                                  188         1           32
# gates: fetch count == 1; BLAKE3 of served + durable cache file == source

[6][7] Upload spool I/O — 64 KiB reads, unbuffered frame writes

Assembly read (stream_from_files, the single read pass over every completed chunked upload) used 64 KiB ReaderStream polls — one blocking-pool dispatch + read(2) each — while every other blob path uses 256 KiB+. Capacity sweep picked 512 KiB. Chunk-spool writes (stream_body_to_path, every chunk PUT on both surfaces) went straight to a bare tokio File — one dispatch + write(2) per ~16-64 KiB HTTP frame; now wrapped in BufWriter::with_capacity(512 KiB) like the dedup handler's spool loop.

cargo run --release --features bench --example bench_upload_spool
# [1] read 16 x 10 MiB parts    wall ms   read syscalls
#   64K  (BEFORE)                 250.8            2580
#   256K                          125.1             660
#   512K (AFTER)                  108.8             340   2.3x
#   1M                            111.3             180
# [2] spool 640 x 16 KiB frames x 20 files
#   bare File (BEFORE)            877.4    12800 syscw
#   BufWriter 512K (AFTER)        157.9      400 syscw    5.6x

[8] S3/Azure chunk writes — HEAD-before-PUT → unconditional PUT — 1.8x

Neither remote backend overrode put_blob_from_bytes_unsynced, so the dedup settle path (every NEW chunk of every upload) routed through put_blob_from_bytes and its "idempotent" HEAD/get_properties probe — 2 round-trips per chunk for chunks the dedup layer already knows are new. Content-addressed keys make re-PUTs overwrite-safe, so the new overrides PUT directly. Azure additionally stopped copying every chunk (data.to_vec() → Bytes into azure_core::Body): 0.44 ms + 4 MiB transient alloc per 4 MiB chunk removed.

cargo run --release --features bench --example bench_s3_put
# 500 x 256 KiB chunks, concurrency 8, 10 ms/request stub
# BEFORE (HEAD+PUT)   1604 ms   500 HEADs + 500 PUTs
# AFTER  (PUT only)    868 ms   500 PUTs             1.8x

[9] Entity → DTO mapping — closed-set interning + 1-alloc formatting

Arc::<str>::from(&'static str) always allocates+copies, so every file row paid 4 allocations for values drawn from a ~60-string closed set (icon class, special class, category, mime), plus 2-alloc etag and 2-alloc size formatting; FolderDto additionally built its etag twice and cloned 4 Strings it could move. Now: LazyLock intern tables (lookup + refcount bump; unknown values fall back to Arc::from, same bytes), single-alloc compute_etag/format_file_size, and Folder::into_parts() moves.

cargo run --release --features bench --example bench_dto_map
# 10k rows                       ns/row   allocs/row
# File→FileDto    BEFORE         1229.2        10.96
# File→FileDto    AFTER          1004.9         3.96
# Folder→FolderDto BEFORE         425.2        11.80
# Folder→FolderDto AFTER          204.5         1.00
# gate: all DTO fields byte-identical BEFORE vs AFTER (10k files + 10k folders)

[10] CardDAV REPORT — dead double vCard generation + O(N²) scan — 9.8x

handle_report pre-generated a vCard for EVERY contact; the adapter then did a linear uid find per contact — O(N²) string compares — and DISCARDED the result (let _ = vcard), regenerating on demand inside write_contact_response anyway. Pure dead work, deleted; contact_to_vcard also switched push_str(&format!(…)) → write! (one temp String per vCard line removed).

cargo run --release --features bench --example bench_carddav_report
# N=5000   getetag                55.7 →  5.7 ms   9.8x
# N=5000   getetag+address-data   76.2 → 15.3 ms   5.0x
# gate: REPORT XML byte-identical BEFORE vs AFTER for all prop sets

[11] Search-results cache — entry count → byte weigher — bounded RSS

The cache was capped at 1000 ENTRIES with a 300 s TTL; each entry holds up to 500 enriched rows (~10 owned Strings each) and keys include user+query+offset+limit, so every keystroke/page/user minted an entry — ~300 MiB of invisible RSS was reachable. Now a byte weigher + 32 MiB budget (OXICLOUD_SEARCH_CACHE_MAX_BYTES), same TTL, same read latency.

cargo run --release --features bench --example bench_search_cache_mem
# 1000 pages x 500 rows          retained bytes    get() p50
# BEFORE (1000 entries)          ~298 MiB (9.3x)      155 ns
# AFTER  (32 MiB weigher)         31.9 MiB            155 ns   parity 1.00x

[12] Cargo — drop aws-config + aws-smithy-types

Both were direct dependencies with ZERO references in the codebase — S3BlobBackend builds its client purely from aws_sdk_s3::config with static credentials. aws-config alone dragged aws-sdk-sso, aws-sdk-ssooidc and aws-sdk-sts into every build. Dependency-graph nodes: 1728 → 1646. tokio's process feature (used by the ffmpeg thumbnailer) was only enabled transitively through aws-config's feature unification — it is now declared explicitly.

Frontend — cached Intl.DateTimeFormat — 51.6x

formatDate (and four sibling callsites) constructed a fresh Intl.DateTimeFormat per call (~131 µs each here) — paid roughly twice per row while rendering/scrolling file lists. Module-scope cache keyed by (locale, options), invalidated on languagechange.

cd frontend && npx vitest run src/lib/utils/formatDate.bench.test.ts
# 20k dates: cached 50.6 ms vs per-call 2612.0 ms (51.6x); output-identity
# matrix across en/es/ar/ja and every option shape used by the app

Audited but NOT adopted (for the record)

  • Fat LTO / panic=abort / OpenAPI LazyLock: refuted by the verification pass (sub-1% plausible gain, or cold paths; catch_unwind shields pdf-extract so panic=abort is off the table).
  • Chained clone-on-hit drive caches, localeCompare→Intl.Collator: measured previously — residual gains are noise or regressions (benches/CHROOT-CACHE.md, benches/NPLUS1-AND-CACHES.md).
  • Follow-ups worth a future round (confirmed real, not yet gated): grouped/swimlane files view is unvirtualized (10k-row DOM); Azure download path buffers whole blobs in RAM (needs an Azurite-gated bench); face-indexing spawns unbounded per-image tasks; WebDAV drive-selector resolution re-runs the grants join per request (cacheable like CHROOT-CACHE); make_file_path split→rejoin + NFC copy per listing row.