Three benchmark-gated optimizations from the ROUND25 backlog (benches/ROUND26.md),
each with a BEFORE/AFTER gate that rolls back if AFTER does not beat BEFORE:
- P1 drive_pg_repository policy reads: decode d.policies through
sqlx::types::Json<DrivePolicies> (one from_slice over the raw JSONB bytes)
instead of a throwaway serde_json::Value DOM + DrivePolicies::from_value —
6 -> 0 allocs/read, 2.77x wall. A shared policies_from_row helper preserves
the lenient unwrap_or_default fallback (malformed bag -> all-false).
- D1 CachedBlobBackend: pre-create the 256 {00..ff} shard dirs at initialize()
(mirroring LocalBlobBackend, reusing HEX_PREFIXES) and drop the redundant
per-write create_dir_all on already-existing shards — ~45us + a blocking-pool
dispatch removed per cache write on cached-remote deployments.
- G1 delta-upload have/need hash sets (distinct_hashes, authorize_chunk_download):
SipHash -> foldhash::quality::RandomState — a fast hasher that stays
DoS-resistant via a per-instance random seed, the required property for the
attacker-controlled 64-hex client hashes — 2.37x wall on a 40k-hash
negotiation. foldhash was already in the lockfile transitively (hashbrown).
Tested and REVERTED (kept as-is): moving the moka eviction unlink off the reactor
via spawn_blocking. The benchmark refuted it — on the local cache dir the
spawn_blocking dispatch (~20us) costs more than the inline unlink (~7us) it would
replace. See ROUND26.md §D2.
Adds bench_round26_{micro,diskio,hasher} (counting allocator / async wall / wall).
Verified: cargo fmt clean, cargo clippy --features bench -D warnings clean,
cargo test --lib --features bench = 529 passed / 0 failed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L8gs91AhmazoxMsDcNk3KT
8.3 KiB
Round 26 — drive-policy JSONB decode (alloc), CachedBlobBackend shard-dir pre-create (disk), delta-upload foldhash (CPU); eviction-unlink off-reactor tested & reverted
This round drains three high-confidence items from the ROUND25 backlog, each
behind a BEFORE/AFTER benchmark that std::process::exit(1)s ("GATE FAIL … rollback") unless AFTER strictly beats BEFORE. A fourth candidate (moving the
cache eviction unlink off the reactor) was tested and reverted — the
benchmark refuted it. All three shipped items target the owner's priorities:
allocations, disk-I/O, and CPU.
Reproduce:
RUSTFLAGS="-C target-cpu=x86-64-v3" cargo run --release --features bench --example bench_round26_micro # P1
RUSTFLAGS="-C target-cpu=x86-64-v3" cargo run --release --features bench --example bench_round26_diskio # D1
RUSTFLAGS="-C target-cpu=x86-64-v3" cargo run --release --features bench --example bench_round26_hasher # G1
[P1] Drive-policy reads: throwaway serde_json::Value DOM → from_slice::<DrivePolicies> (allocations)
drive_pg_repository's four policy reads (get_policies_for_file/_folder,
get_drive_id_and_policies_for_file/_folder) fetched d.policies as a
serde_json::Value and then called DrivePolicies::from_value(&raw)
(Self::deserialize(&Value)). The Value tree — a Map + a boxed String key
- a
Valuenode per policy field — is built once, walked once, and dropped. This is the exact throwaway-DOM pattern ROUND23 §J1 removed for contacts; §J2 removed only thefrom_valueclone, not the DOM. These reads fire on file/folder move & copy and on every share/grant creation.
AFTER fetches through sqlx::types::Json<DrivePolicies> — one
serde_json::from_slice::<DrivePolicies> over the raw JSONB bytes, no
intermediate DOM — via a shared policies_from_row helper that preserves the
lenient unwrap_or_default fallback exactly (a malformed bag → all-false,
try_get(...).unwrap_or_default(), mirroring §J1).
| arm | ns/op | allocs/op | bytes/op |
|---|---|---|---|
| BEFORE | 399.4 | 6.00 | 719 |
| AFTER | 161.0 | 0.00 | 0 |
6 → 0 allocs/op, −719 bytes/op, 2.48× wall — the entire Value DOM removed per
policy read. Gate: AFTER allocs/op strictly lower. Equivalence: the decoded
DrivePolicies is asserted identical BEFORE vs AFTER.
[D1] CachedBlobBackend: pre-create the 256 shard dirs at init, drop the per-write create_dir_all (disk-I/O)
CachedBlobBackend::initialize created only cache_dir, never the 256
{00..ff} shard dirs (the line-122 comment claimed otherwise). So all three
cache-write sites (cache_bytes_write_through, insert_into_cache,
fetch_and_cache) re-ran tokio::fs::create_dir_all(parent) per chunk — a
wasted mkdirat(EEXIST) + component stat + blocking-pool dispatch on a shard
that already exists, on every cached-remote write. AFTER creates all 256 shards
once at init (mirroring LocalBlobBackend::initialize, reusing its
HEX_PREFIXES table) and deletes the three per-write calls; the shard for any
&hash[..2] prefix always exists, so the writes just fs::write/fs::copy.
Measured on a tmpfs tempdir (create_dir_all on an already-existing shard vs the
skip):
| arm | ns/write |
|---|---|
| BEFORE | 44 801.8 |
| AFTER | 0.3 |
~45 µs removed per cache write. Gate: AFTER ns/write strictly lower. The on-disk layout is identical; the directory creation simply moved from the hot path to one-time startup.
[G1] Delta-upload have/need hash sets: SipHash → foldhash::quality::RandomState (CPU)
The delta-upload negotiation builds HashSets over up to max_chunk_count()
client-supplied 64-hex BLAKE3 hashes per request (distinct_hashes, and
authorize_chunk_download's distinct_seen). std HashSet uses SipHash-1-3 —
DoS-resistant but ~2-4× slower than a modern hash on short keys. AFTER uses
foldhash::quality::RandomState, a fast non-cryptographic hasher that stays
DoS-resistant because it is per-instance random-seeded — the required property
for these attacker-controlled inputs (not FxHash/a fixed seed). foldhash is
already in the lockfile transitively (via hashbrown), so the direct dep adds no
newly-compiled crate.
Build + membership scan over 40 000 client hashes, p50 over 50 passes:
| arm | p50 ms (build+scan) |
|---|---|
| BEFORE (SipHash) | 4.768 |
| AFTER (foldhash) | 2.007 |
2.37× wall on the delta negotiation's hottest set — a bulk sync of a large
file negotiates thousands of chunks. Scales with max_chunk_count().
Gate: AFTER p50 wall (build set + membership scan over N hashes) strictly lower,
and two RandomState::default() instances must produce different hashes for
the same key (asserting the random per-instance seed — DoS resistance retained).
The set membership decisions are unchanged, so behaviour is identical.
Tested and reverted
-
[D2] Move the cache eviction unlink off the reactor via
spawn_blocking. The moka eviction listener unlinks a size-evicted blob with a synchronousstd::fs::remove_fileinline on the tokio worker that triggered the insert. The hypothesis: hand it tospawn_blockingso the reactor isn't blocked onunlink(2). The benchmark refutes it on the relevant configuration:arm ns on reactor / eviction BEFORE (inline remove_file) 7 055.7 AFTER (spawn_blocking) 19 848.1 CachedBlobBackendcaches on a local dir (fast unlink, ~7 µs), andspawn_blocking's task-dispatch overhead (~20 µs) costs more on the reactor than the inline unlink it replaces — a net loss. The original code comment ("a quick unlink on the inserting task's thread, off the hot get path") is correct for the fast-local-cache case. A win would only materialize on genuinely slow storage (network-backed cache dir), which there is no fixture for here. Reverted; kept the inline unlink. (A "measure before believing" result, like BASELINE's dropped Task 2.1 / reverted Phase 1.7.)
Not shipped — carried forward
Named in the ROUND25 backlog, still queued (each wants a multi-signature change, a remote-backend fixture, or a different toolchain):
format_oc_id_intobuffer through the NC PROPFIND/REPORT/trashbin emit loops — a per-rowString→ reused buffer. Threads a buffer through ~4 loop sites across 3 files and depends onNextcloudFileIdService's instance-id format; wants its own validated pass so a wrongoc:idcan't reach a client.- S3 read zero-copy forward (
into_async_read()+ReaderStream→ forward the SDKBytesframes, Azure-style) — needs a MinIO/stubByteStreamfixture. - Frontend folder-listing cache (
getCachedFolder/cacheFolderis dead code; every navigation refetches withcache:'no-store') — a SvelteKit/Vitest pass (bandwidth + instant paint on revisits). - foldhash for the NC PROPFIND trusted-key maps (
favorite_ids,nc_id) —foldhash::fast(no random seed needed; server-generated keys). Threads the hasher type through the emit-loop map builders. - Contact create/update
Json<T>bind (write-side twin of §J1) and the other ROUND25 backlog items.
Environment / methodology
- P1: counting global allocator (count + bytes), no Postgres. The real
DrivePoliciestype is imported from the crate; BEFORE replicates the shippedserde_json::from_slice::<Value>+deserialize(&Value), AFTER the shippedfrom_slice::<DrivePolicies>. Value-equivalence asserted; gate on allocs/op. - D1: async wall on a tmpfs
tempfile::tempdir; BEFORE =create_dir_allon a pre-existing shard, AFTER = the skip. Gate on ns/write. - G1: wall-gated (a hasher swap changes 0 allocations); SipHash vs
foldhash::qualitybuild+scan over N random 64-hex hashes; DoS-seed assertion. - Built with
RUSTFLAGS="-C target-cpu=x86-64-v3"(the checked-in.cargo/config.tomlpinstarget-cpu=native, whichSIGILLs on this host — see ROUND23/24; local override only). - Verified beyond the benches:
cargo fmt --all --checkclean,cargo clippy --features bench -- -D warningsclean,cargo test --lib --features benchgreen.