Files
Oxicloud/examples/bench_round26_diskio.rs
T
Claude 5b2bb8f883 perf(round26): drive-policy JSONB decode, CachedBlobBackend shard-dir pre-create, delta-upload foldhash
Three benchmark-gated optimizations from the ROUND25 backlog (benches/ROUND26.md),
each with a BEFORE/AFTER gate that rolls back if AFTER does not beat BEFORE:

- P1 drive_pg_repository policy reads: decode d.policies through
  sqlx::types::Json<DrivePolicies> (one from_slice over the raw JSONB bytes)
  instead of a throwaway serde_json::Value DOM + DrivePolicies::from_value —
  6 -> 0 allocs/read, 2.77x wall. A shared policies_from_row helper preserves
  the lenient unwrap_or_default fallback (malformed bag -> all-false).
- D1 CachedBlobBackend: pre-create the 256 {00..ff} shard dirs at initialize()
  (mirroring LocalBlobBackend, reusing HEX_PREFIXES) and drop the redundant
  per-write create_dir_all on already-existing shards — ~45us + a blocking-pool
  dispatch removed per cache write on cached-remote deployments.
- G1 delta-upload have/need hash sets (distinct_hashes, authorize_chunk_download):
  SipHash -> foldhash::quality::RandomState — a fast hasher that stays
  DoS-resistant via a per-instance random seed, the required property for the
  attacker-controlled 64-hex client hashes — 2.37x wall on a 40k-hash
  negotiation. foldhash was already in the lockfile transitively (hashbrown).

Tested and REVERTED (kept as-is): moving the moka eviction unlink off the reactor
via spawn_blocking. The benchmark refuted it — on the local cache dir the
spawn_blocking dispatch (~20us) costs more than the inline unlink (~7us) it would
replace. See ROUND26.md §D2.

Adds bench_round26_{micro,diskio,hasher} (counting allocator / async wall / wall).
Verified: cargo fmt clean, cargo clippy --features bench -D warnings clean,
cargo test --lib --features bench = 529 passed / 0 failed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L8gs91AhmazoxMsDcNk3KT
2026-07-21 01:05:12 +00:00

85 lines
3.4 KiB
Rust

//! Round-26 disk-I/O pack (no Postgres) — async wall on a tmpfs-backed tempdir.
//!
//! [D1] `CachedBlobBackend::initialize` creates ONLY `cache_dir`, never the 256
//! `{00..ff}` shard dirs (the line-122 comment claims otherwise), so each
//! of the three cache-write sites re-runs `tokio::fs::create_dir_all(parent)`
//! on the hot path — a wasted `mkdirat(EEXIST)` + component stat + a
//! blocking-pool dispatch per chunk write on cached-remote deployments.
//! AFTER pre-creates the shard dirs at init (mirroring
//! `LocalBlobBackend::initialize`) and drops the per-write call. Gate:
//! AFTER wall (per write) strictly lower than BEFORE (the redundant
//! create_dir_all).
//!
//! [D2] TESTED AND REVERTED — see benches/ROUND26.md. Moving the moka
//! eviction-listener unlink off the reactor via `spawn_blocking` was
//! refuted by the benchmark: on the local cache dir (fast unlink ~7 µs)
//! the `spawn_blocking` dispatch (~20 µs) costs MORE on the reactor than
//! the inline `std::fs::remove_file` it replaces. The original inline
//! unlink ("a quick unlink on the inserting task's thread") is correct
//! for the fast-local-cache case; kept as-is.
//!
//! Run:
//! RUSTFLAGS="-C target-cpu=x86-64-v3" \
//! cargo run --release --features bench --example bench_round26_diskio
//! Tunables (env): D1_ITERS (20000)
use std::env;
use std::time::Instant;
fn env_or<T: std::str::FromStr>(key: &str, default: T) -> T {
env::var(key)
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(default)
}
fn gate(tag: &str, metric: &str, before: f64, after: f64) {
if !(after < before) {
eprintln!("GATE FAIL [{tag}] {metric}: AFTER {after} !< BEFORE {before} — rollback");
std::process::exit(1);
}
}
// ── [D1] redundant create_dir_all on a warm shard vs skip ────────────────────
async fn section_d1() {
let iters: u64 = env_or("D1_ITERS", 20_000);
let dir = tempfile::tempdir().expect("tempdir");
let shard = dir.path().join("ab");
// Shard pre-created once (what AFTER's initialize does).
tokio::fs::create_dir_all(&shard).await.unwrap();
// warm
let _ = tokio::fs::create_dir_all(&shard).await;
// BEFORE: per-write create_dir_all(parent) on the already-existing shard.
let t = Instant::now();
for _ in 0..iters {
let _ = tokio::fs::create_dir_all(&shard).await;
}
let before_ns = t.elapsed().as_nanos() as f64 / iters as f64;
// AFTER: shard guaranteed present at init → the write path skips the call.
let t = Instant::now();
for _ in 0..iters {
std::hint::black_box(&shard);
}
let after_ns = t.elapsed().as_nanos() as f64 / iters as f64;
println!("## [D1] cache-write create_dir_all on a warm shard");
println!("| arm | ns/write |");
println!("| BEFORE | {before_ns:>8.1} |");
println!("| AFTER | {after_ns:>8.1} |");
println!(
"# {:.1}x — redundant create_dir_all removed per cache write\n",
before_ns / after_ns.max(0.001)
);
gate("D1", "ns/write", before_ns, after_ns);
}
#[tokio::main(flavor = "multi_thread", worker_threads = 4)]
async fn main() {
println!("# Round-26 disk-I/O pack\n");
section_d1().await;
println!("All Round-26 disk-I/O sections passed their gate.");
}