perf: round 23 — Postgres query-shape pass: typed JSONB decode, drive-policy borrow-deserialize, user-profile join!, subject-group CTE reuse, dedup unzip

Benchmark-gated, same rule as ROUND2-22: BEFORE/AFTER with a value-equivalence
gate and rollback-on-regression. Two harnesses — bench_round23_micro (no
Postgres; deterministic allocation gate) and bench_round23_queries (live
Postgres; p50 latency + strict equivalence gate against seeded fixtures). See
benches/ROUND23.md.

- J1: contact_pg_repository::row_to_contact (+ the inlined contact_group sibling)
  decode the 3 JSONB columns via sqlx::types::Json<T> (one from_slice pass)
  instead of row.get::<serde_json::Value> + from_value (a throwaway Value DOM
  per column, walked a second time). Per contact row of every list / multiget /
  CardDAV sync. Micro 84 -> 33 allocs/op (2.15x); PG 3794 -> 2360 ns/contact
  (1.61x) on 500 real rows.
- J2: DrivePolicies::from_value deserializes straight from the borrow
  (T::deserialize(&Value)) instead of from_value(value.clone()) — dropping the
  full-DOM clone on every drive-policy read (move/copy, share, grant); one-line
  body change, all 7 callers unchanged. Micro 5 -> 0 allocs/op (11.51x).
- P1: get_user_profile overlaps the two independent caller+target reads with
  tokio::join! (self-case still a single fetch; caller-error precedence
  preserved via caller_res? first) instead of two serial round-trips. PG
  577 -> 312 us/call (1.85x).
- G1: subject_group remove_member computes the child's transitive-user recursive
  CTE once and reuses it for both the would-empty pre-check and the cache
  invalidation, instead of running the identical CTE twice (the edge delete is
  above the child, so its descendants can't change). PG 829 -> 412 us/removal
  (2.01x).
- U1: dedup_service (store_loose_chunks final registration + the ingest
  run_rollback) reshapes the owned, dead-after Vec<(String,i64)> via
  into_iter().unzip() instead of cloning every 64-byte hash for the
  sync_blobs(&[String]) + UNNEST bind. Micro 256 -> 0 hash clones.

Verified: cargo clippy --features bench --all-targets -D warnings clean, cargo
fmt --all --check clean, cargo test --lib --features bench = 529 passed / 0
failed. The PG benches run against a local PostgreSQL 16 (schema applied from
migrations/); every equivalence gate passes.

The download_zip per-item N+1 (the audit's highest raw-latency candidate) is
deferred to a dedicated pass: its fix moves the sole authorization inside the
stream call, so it needs an AuthZ-ordering + anti-enumeration proof, not a perf
banner.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DKyQ4AnYtgp1JtjzweyMeo
This commit is contained in:
Claude
2026-07-20 15:25:42 +00:00
parent 992bdae898
commit 1ec7030cc7
10 changed files with 1107 additions and 42 deletions
+9 -4
View File
@@ -209,8 +209,11 @@ impl IngestGuard {
// sweep can reclaim the bytes — a backend file with no PG row would be
// invisible to it. ON CONFLICT DO NOTHING keeps a concurrent
// uploader's row (and its references) intact.
let hashes: Vec<String> = written.iter().map(|(h, _)| h.clone()).collect();
let sizes: Vec<i64> = written.iter().map(|(_, s)| *s).collect();
// `written` is owned and dead after this rollback — unzip it (moving each
// 64-byte hash String out) instead of cloning every hash purely to
// reshape for `sync_blobs(&[String])` + the UNNEST bind.
// (benches/ROUND23.md §U1)
let (hashes, sizes): (Vec<String>, Vec<i64>) = written.into_iter().unzip();
if let Err(e) = backend.sync_blobs(&hashes).await {
tracing::warn!(
"Ingest rollback: sync of {} chunks failed: {e}",
@@ -896,8 +899,10 @@ impl DedupService {
if !new_rows.is_empty() {
// Durability before visibility — same invariant as the ingest
// engine: no PG row may ever point at unsynced bytes.
let hashes: Vec<String> = new_rows.iter().map(|(h, _)| h.clone()).collect();
let sizes: Vec<i64> = new_rows.iter().map(|(_, s)| *s).collect();
// `new_rows` is owned and dead after this block — unzip (move the
// hash Strings out) instead of cloning each one for the reshape +
// UNNEST bind. (benches/ROUND23.md §U1)
let (hashes, sizes): (Vec<String>, Vec<i64>) = new_rows.into_iter().unzip();
self.backend.sync_blobs(&hashes).await?;
sqlx::query(
"INSERT INTO storage.blobs (hash, size, ref_count, orphaned_at)