perf: round 15 — grouped-listing O(N²) rebucket, exif/reseed allocs, tantivy zero-hit snippet skip

Benchmark-gated, same rule as rounds 2–14: every change ships with a
BEFORE/AFTER benchmark and an equivalence/safety gate; an AFTER that doesn't
beat its BEFORE is rolled back. The rule is encoded per harness (GATE FAIL
non-zero exit in the Rust examples, threshold expect() in vitest).

F1 — Grouped listings (trash / recent / favorites / shared-with-me)
re-bucketed the WHOLE accumulated list on every infinite-scroll page.
ResourceSectionsBuilder (new, off the reactive graph) re-buckets only the
fresh page and hands VirtualList the same rows array reference for untouched
buckets. 50×50 (2 500-item) drain: 63 750 → 2 500 bucketOf calls (25.5×),
12.5 → 1.3 ms wall (9.9×); O(N²/page) → O(N). Deep-equal to the full-rebuild
reference at every page for both a contiguous (date) and a non-contiguous
(trash-by-drive) group-by; reference-stability + fallback gated.

B1 — exif Make/Model: the display String was thrown away to allocate the
trimmed copy; display_value_trimmed trims in place (drain + truncate), 2 → 1
alloc per field (8 → 4 allocs/op, 1.26×).

B2 — content-index worker: text_extractor::supports (lowercases MIME +
extension) was called twice per file per drain batch; classify once into a
Vec<bool> and thread it through both uses. 256-file batch: 704 → 353 allocs,
34.5 → 16.7 µs (2.07×).

B3 — tantivy: skip SnippetGenerator::create on a zero-hit content search
(return Ok(vec![]) once top_docs.is_empty()); the per-hit loop was empty.
400-doc index: 1 575.6 → 1 237.2 ns (1.27×), widens with index size.

Harnesses: examples/bench_round15_micro.rs, examples/bench_round15_tantivy.rs,
frontend resourceSections.bench.test.ts; writeup in benches/ROUND15.md. Also
normalizes two round14 bench examples that were committed unformatted
(cargo fmt --all).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012o47jSrtL7xuNGTHXmtiYL
This commit is contained in:
Claude
2026-07-19 11:36:47 +00:00
parent 76b9113c96
commit 3be85fa9f0
12 changed files with 1244 additions and 67 deletions
+16 -4
View File
@@ -57,7 +57,11 @@ fn bytes_to_embedding(b: &[u8]) -> Vec<f32> {
/// hydrate the full 10-column row (decoding the 2 KiB embedding like
/// `row_to_face`), then filter `user_id == caller` in Rust and keep only
/// `(id, person_id, bbox)`.
async fn boxes_before(pool: &PgPool, file_id: Uuid, caller: Uuid) -> Vec<(Uuid, Option<Uuid>, Vec<f32>)> {
async fn boxes_before(
pool: &PgPool,
file_id: Uuid,
caller: Uuid,
) -> Vec<(Uuid, Option<Uuid>, Vec<f32>)> {
let rows = sqlx::query(
"SELECT id, file_id, user_id, person_id, bbox, det_score, quality, embedding, blob_hash, created_at
FROM faces.faces WHERE file_id = $1",
@@ -83,7 +87,11 @@ async fn boxes_before(pool: &PgPool, file_id: Uuid, caller: Uuid) -> Vec<(Uuid,
}
/// AFTER: narrow projection, caller filter in SQL.
async fn boxes_after(pool: &PgPool, file_id: Uuid, caller: Uuid) -> Vec<(Uuid, Option<Uuid>, Vec<f32>)> {
async fn boxes_after(
pool: &PgPool,
file_id: Uuid,
caller: Uuid,
) -> Vec<(Uuid, Option<Uuid>, Vec<f32>)> {
let rows = sqlx::query(
"SELECT id, person_id, bbox FROM faces.faces WHERE file_id = $1 AND user_id = $2",
)
@@ -200,8 +208,12 @@ async fn section_face_boxes(pool: &PgPool) {
let wire_after = n * (16 + 16 + 16 + 8);
println!("\n## [Q1] Lightbox face boxes — group photo, {n} faces");
println!("| arm | mean ms | p50 ms | p95 ms | ~bytes/req |");
println!("| BEFORE wide row (incl. embedding) | {wm:>7.3} | {wp50:>6.3} | {wp95:>6.3} | {wire_before:>9} |");
println!("| AFTER narrow (id,person,bbox) | {nm:>7.3} | {np50:>6.3} | {np95:>6.3} | {wire_after:>9} |");
println!(
"| BEFORE wide row (incl. embedding) | {wm:>7.3} | {wp50:>6.3} | {wp95:>6.3} | {wire_before:>9} |"
);
println!(
"| AFTER narrow (id,person,bbox) | {nm:>7.3} | {np50:>6.3} | {np95:>6.3} | {wire_after:>9} |"
);
println!(
"# {:.2}x faster; ~{} KiB embedding/columns off the wire per lightbox open (scales with face count)",
wm / nm,