perf(thumbnails): SIMD resize via fast_image_resize (PNG 2.6x, RAM 2.5x lower)
Replace the image crate's scalar resampler with fast_image_resize (AVX2/SSE4.1/NEON) in a shared encode_thumbnail() helper. render_all now converts to RGB8 once and SIMD-resizes the shared buffer per size. Lanczos3 for downscaling, CatmullRom when upscaling (Lanczos rings on enlargement). Also folds the duplicated path-variant generate_all_sizes_background into the shared render path -- it had missed BOTH shrink-on-load and SIMD resizing -- so every thumbnail path now goes through one optimised routine (no duplication). Measured on 14 cores vs the post-1.5 state (benches/BASELINE.md): - PNG 2.60x faster (33.6->12.9ms), GIF/WebP 1.25-1.6x: full-resolution decode paths where the resize dominates, so SIMD helps most - JPEG only ~7% (shrink-on-load already shrank the bitmap) but peak heap fell another ~2.5x (17.6->7.1MB): tight RGB buffers, RGB conversion once - quality SSIM 0.986-0.994 at identical dims (>=0.98 gate) Thumbnails are now exactly max_dim on the long side (e.g. 400x266) vs the old fit-within 399x266 -- a <=1px change, invisible under object-fit: cover. Bench example gains an exact-dims quality reference + semaphore throughput table. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -43,6 +43,9 @@ image = { version = "0.25.10", default-features = false, features = ["jpeg", "pn
|
||||
# Rust, no C toolchain. The `image` crate's zune-jpeg backend can't scale during
|
||||
# decode; this can, cutting decode time/RAM ~order-of-magnitude on large photos.
|
||||
jpeg-decoder = "0.3"
|
||||
# SIMD (AVX2/SSE4.1/NEON) image resizing for thumbnails — far faster than the
|
||||
# `image` crate's scalar resampler, and it speeds every format (incl. PNG/WebP).
|
||||
fast_image_resize = "5"
|
||||
id3 = "1.17"
|
||||
mp3-duration = "0.1"
|
||||
kamadak-exif = "0.6.1"
|
||||
|
||||
Reference in New Issue
Block a user