perf(thumbnails): shrink-on-load JPEG decode (1.8-2× faster, 5-15× less RAM)

Decode JPEGs at the smallest DCT scale (1/8·1/4·1/2·1/1) whose long axis is
still ≥ the largest needed thumbnail (800px), via jpeg-decoder, instead of a
full-resolution decode through the image crate. The full-res bitmap — the
dominant time and RAM cost — is never materialised. PNG/GIF/WebP and unusual
JPEG colour spaces (CMYK / 16-bit grey) fall back to a full decode.

Extracts the shared decode + EXIF-orientation logic into decode_oriented(),
removing the duplication that existed between render_thumbnail_from_data and
render_all_thumbnails_from_data.

Measured on 14 cores (see benches/BASELINE.md):
- render_all 1.8-2.0× faster (12MP 111->61ms, 48MP 398->203ms)
- peak heap 5.5-14.8× lower, now decoupled from source MP (~18-25MB regardless)
- saturated throughput 3-3.6× (parallel efficiency 4.9×->8.5×)
- quality SSIM 0.987-0.999 (>=0.98 gate), PSNR 47-55dB

Also adds the Phase 0 benchmark harness (gated behind the `bench` feature, zero
prod impact): deterministic image corpus (src/bench_support.rs), criterion
latency bench (benches/thumbnails.rs), and a peak-RAM/throughput/SSIM harness
(examples/bench_thumbnails_mem.rs). Baseline + before/after in benches/BASELINE.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
DioCrafts
2026-06-21 15:13:03 +02:00
parent 08b36cf0d4
commit fd5808c157
9 changed files with 1415 additions and 39 deletions
+317
View File
@@ -0,0 +1,317 @@
//! Phase 0 perf-benchmark support — deterministic image corpus.
//!
//! Shared by `benches/thumbnails.rs` (Task 0.2, criterion latency + output
//! size) and `examples/bench_thumbnails_mem.rs` (Task 0.3, peak RAM +
//! throughput). Gated behind the `bench` feature so it never touches normal
//! builds.
//!
//! The corpus is **generated deterministically** (a low-frequency gradient plus
//! seeded high-frequency xorshift noise) so it is reproducible, license-free and
//! gives the decoder/resizer realistic work without committing large binaries to
//! git. Files are written to `benches/corpus/` (git-ignored) on first run and
//! reused afterwards.
//!
//! Files already present on disk are **always preferred** over generation — so
//! you can drop your own real photos into `benches/corpus/` using the documented
//! filenames (see [`CASE_SPECS`]) to benchmark against real-world data.
use std::io::Cursor;
use std::path::PathBuf;
use image::codecs::jpeg::JpegEncoder;
use image::{DynamicImage, ImageFormat, Rgb, RgbImage};
/// One corpus entry: the encoded file bytes plus its probed dimensions.
pub struct CorpusCase {
/// Stable identifier (e.g. `"jpeg_12mp"`), used as the bench label.
pub name: &'static str,
/// Container format (`"jpeg"`, `"png"`, `"gif"`, `"webp"`).
pub format: &'static str,
/// Actual decoded width (probed from the bytes).
pub width: u32,
/// Actual decoded height (probed from the bytes).
pub height: u32,
/// Encoded file bytes (what the thumbnail pipeline receives as `&[u8]`).
pub bytes: Vec<u8>,
}
impl CorpusCase {
/// Megapixels of the source image (decoded resolution).
pub fn megapixels(&self) -> f64 {
(self.width as f64 * self.height as f64) / 1_000_000.0
}
}
/// Declarative description of a synthetic corpus image.
struct Spec {
name: &'static str,
filename: &'static str,
format: &'static str,
width: u32,
height: u32,
/// JPEG quality (ignored for non-JPEG formats).
quality: u8,
/// When set, injects an EXIF Orientation tag into the JPEG (e.g. 6 = rotate
/// 90° CW) to exercise the orientation-correction path.
exif_orientation: Option<u16>,
}
/// The corpus matrix: a spread of sizes and formats so we never trust an
/// average. Drop a real photo at `benches/corpus/<filename>` to override any
/// entry with real-world data.
const CASE_SPECS: &[Spec] = &[
// JPEG photo-like sources at the three sizes that matter for "hundreds of
// phone/camera photos". These dominate the real upload load.
Spec {
name: "jpeg_12mp",
filename: "jpeg_12mp.jpg",
format: "jpeg",
width: 4000,
height: 3000,
quality: 90,
exif_orientation: None,
},
Spec {
name: "jpeg_24mp",
filename: "jpeg_24mp.jpg",
format: "jpeg",
width: 6000,
height: 4000,
quality: 90,
exif_orientation: None,
},
// 8000×6000 = 48 MP, just under the 50 MP MAX_DECODE_PIXELS guard.
Spec {
name: "jpeg_48mp",
filename: "jpeg_48mp.jpg",
format: "jpeg",
width: 8000,
height: 6000,
quality: 90,
exif_orientation: None,
},
// Non-JPEG decode paths (no DCT shrink-on-load possible — useful contrast).
Spec {
name: "png_large",
filename: "png_large.png",
format: "png",
width: 3000,
height: 2000,
quality: 0,
exif_orientation: None,
},
Spec {
name: "gif_large",
filename: "gif_large.gif",
format: "gif",
width: 600,
height: 600,
quality: 0,
exif_orientation: None,
},
Spec {
name: "webp_large",
filename: "webp_large.webp",
format: "webp",
width: 1280,
height: 853,
quality: 0,
exif_orientation: None,
},
// Small source: exercises the (future) no-upscale clamp — Large=800 target
// is bigger than the 300 px source.
Spec {
name: "small_300",
filename: "small_300.jpg",
format: "jpeg",
width: 300,
height: 300,
quality: 90,
exif_orientation: None,
},
// EXIF orientation ≠ 1: exercises the rotate/flip correction path.
Spec {
name: "jpeg_exif_orient",
filename: "jpeg_exif_orient.jpg",
format: "jpeg",
width: 4000,
height: 3000,
quality: 90,
exif_orientation: Some(6),
},
];
/// Absolute path to `benches/corpus/` next to this crate's `Cargo.toml`.
pub fn corpus_dir() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("benches")
.join("corpus")
}
/// Load the corpus, generating any missing files on disk first.
///
/// Existing files win, so user-provided real photos are used as-is. A case that
/// fails to generate (e.g. an unavailable encoder) is logged and skipped rather
/// than aborting the whole baseline.
pub fn load_or_generate() -> Vec<CorpusCase> {
let dir = corpus_dir();
if let Err(e) = std::fs::create_dir_all(&dir) {
panic!("bench_support: cannot create {}: {e}", dir.display());
}
let mut out = Vec::new();
for spec in CASE_SPECS {
let path = dir.join(spec.filename);
let bytes = if path.exists() {
match std::fs::read(&path) {
Ok(b) => b,
Err(e) => {
eprintln!("bench_support: skipping {} (read failed: {e})", spec.name);
continue;
}
}
} else {
match generate(spec) {
Ok(b) => {
if let Err(e) = std::fs::write(&path, &b) {
eprintln!("bench_support: could not cache {} ({e})", path.display());
}
b
}
Err(e) => {
eprintln!(
"bench_support: skipping {} (generate failed: {e})",
spec.name
);
continue;
}
}
};
let (width, height) = probe_dimensions(&bytes).unwrap_or((spec.width, spec.height));
out.push(CorpusCase {
name: spec.name,
format: spec.format,
width,
height,
bytes,
});
}
out
}
/// Probe the decoded dimensions of encoded image bytes without a full decode.
fn probe_dimensions(bytes: &[u8]) -> Option<(u32, u32)> {
image::ImageReader::new(Cursor::new(bytes))
.with_guessed_format()
.ok()?
.into_dimensions()
.ok()
}
/// Render and encode one spec into file bytes.
fn generate(spec: &Spec) -> Result<Vec<u8>, String> {
let img = synthesize(spec.width, spec.height, seed_for(spec.name));
match spec.format {
"jpeg" => {
let mut buf = Vec::new();
let encoder = JpegEncoder::new_with_quality(&mut buf, spec.quality);
img.write_with_encoder(encoder)
.map_err(|e| format!("jpeg encode: {e}"))?;
match spec.exif_orientation {
Some(o) => inject_exif_orientation(&buf, o),
None => Ok(buf),
}
}
other => {
let fmt = match other {
"png" => ImageFormat::Png,
"gif" => ImageFormat::Gif,
"webp" => ImageFormat::WebP,
_ => return Err(format!("unknown format {other}")),
};
let mut buf = Vec::new();
DynamicImage::ImageRgb8(img)
.write_to(&mut Cursor::new(&mut buf), fmt)
.map_err(|e| format!("{other} encode: {e}"))?;
Ok(buf)
}
}
}
/// Build a photo-like RGB image: a smooth diagonal gradient (low frequency)
/// plus seeded ±32 white noise (high frequency). Deterministic for a given
/// seed, so corpus bytes are byte-stable across runs and machines.
fn synthesize(width: u32, height: u32, seed: u64) -> RgbImage {
let mut img = RgbImage::new(width, height);
let mut state = seed | 1; // xorshift requires a non-zero state
let (w, h) = (width.max(1), height.max(1));
for y in 0..height {
let gy = (y as i32 * 255 / h as i32).clamp(0, 255);
for x in 0..width {
let gx = (x as i32 * 255 / w as i32).clamp(0, 255);
let noise = (xorshift(&mut state) & 0x3F) as i32 - 32; // -32..=31
let r = (gx + noise).clamp(0, 255) as u8;
let g = (gy + noise).clamp(0, 255) as u8;
let b = (((gx + gy) / 2) + noise).clamp(0, 255) as u8;
img.put_pixel(x, y, Rgb([r, g, b]));
}
}
img
}
/// Tiny xorshift64 PRNG — fast, deterministic, no dependency.
fn xorshift(state: &mut u64) -> u64 {
let mut x = *state;
x ^= x << 13;
x ^= x >> 7;
x ^= x << 17;
*state = x;
x
}
/// Per-case fixed seed so each image has distinct noise but stays reproducible.
fn seed_for(name: &str) -> u64 {
// FNV-1a over the name → splitmix-ish spread.
let mut hash: u64 = 0xcbf29ce484222325;
for b in name.bytes() {
hash ^= b as u64;
hash = hash.wrapping_mul(0x100000001b3);
}
hash.wrapping_mul(0x9E3779B97F4A7C15) | 1
}
/// Splice a minimal, standard EXIF APP1 segment carrying a single Orientation
/// tag into a baseline JPEG, right after the SOI marker. Little-endian TIFF.
fn inject_exif_orientation(jpeg: &[u8], orientation: u16) -> Result<Vec<u8>, String> {
if jpeg.len() < 2 || jpeg[0] != 0xFF || jpeg[1] != 0xD8 {
return Err("not a JPEG (missing SOI)".into());
}
// TIFF body (little-endian "II").
let mut tiff = Vec::new();
tiff.extend_from_slice(b"II");
tiff.extend_from_slice(&0x2Au16.to_le_bytes()); // magic 42
tiff.extend_from_slice(&8u32.to_le_bytes()); // IFD0 offset
tiff.extend_from_slice(&1u16.to_le_bytes()); // 1 directory entry
tiff.extend_from_slice(&0x0112u16.to_le_bytes()); // tag: Orientation
tiff.extend_from_slice(&3u16.to_le_bytes()); // type: SHORT
tiff.extend_from_slice(&1u32.to_le_bytes()); // count
tiff.extend_from_slice(&(orientation as u32).to_le_bytes()); // value (SHORT in low bytes)
tiff.extend_from_slice(&0u32.to_le_bytes()); // next IFD = none
let mut payload = Vec::with_capacity(6 + tiff.len());
payload.extend_from_slice(b"Exif\0\0");
payload.extend_from_slice(&tiff);
let seg_len = u16::try_from(2 + payload.len()).map_err(|_| "EXIF segment too large")?;
let mut out = Vec::with_capacity(jpeg.len() + 4 + payload.len());
out.extend_from_slice(&jpeg[0..2]); // SOI
out.extend_from_slice(&[0xFF, 0xE1]); // APP1 marker
out.extend_from_slice(&seg_len.to_be_bytes()); // APP1 length (big-endian)
out.extend_from_slice(&payload);
out.extend_from_slice(&jpeg[2..]); // rest of the original JPEG
Ok(out)
}
+118 -34
View File
@@ -551,11 +551,31 @@ impl ThumbnailService {
Ok(bytes)
}
fn render_thumbnail_from_data(
/// Cheap magic-byte check for a JPEG container (SOI + first marker).
fn is_jpeg(data: &[u8]) -> bool {
data.len() >= 3 && data[0] == 0xFF && data[1] == 0xD8 && data[2] == 0xFF
}
/// Decode `data` into a `DynamicImage` sized for a thumbnail whose longest
/// side is `target_long`, with EXIF orientation already applied.
///
/// For JPEG this is **shrink-on-load**: the decoder emits the image at the
/// smallest DCT scale (1/8·1/4·1/2·1/1) whose longest axis is still ≥
/// `target_long`, so a 12 MP photo decodes ~16× fewer pixels for an 800 px
/// thumbnail — and the full-resolution bitmap (the dominant time/RAM cost)
/// is never materialised. PNG/GIF/WebP (no DCT scaling) and unusual JPEG
/// colour spaces (CMYK / 16-bit grey) fall back to a full decode.
fn decode_oriented(
data: &[u8],
size: ThumbnailSize,
) -> Result<Vec<u8>, ThumbnailError> {
let max_dim = size.max_dimension();
target_long: u32,
) -> Result<image::DynamicImage, ThumbnailError> {
// JPEG fast path: shrink-on-load. A non-JPEG, or a JPEG colour space we
// don't map (CMYK / 16-bit grey → `None`), falls through to a full decode.
if Self::is_jpeg(data)
&& let Some(img) = Self::decode_jpeg_scaled(data, target_long)?
{
return Ok(Self::apply_exif_orientation(data, img));
}
let (w, h) = image::ImageReader::new(std::io::Cursor::new(data))
.with_guessed_format()
@@ -568,17 +588,74 @@ impl ThumbnailService {
w as u64 * h as u64 / 1_000_000
)));
}
let img =
image::load_from_memory(data).map_err(|e| ThumbnailError::ImageError(e.to_string()))?;
Ok(Self::apply_exif_orientation(data, img))
}
let img = {
use crate::infrastructure::services::exif_service::{ExifService, apply_orientation};
let orientation = ExifService::extract(data)
.and_then(|m| m.orientation)
.unwrap_or(1);
apply_orientation(img, orientation)
/// JPEG shrink-on-load. Returns `Ok(None)` for colour spaces we don't map
/// (CMYK / 16-bit grey), signalling the caller to fall back to a full decode.
fn decode_jpeg_scaled(
data: &[u8],
target_long: u32,
) -> Result<Option<image::DynamicImage>, ThumbnailError> {
let mut decoder = jpeg_decoder::Decoder::new(std::io::Cursor::new(data));
// read_info() first so the dimension guard sees the *original* size.
decoder
.read_info()
.map_err(|e| ThumbnailError::ImageError(e.to_string()))?;
let info = decoder
.info()
.ok_or_else(|| ThumbnailError::ImageError("missing JPEG metadata".into()))?;
let (orig_w, orig_h) = (info.width as u64, info.height as u64);
if orig_w * orig_h > MAX_DECODE_PIXELS {
return Err(ThumbnailError::ImageError(format!(
"Image too large for thumbnail: {orig_w}×{orig_h} ({} MP, max {MAX_DECODE_PIXELS})",
orig_w * orig_h / 1_000_000
)));
}
// Request a `target_long` square box: jpeg-decoder picks the smallest
// scale whose longest axis is still ≥ target_long (its "≥ in at least
// one axis" rule reduces to the long axis since it dominates), so the
// later resample step only ever downscales — never upscales/blurs.
let req = target_long.min(u16::MAX as u32) as u16;
let (sw, sh) = decoder
.scale(req, req)
.map_err(|e| ThumbnailError::ImageError(e.to_string()))?;
let pixels = decoder
.decode()
.map_err(|e| ThumbnailError::ImageError(e.to_string()))?;
let (sw, sh) = (sw as u32, sh as u32);
let img = match decoder.info().map(|i| i.pixel_format) {
Some(jpeg_decoder::PixelFormat::RGB24) => {
image::RgbImage::from_raw(sw, sh, pixels).map(image::DynamicImage::ImageRgb8)
}
Some(jpeg_decoder::PixelFormat::L8) => {
image::GrayImage::from_raw(sw, sh, pixels).map(image::DynamicImage::ImageLuma8)
}
_ => None,
};
Ok(img)
}
/// Read EXIF orientation from the original bytes and rotate/flip the image.
/// (Applied after shrink-on-load, so the rotation works on the small bitmap.)
fn apply_exif_orientation(data: &[u8], img: image::DynamicImage) -> image::DynamicImage {
use crate::infrastructure::services::exif_service::{ExifService, apply_orientation};
let orientation = ExifService::extract(data)
.and_then(|m| m.orientation)
.unwrap_or(1);
apply_orientation(img, orientation)
}
fn render_thumbnail_from_data(
data: &[u8],
size: ThumbnailSize,
) -> Result<Vec<u8>, ThumbnailError> {
let max_dim = size.max_dimension();
let img = Self::decode_oriented(data, max_dim)?;
let (orig_width, orig_height) = (img.width(), img.height());
let (new_width, new_height) = if orig_width > orig_height {
@@ -608,29 +685,11 @@ impl ThumbnailService {
fn render_all_thumbnails_from_data(
data: &[u8],
) -> Result<Vec<(ThumbnailSize, Bytes)>, ThumbnailError> {
let (w, h) = image::ImageReader::new(std::io::Cursor::new(data))
.with_guessed_format()
.map_err(|e| ThumbnailError::ImageError(e.to_string()))?
.into_dimensions()
.map_err(|e| ThumbnailError::ImageError(e.to_string()))?;
if (w as u64) * (h as u64) > MAX_DECODE_PIXELS {
return Err(ThumbnailError::ImageError(format!(
"Image too large for thumbnail: {w}×{h} ({} MP, max {MAX_DECODE_PIXELS})",
w as u64 * h as u64 / 1_000_000
)));
}
let img =
image::load_from_memory(data).map_err(|e| ThumbnailError::ImageError(e.to_string()))?;
let img = {
use crate::infrastructure::services::exif_service::{ExifService, apply_orientation};
let orientation = ExifService::extract(data)
.and_then(|m| m.orientation)
.unwrap_or(1);
apply_orientation(img, orientation)
};
// Decode once, shrunk-on-load for the largest size (800 px); all three
// sizes are then resampled from this single shared bitmap. Sizing the
// shared decode to Large keeps quality for every size while paying the
// (now much smaller) decode cost only once.
let img = Self::decode_oriented(data, ThumbnailSize::Large.max_dimension())?;
let (orig_w, orig_h) = (img.width(), img.height());
ThumbnailSize::all()
@@ -1271,6 +1330,31 @@ impl ThumbnailPort for ThumbnailService {
}
}
/// Benchmark-only public surface (Phase 0 perf harness).
///
/// The real render functions are private (`render_thumbnail_from_data` /
/// `render_all_thumbnails_from_data`). Benches and examples are separate crate
/// targets and can only see `pub` items, so these thin wrappers — gated behind
/// the `bench` feature so they never exist in production builds — expose the
/// exact CPU-bound work (decode → orientation → resize → JPEG encode) for
/// before/after measurement. Errors are flattened to `String` to avoid leaking
/// `ThumbnailError` into the public API.
#[cfg(feature = "bench")]
impl ThumbnailService {
/// Render a single thumbnail size, returning the encoded JPEG bytes.
pub fn bench_render_thumbnail(data: &[u8], size: ThumbnailSize) -> Result<Vec<u8>, String> {
Self::render_thumbnail_from_data(data, size).map_err(|e| e.to_string())
}
/// Render all sizes in one decode (the upload-time path), returning each
/// size paired with its encoded byte length (output-size baseline).
pub fn bench_render_all(data: &[u8]) -> Result<Vec<(ThumbnailSize, usize)>, String> {
Self::render_all_thumbnails_from_data(data)
.map(|v| v.into_iter().map(|(s, b)| (s, b.len())).collect())
.map_err(|e| e.to_string())
}
}
/// Thumbnail service errors
#[derive(Debug, thiserror::Error)]
pub enum ThumbnailError {
+6
View File
@@ -12,6 +12,12 @@ pub mod interfaces;
#[cfg(integration_tests)]
pub mod integration_test_support;
// Phase 0 perf-benchmark support: deterministic image corpus generation/loading
// shared by `benches/thumbnails.rs` and `examples/bench_thumbnails_mem.rs`.
// Gated behind the `bench` feature so it adds nothing to normal builds.
#[cfg(feature = "bench")]
pub mod bench_support;
// Common public re-exports
pub use application::services::folder_service::FolderService;
pub use application::services::i18n_application_service::I18nApplicationService;