feat(storage): add storage.file_attached_blobs, the file-keyed half

Step 9 of docs/plan/derived-blobs.md. content_derived_blobs holds bytes
that are a pure function of a file's content, so they are keyed by that
content and shared by every file holding it. This table holds the
opposite: bytes a user supplied or chose, which must never be shared
across files. The key is what enforces it.

That difference is a security boundary, not a modelling preference. A
content-keyed client preview would let user A upload a file plus a
preview that misrepresents it; when user B later uploads the same bytes,
dedup matches and B is served A's preview. Content-keying is only safe
when the server can derive the bytes — there is nothing to poison,
because the same input yields the same output for everyone.

Required now rather than deferred: the SPA already generates and PUTs
previews for PDFs, and there is no server-side regeneration path for
them, so the sidecar migration has nowhere else to put those bytes.

uploaded_by is NOT NULL with no foreign key, per the provenance
convention rather than the plan's sketch. A FK with ON DELETE SET NULL
discards the audit trail exactly when it matters, and without an
ON DELETE clause it would block deleting a user outright. Deleting the
uploader must not rewrite history.

FileAttachedReferenceSource is registered in built_in_registry before
anything writes to the table, so dedup_gc's reap predicate already knows
it exists — otherwise the first sweep after the first attachment would
delete it. Manifest level only, like the derived source: these blobs are
almost always single-chunk, so contributing at chunk level would
double-count against the aliased hash.

copy_file_satellites gains one arm: attachments are DUPLICATED, since
the key is file_id and the copy is a different file, with uploaded_by
carried over — the person who supplied the bytes did not change because
someone copied the file. Each duplicate takes its own reference, so the
bytes stay deduplicated while the mapping does not.

Both golden SQL tests updated: the new fragment lands inside the reap
predicate's NOT(...) group and as a summed term in the manifest
recompute. Verified on a scratch PG with every migration applied — the
attachment duplicates to 2 rows holding 2 references with provenance
intact, while the content-keyed thumbnail stays 1 row reachable from
both files.
This commit is contained in:
Edouard Vanbelle
2026-08-25 07:21:32 +02:00
parent cf21ee2f8f
commit fac82fea23
4 changed files with 279 additions and 3 deletions
@@ -29,6 +29,7 @@ use crate::domain::errors::DomainError;
const FILES_ALIAS: &str = "cnt_f";
const MANIFEST_ALIAS: &str = "cnt_m";
const DERIVED_ALIAS: &str = "cnt_d";
const ATTACHED_ALIAS: &str = "cnt_a";
/// Fragment for [`FilesReferenceSource`], as a free function so the SQL
/// shape can be tested without constructing a pool — it is a property of
@@ -117,6 +118,33 @@ fn content_derived_exists_sql(level: RefLevel, outer_hash_expr: &str) -> Option<
}
}
/// Fragment for [`FileAttachedReferenceSource`].
///
/// **Manifest level only**, for the same reason as the derived source: an
/// attached artifact's `blob_hash` names a Blob, never a chunk, and these are
/// almost always single-chunk — so contributing at the chunk level would
/// double-count against the aliased hash.
fn file_attached_ref_sql(level: RefLevel, outer_hash_expr: &str) -> Option<String> {
match level {
RefLevel::Chunk => None,
RefLevel::Manifest => Some(format!(
"(SELECT COUNT(*) FROM storage.file_attached_blobs {ATTACHED_ALIAS} \
WHERE {ATTACHED_ALIAS}.blob_hash = {outer_hash_expr})"
)),
}
}
/// Short-circuiting existence form, used by `dedup_gc`'s reap predicate.
fn file_attached_exists_sql(level: RefLevel, outer_hash_expr: &str) -> Option<String> {
match level {
RefLevel::Chunk => None,
RefLevel::Manifest => Some(format!(
"EXISTS (SELECT 1 FROM storage.file_attached_blobs {ATTACHED_ALIAS} \
WHERE {ATTACHED_ALIAS}.blob_hash = {outer_hash_expr})"
)),
}
}
/// Every built-in blob-reference source, in one place.
///
/// THE definition of "what references a blob". `DedupService::new` uses it
@@ -131,7 +159,10 @@ pub fn built_in_registry(pool: Arc<PgPool>) -> BlobReferenceRegistry {
// Registered before anything writes a derived blob: dedup_gc's reap
// predicate must already know this table exists, or the first sweep
// after the first thumbnail deletes it.
registry.register(Arc::new(ContentDerivedReferenceSource::new(pool)));
registry.register(Arc::new(ContentDerivedReferenceSource::new(pool.clone())));
// Same rule as above: registered before the first attachment is written,
// so dedup_gc's reap predicate already knows the table exists.
registry.register(Arc::new(FileAttachedReferenceSource::new(pool)));
registry
}
@@ -388,6 +419,92 @@ impl BlobReferenceSource for ContentDerivedReferenceSource {
}
}
// ─── storage.file_attached_blobs ─────────────────────────────────────────
/// References held by `storage.file_attached_blobs.blob_hash` — bytes a user
/// supplied for one specific file.
///
/// Structurally the twin of [`ContentDerivedReferenceSource`]: same level,
/// same shape, different table. The difference that matters is upstream — the
/// row is keyed by `file_id` rather than by content, so the same bytes
/// attached to two files are two rows and therefore two references. Dedup
/// still applies to the bytes; what must not be shared is the mapping.
///
/// `file_id` is deliberately not a reference at this layer: it is an
/// `ON DELETE CASCADE` foreign key, so the row disappears with the file, and
/// the blob reference it held is released by the owning service's
/// `on_file_deleted` hook.
pub struct FileAttachedReferenceSource {
pool: Arc<PgPool>,
}
impl FileAttachedReferenceSource {
pub fn new(pool: Arc<PgPool>) -> Self {
Self { pool }
}
}
#[async_trait]
impl BlobReferenceSource for FileAttachedReferenceSource {
fn source_name(&self) -> &'static str {
"file_attached"
}
fn ref_count_sql(&self, level: RefLevel, outer_hash_expr: &str) -> Option<String> {
file_attached_ref_sql(level, outer_hash_expr)
}
fn ref_exists_sql(&self, level: RefLevel, outer_hash_expr: &str) -> Option<String> {
file_attached_exists_sql(level, outer_hash_expr)
}
async fn count_references(&self, blob_hash: &str) -> Result<u64, DomainError> {
let n: i64 = sqlx::query_scalar(
"SELECT COUNT(*) FROM storage.file_attached_blobs WHERE blob_hash = $1",
)
.bind(blob_hash)
.fetch_one(self.pool.as_ref())
.await
.map_err(|e| {
DomainError::internal_error("BlobRefSource", format!("attached count: {e}"))
})?;
Ok(n.max(0) as u64)
}
async fn list_referenced_blobs(
&self,
cursor: Option<Vec<u8>>,
limit: usize,
) -> Result<(Vec<String>, Option<Vec<u8>>), DomainError> {
// Paged by `blob_hash`, same as the derived source: it IS the value
// returned, and DISTINCT collapses one Blob attached to several files.
let after: Option<String> = match cursor {
Some(bytes) => Some(String::from_utf8(bytes).map_err(|e| {
DomainError::internal_error("BlobRefSource", format!("bad attached cursor: {e}"))
})?),
None => None,
};
let rows: Vec<(String,)> = sqlx::query_as(
"SELECT DISTINCT blob_hash FROM storage.file_attached_blobs
WHERE ($1::text IS NULL OR blob_hash > $1)
ORDER BY blob_hash
LIMIT $2",
)
.bind(after)
.bind(limit as i64)
.fetch_all(self.pool.as_ref())
.await
.map_err(|e| DomainError::internal_error("BlobRefSource", format!("attached page: {e}")))?;
let next = rows
.last()
.map(|(h,)| h.clone().into_bytes())
.filter(|_| rows.len() == limit);
Ok((rows.into_iter().map(|(h,)| h).collect(), next))
}
}
#[cfg(test)]
mod tests {
use super::*;
+2 -1
View File
@@ -3550,7 +3550,8 @@ mod tests {
FROM storage.chunk_manifests m
WHERE m.ref_count <= 0
OR NOT (EXISTS (SELECT 1 FROM storage.files cnt_f WHERE cnt_f.blob_hash = m.file_hash)
OR EXISTS (SELECT 1 FROM storage.content_derived_blobs cnt_d WHERE cnt_d.blob_hash = m.file_hash))
OR EXISTS (SELECT 1 FROM storage.content_derived_blobs cnt_d WHERE cnt_d.blob_hash = m.file_hash)
OR EXISTS (SELECT 1 FROM storage.file_attached_blobs cnt_a WHERE cnt_a.blob_hash = m.file_hash))
LIMIT $1
)
RETURNING file_hash, chunk_hashes, total_size"#;
@@ -432,7 +432,8 @@ mod tests {
m.chunk_count AS chunk_count,
((SELECT COUNT(*) FROM storage.files cnt_f
WHERE cnt_f.blob_hash = m.file_hash)
+ (SELECT COUNT(*) FROM storage.content_derived_blobs cnt_d WHERE cnt_d.blob_hash = m.file_hash))::bigint AS actual_ref_count
+ (SELECT COUNT(*) FROM storage.content_derived_blobs cnt_d WHERE cnt_d.blob_hash = m.file_hash)
+ (SELECT COUNT(*) FROM storage.file_attached_blobs cnt_a WHERE cnt_a.blob_hash = m.file_hash))::bigint AS actual_ref_count
FROM storage.chunk_manifests m
WHERE ($1::text IS NULL OR m.file_hash > $1)
ORDER BY m.file_hash