Files
Oxicloud/docs/delta-upload-protocol.md
T
Claude 44967da7f1 Delta-upload protocol: negotiate chunks by hash, upload only what changed
Phase 1 of the delta-sync plan, server side. The CDC store already shares
unchanged chunks between file versions after the bytes arrive; these three
stateless endpoints move that detection to the client, so editing a few
bytes of a large file uploads ~1 MiB instead of the whole file:

- POST /api/files/delta/negotiate — given the file's chunk hashes, answer
  which ones the caller must upload. User-scoped and purely advisory.
- PUT  /api/files/delta/chunks — missing chunks as [u32 BE len][bytes]
  frames (streaming parse, ≤1 MiB per frame, chunk_max_bytes per request).
  Every hash is recomputed server-side; chunks land as ref_count=0 orphans
  that a commit pins or the periodic GC sweeps — no session table.
- POST /api/files/delta/commit — pin one reference per distinct chunk with
  a single UPDATE…RETURNING restricted to chunks the caller is entitled
  to (reachable through their own non-trashed files, or unreferenced
  orphans); anything else returns 409 {still_missing} for the client to
  upload and retry. The pinned sequence is then RE-READ and the whole-file
  BLAKE3 recomputed before any manifest exists — a declared file_hash is
  never trusted, because a forged manifest would poison future whole-file
  dedup hits for other users. The manifest accounting is shared with the
  byte path (attach_manifest, extracted from store_from_stream); the file
  row is created (201) or its content swapped by file_id (200). Owners of
  the exact file_hash short-circuit to a pure reference bump.

Supporting pieces: GIN index on chunk_manifests.chunk_hashes (containment
probes were sequential scans), claimable/pin/release/store-loose/verify
primitives on DedupService, update-by-id with Update-permission AuthZ on
FileUploadService, per-caller rate limiter (240/min), audit events with
stable reasons (rate_limited, chunk_verification_failed,
file_hash_mismatch), OpenAPI + docs/delta-upload-protocol.md, framing
parser unit tests and a PG-gated integration suite covering the
entitlement matrix (owned/foreign/orphan/unknown), orphan registration
and the verification read.

Verified end-to-end against PostgreSQL 16 with a node client hashing via
the vendored WASM: a 24 MB file delta-committed in 96 fixed-size chunks;
a 3-byte edit then negotiated missing 1/96 and synced with 278 KB on the
wire vs 24 MB (98.9% saved), downloading byte-identical. A second user
probing the same chunks got nothing (negotiate: all missing; commit: 409
with all 96 still withheld); a forged file_hash returned 400 plus the
audit line; a commit referencing one never-uploaded chunk returned 409
naming exactly that hash; the GIN index serves containment probes
(Bitmap Index Scan) once the planner favors it.

https://claude.ai/code/session_01WdNenpnujNR2sc32XVvwfS
2026-06-11 14:34:02 +00:00

5.4 KiB
Raw Blame History

Delta-Upload Protocol

Upload only what changed. The server's dedup store already splits every file into content-defined chunks (FastCDC, 64 KB – 1 MiB, avg 256 KB, BLAKE3-addressed) and shares unchanged chunks between file versions — but a classic upload still transfers every byte just for the server to discard the known ones. This protocol moves the "which chunks are new?" question to the client, so unchanged bytes never cross the wire.

Editing a few bytes of a 500 MB file re-uploads ~1 MiB instead of 500 MB.

Who can use it

Any authenticated API client. The OxiCloud web frontend adopts it in a later phase; generic WebDAV/NextCloud clients cannot (their protocols have no delta concept) — they keep uploading full bytes, and the server keeps deduplicating those on write.

Chunk boundaries are the client's choice: matching the server's FastCDC parameters maximizes cross-version sharing, but any split with chunks of 1 byte … 1 MiB is valid — correctness is guaranteed by server-side verification, not by the chunking scheme.

The three steps

1. POST /api/files/delta/negotiate

{ "chunks": [ { "h": "<blake3-hex>", "s": 262144 }, … ] }

Response — the distinct chunk hashes the caller must upload, in first-occurrence order:

{ "missing": [ "<blake3-hex>", … ] }

The answer is user-scoped: a chunk counts as available only when one of the caller's own (non-trashed) files already references it. The endpoint is purely advisory — the commit re-checks entitlement atomically, so a stale or spoofed answer can never leak content.

2. PUT /api/files/delta/chunks

Body: application/octet-stream, a sequence of frames

[u32 length, big-endian][length bytes]  …repeated…
  • one frame per chunk, each 1 byte … 1 MiB (the CDC maximum),
  • whole request capped by OXICLOUD_CHUNK_MAX_BYTES (default 100 MB) — split larger deltas across requests.

The server recomputes BLAKE3 of every frame itself (a declared hash is never trusted for content addressing) and registers the chunks as unreferenced orphans (ref_count = 0). Response:

{ "received": [ { "h": "<server-computed>", "s": 262144 }, … ] }

Compare against your own hashes to catch corruption before committing. Abandoned uploads need no cleanup call: the periodic GC sweeps zero-reference chunks.

3. POST /api/files/delta/commit

{
  "file_hash": "<blake3-hex of the whole file>",
  "chunks": [ { "h": "…", "s": 262144 }, … ],   // full sequence, in order
  "name": "video.mp4", "folder_id": "<uuid>"     // create mode
  // — or —
  "file_id": "<uuid>"                            // update (replace content)
}

Server-side, in order:

  1. AuthZ — Create on the folder (create mode) or Update on the file (update mode); quota on the logical size.
  2. Pin — one atomic UPDATE … RETURNING takes a reference on every distinct chunk the caller is entitled to: chunks reachable through the caller's own files, or unreferenced orphans (the just-uploaded state). Anything else → 409 { "still_missing": […] }: upload exactly those and retry the same commit.
  3. Verify — the pinned sequence is re-read and the whole-file BLAKE3 recomputed. A mismatch releases the pins and returns 400 (and an audit event): the declared file_hash is never trusted, because a forged manifest would poison future whole-file dedup hits for other users uploading the genuine content.
  4. Attach — the manifest is inserted with the same accounting as the streaming byte path (a concurrent identical commit resolves via ON CONFLICT: the loser's references are released and it becomes a dedup hit).
  5. Row — the file is created (201, body = FileDto) or its content swapped (200).

If the caller already owns the exact file_hash, the commit short-circuits to a pure reference bump — chunks aren't even looked at (same as POST /api/files/by-hash).

Security model

  • No content oracle. Possession is proven per chunk: without bytes you can only claim what your own files already reference. Probing someone else's chunk hashes yields still_missing, indistinguishable from the hash never existing.
  • No manifest poisoning. file_hash and every chunk hash are recomputed server-side before becoming addressable.
  • Bounded resources. Per-frame cap 1 MiB, per-request cap OXICLOUD_CHUNK_MAX_BYTES, whole-file cap OXICLOUD_MAX_UPLOAD_SIZE, per-caller rate limit (240 delta requests/min), quota enforced at commit. Orphan chunks are GC-swept.
  • Audit. Rejections emit delta_upload.rejected with stable reason keys: rate_limited, chunk_verification_failed, file_hash_mismatch. AuthZ denials surface as the engine's standard authz.denied.

Error summary

Status Meaning Client action
400 malformed framing/hashes/sizes, or file_hash mismatch fix and retry from step 1
404 folder/file not found or not accessible —
409 {"still_missing": […]} PUT those chunks, retry the commit
429 rate limited back off
507 quota exceeded —

Cost notes

  • negotiate is one indexed query (GIN over manifest chunk arrays).
  • commit performs one sequential server-side read of the full logical file for verification — cheap on local backends, a full object read on S3/Azure. Still strictly cheaper than receiving the bytes, and the client's bandwidth saving is unaffected.