docs: update documentation to reflect 100% blob storage model

Rewrite documentation to match the new architecture where all
file metadata lives in PostgreSQL and content is stored as
content-addressed blobs via DedupService.

Updated files:
- internal-architecture.md: complete rewrite — new DB schema,
  blob repos (FolderDb, FileBlobRead/Write, TrashDb), updated
  DI container, service groups, architecture diagram, data flows
- file-system-safety.md: repurposed as storage-safety.md —
  covers PostgreSQL ACID guarantees + DedupService atomic writes
- caching-architecture.md: updated repo references to blob repos,
  removed write-behind cache section, updated upload/download flows
- trash-feature-summary.md: rewritten for soft-delete model
  (is_trashed flag, trash_items VIEW, TrashDbRepository)
- share-integration.md: clarified ShareFsRepository scope,
  updated DI snippet, added blob storage context note
- deduplication.md: updated DI snippet (dedup injected into repos)
- deployment.md: updated feature matrix (file storage requires DB)
- important-delta-sync-implementation.md: updated DI references

Removed legacy references: IdMappingPort, StorageMediator,
WriteBehindCache, FsFileRepository, FsFolderRepository,
TrashFsRepository, folder_ids.json, file_ids.json.
This commit is contained in:
Dionisio
2026-02-14 18:27:30 +01:00
parent 5d2bc36d74
commit e071841ec2
8 changed files with 1033 additions and 1115 deletions
+125 -156
View File
@@ -1,156 +1,125 @@
# 03 - File System Safety
OxiCloud ensures data integrity and durability during file operations through atomic writes, fsync, and directory synchronization. The goal: writes either complete fully or not at all, data reaches persistent storage, and the system recovers from crashes or power loss.
---
## The Problem: Buffered I/O
Standard filesystem operations use buffered I/O by default:
```rust
// This operation may not immediately persist to disk
fs::write(path, content)
```
When an application writes data, the OS typically:
1. Accepts the write into memory buffers
2. Acknowledges completion to the application
3. Schedules the actual disk write for later
A crash during that window means data loss -- the data exists only in memory buffers that haven't been flushed.
---
## OxiCloud's Approach
All safety mechanisms live in the **FileSystemUtils** service.
### Atomic Write Pattern
Files are written using write-then-rename:
```rust
/// Writes data to a file with fsync to ensure durability
/// Uses a safe atomic write pattern: write to temp file, fsync, rename
pub async fn atomic_write<P: AsRef<Path>>(path: P, contents: &[u8]) -> Result<(), IoError>
```
Steps:
1. Write to a temporary file in the same directory
2. Call `fsync` to ensure data is on disk
3. Atomically rename the temp file to the target file
4. Sync the parent directory to ensure the rename is persisted
### Directory Synchronization
```rust
/// Creates directories with fsync
pub async fn create_dir_with_sync<P: AsRef<Path>>(path: P) -> Result<(), IoError>
```
Directories are created, their entries persisted to disk, and parent directories synchronized too.
### Rename and Delete Operations
```rust
/// Renames a file or directory with proper syncing
pub async fn rename_with_sync<P: AsRef<Path>, Q: AsRef<Path>>(from: P, to: Q) -> Result<(), IoError>
/// Removes a file with directory syncing
pub async fn remove_file_with_sync<P: AsRef<Path>>(path: P) -> Result<(), IoError>
```
Both complete the operation itself, then update and sync the parent directory entry.
---
## Implementation Details
### fsync on Files
```rust
// Write file content
file.write_all(contents).await?;
// Ensure data is synced to disk
file.flush().await?;
file.sync_all().await?;
```
`sync_all()` instructs the OS to flush data and metadata to the physical storage device.
### fsync on Directories
```rust
// Sync a directory to ensure its contents (entries) are durable
async fn sync_directory<P: AsRef<Path>>(path: P) -> Result<(), IoError> {
let dir_file = OpenOptions::new().read(true).open(path).await?;
dir_file.sync_all().await
}
```
Required after any operation that modifies directory entries (create, rename, delete).
---
## Usage in the Codebase
### File Write Repository
```rust
// Write the file to disk using atomic write with fsync
tokio::time::timeout(
self.config.timeouts.file_write_timeout(),
FileSystemUtils::atomic_write(&abs_path, &content)
).await
```
### File Move Operations
```rust
// Move the file physically with fsync
time::timeout(
self.config.timeouts.file_timeout(),
FileSystemUtils::rename_with_sync(&old_abs_path, &new_abs_path)
).await
```
### Directory Creation
```rust
// Ensure the parent directory exists with proper syncing
self.ensure_parent_directory(&abs_path).await?;
// Implementation uses FileSystemUtils
async fn ensure_parent_directory(&self, abs_path: &PathBuf) -> FileRepositoryResult<()> {
if let Some(parent) = abs_path.parent() {
time::timeout(
self.config.timeouts.dir_timeout(),
FileSystemUtils::create_dir_with_sync(parent)
).await
}
}
```
---
## Benefits
1. **Data durability** -- critical data is synced to persistent storage
2. **Crash resilience** -- recovery from unexpected failures without data loss
3. **Consistency** -- file operations maintain a consistent filesystem state
4. **Atomic operations** -- file writes appear as all-or-nothing
---
## Performance Considerations
Syncing to disk costs more than buffered writes. OxiCloud mitigates this by:
1. Applying these measures only to critical operations
2. Using timeouts to prevent indefinite blocking
3. Implementing parallel processing for large files
The tradeoff favors safety for critical data while maintaining good performance for most operations.
# 03 - Storage Safety
OxiCloud ensures data integrity and durability through a combination of PostgreSQL transactional guarantees and atomic blob writes. The goal: writes either complete fully or not at all, data reaches persistent storage, and the system recovers from crashes or power loss.
---
## Storage Model
OxiCloud uses a **100% blob storage model**:
- **Metadata** (file names, folder hierarchy, sizes, MIME types, trash status) lives in **PostgreSQL** — protected by ACID transactions.
- **File content** is stored as content-addressed blobs via **DedupService** at `.blobs/{prefix}/{hash}.blob` — protected by atomic writes and fsync.
---
## PostgreSQL Safety (Metadata)
All file and folder metadata operations use PostgreSQL transactions:
- **Single-row operations** (INSERT, UPDATE, DELETE) are inherently atomic.
- **Multi-step operations** (e.g., move file: UPDATE folder_id + UPDATE path) use explicit transactions via `sqlx`.
- **Foreign key constraints** prevent orphaned records (e.g., files referencing non-existent folders).
- **Unique constraints** prevent duplicate names within the same parent folder.
- **Soft-delete** for trash (`is_trashed = TRUE`) preserves data until explicit permanent deletion.
The `storage.trash_items` VIEW provides a unified read interface over trashed files and folders without duplicating data.
---
## Blob Storage Safety (Content)
### DedupService Atomic Writes
**File:** `src/infrastructure/services/dedup_service.rs`
When storing file content, DedupService uses the following pattern:
1. **Hash computation** — SHA-256 hash of content determines the blob path
2. **Deduplication check** — if a blob with the same hash exists, only increment the reference counter (no write needed)
3. **Atomic write** — if new content:
- Write to a temporary file (`.blob.tmp`)
- Call `fsync` to ensure data reaches persistent storage
- Atomically rename temp file to final path (`.blobs/{prefix}/{hash}.blob`)
4. **Reference counting** — track how many files reference each blob
This ensures that a blob either fully exists or doesn't — no partial writes.
### FileSystemUtils
**File:** `src/infrastructure/services/file_system_utils.rs`
Low-level utilities used internally by DedupService and other infrastructure services:
```rust
/// Atomic write: temp file → fsync → rename
pub async fn atomic_write<P: AsRef<Path>>(path: P, contents: &[u8]) -> Result<(), IoError>
/// Directory creation with fsync
pub async fn create_dir_with_sync<P: AsRef<Path>>(path: P) -> Result<(), IoError>
/// Rename with directory sync
pub async fn rename_with_sync<P, Q>(from: P, to: Q) -> Result<(), IoError>
/// Delete with directory sync
pub async fn remove_file_with_sync<P: AsRef<Path>>(path: P) -> Result<(), IoError>
```
### fsync Guarantees
- `sync_all()` on written files ensures data and metadata reach the physical storage device
- Directory entries are synced after create/rename/delete operations
- Prevents data loss during crashes or power failures between OS buffer flush and disk write
---
## Transaction Flow: File Upload
```
1. DedupService.store_bytes(content)
→ Compute SHA-256 hash
→ Check if blob exists (dedup hit → increment ref, return hash)
→ Write to .blobs/{prefix}/{hash}.blob.tmp
→ fsync + rename → .blobs/{prefix}/{hash}.blob
2. FileBlobWriteRepository.save_file()
→ BEGIN TRANSACTION
→ INSERT INTO storage.files (name, folder_id, blob_hash, size, ...)
→ COMMIT
```
If step 1 fails, no metadata is written. If step 2 fails, the blob exists but is unreferenced (cleaned up by garbage collection). Data is never in an inconsistent state.
## Transaction Flow: File Deletion
```
1. FileBlobWriteRepository.delete_file_permanently()
→ BEGIN TRANSACTION
→ DELETE FROM storage.files WHERE id = $1 (captures blob_hash first)
→ COMMIT
2. DedupService.decrement_ref(blob_hash)
→ Decrement reference counter
→ If counter reaches 0, delete the blob file
```
If step 2 fails, an unreferenced blob may remain on disk (occupies space but is not a correctness issue). Future garbage collection can clean these up.
---
## Benefits
1. **ACID transactions** — metadata operations are atomic, consistent, isolated, and durable
2. **Content-addressable storage** — identical content is stored once, referenced by hash
3. **Crash resilience** — atomic blob writes + PostgreSQL WAL ensure recovery
4. **No partial writes** — temp file + rename pattern guarantees all-or-nothing
5. **Referential integrity** — foreign keys prevent orphaned metadata
---
## Performance Considerations
- PostgreSQL connection pooling (`sqlx::PgPool`) amortizes connection overhead
- Dedup hash computation is CPU-bound but avoids unnecessary disk writes for duplicate content
- Blob fsync adds latency vs. buffered writes, but ensures durability for critical user data
- Content cache (in-memory LRU) serves repeat reads without disk or DB access