docs: update documentation to reflect 100% blob storage model
Rewrite documentation to match the new architecture where all file metadata lives in PostgreSQL and content is stored as content-addressed blobs via DedupService. Updated files: - internal-architecture.md: complete rewrite — new DB schema, blob repos (FolderDb, FileBlobRead/Write, TrashDb), updated DI container, service groups, architecture diagram, data flows - file-system-safety.md: repurposed as storage-safety.md — covers PostgreSQL ACID guarantees + DedupService atomic writes - caching-architecture.md: updated repo references to blob repos, removed write-behind cache section, updated upload/download flows - trash-feature-summary.md: rewritten for soft-delete model (is_trashed flag, trash_items VIEW, TrashDbRepository) - share-integration.md: clarified ShareFsRepository scope, updated DI snippet, added blob storage context note - deduplication.md: updated DI snippet (dedup injected into repos) - deployment.md: updated feature matrix (file storage requires DB) - important-delta-sync-implementation.md: updated DI references Removed legacy references: IdMappingPort, StorageMediator, WriteBehindCache, FsFileRepository, FsFolderRepository, TrashFsRepository, folder_ids.json, file_ids.json.
This commit is contained in:
+125
-156
@@ -1,156 +1,125 @@
|
||||
# 03 - File System Safety
|
||||
|
||||
OxiCloud ensures data integrity and durability during file operations through atomic writes, fsync, and directory synchronization. The goal: writes either complete fully or not at all, data reaches persistent storage, and the system recovers from crashes or power loss.
|
||||
|
||||
---
|
||||
|
||||
## The Problem: Buffered I/O
|
||||
|
||||
Standard filesystem operations use buffered I/O by default:
|
||||
|
||||
```rust
|
||||
// This operation may not immediately persist to disk
|
||||
fs::write(path, content)
|
||||
```
|
||||
|
||||
When an application writes data, the OS typically:
|
||||
|
||||
1. Accepts the write into memory buffers
|
||||
2. Acknowledges completion to the application
|
||||
3. Schedules the actual disk write for later
|
||||
|
||||
A crash during that window means data loss -- the data exists only in memory buffers that haven't been flushed.
|
||||
|
||||
---
|
||||
|
||||
## OxiCloud's Approach
|
||||
|
||||
All safety mechanisms live in the **FileSystemUtils** service.
|
||||
|
||||
### Atomic Write Pattern
|
||||
|
||||
Files are written using write-then-rename:
|
||||
|
||||
```rust
|
||||
/// Writes data to a file with fsync to ensure durability
|
||||
/// Uses a safe atomic write pattern: write to temp file, fsync, rename
|
||||
pub async fn atomic_write<P: AsRef<Path>>(path: P, contents: &[u8]) -> Result<(), IoError>
|
||||
```
|
||||
|
||||
Steps:
|
||||
1. Write to a temporary file in the same directory
|
||||
2. Call `fsync` to ensure data is on disk
|
||||
3. Atomically rename the temp file to the target file
|
||||
4. Sync the parent directory to ensure the rename is persisted
|
||||
|
||||
### Directory Synchronization
|
||||
|
||||
```rust
|
||||
/// Creates directories with fsync
|
||||
pub async fn create_dir_with_sync<P: AsRef<Path>>(path: P) -> Result<(), IoError>
|
||||
```
|
||||
|
||||
Directories are created, their entries persisted to disk, and parent directories synchronized too.
|
||||
|
||||
### Rename and Delete Operations
|
||||
|
||||
```rust
|
||||
/// Renames a file or directory with proper syncing
|
||||
pub async fn rename_with_sync<P: AsRef<Path>, Q: AsRef<Path>>(from: P, to: Q) -> Result<(), IoError>
|
||||
|
||||
/// Removes a file with directory syncing
|
||||
pub async fn remove_file_with_sync<P: AsRef<Path>>(path: P) -> Result<(), IoError>
|
||||
```
|
||||
|
||||
Both complete the operation itself, then update and sync the parent directory entry.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Details
|
||||
|
||||
### fsync on Files
|
||||
|
||||
```rust
|
||||
// Write file content
|
||||
file.write_all(contents).await?;
|
||||
|
||||
// Ensure data is synced to disk
|
||||
file.flush().await?;
|
||||
file.sync_all().await?;
|
||||
```
|
||||
|
||||
`sync_all()` instructs the OS to flush data and metadata to the physical storage device.
|
||||
|
||||
### fsync on Directories
|
||||
|
||||
```rust
|
||||
// Sync a directory to ensure its contents (entries) are durable
|
||||
async fn sync_directory<P: AsRef<Path>>(path: P) -> Result<(), IoError> {
|
||||
let dir_file = OpenOptions::new().read(true).open(path).await?;
|
||||
dir_file.sync_all().await
|
||||
}
|
||||
```
|
||||
|
||||
Required after any operation that modifies directory entries (create, rename, delete).
|
||||
|
||||
---
|
||||
|
||||
## Usage in the Codebase
|
||||
|
||||
### File Write Repository
|
||||
|
||||
```rust
|
||||
// Write the file to disk using atomic write with fsync
|
||||
tokio::time::timeout(
|
||||
self.config.timeouts.file_write_timeout(),
|
||||
FileSystemUtils::atomic_write(&abs_path, &content)
|
||||
).await
|
||||
```
|
||||
|
||||
### File Move Operations
|
||||
|
||||
```rust
|
||||
// Move the file physically with fsync
|
||||
time::timeout(
|
||||
self.config.timeouts.file_timeout(),
|
||||
FileSystemUtils::rename_with_sync(&old_abs_path, &new_abs_path)
|
||||
).await
|
||||
```
|
||||
|
||||
### Directory Creation
|
||||
|
||||
```rust
|
||||
// Ensure the parent directory exists with proper syncing
|
||||
self.ensure_parent_directory(&abs_path).await?;
|
||||
|
||||
// Implementation uses FileSystemUtils
|
||||
async fn ensure_parent_directory(&self, abs_path: &PathBuf) -> FileRepositoryResult<()> {
|
||||
if let Some(parent) = abs_path.parent() {
|
||||
time::timeout(
|
||||
self.config.timeouts.dir_timeout(),
|
||||
FileSystemUtils::create_dir_with_sync(parent)
|
||||
).await
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Benefits
|
||||
|
||||
1. **Data durability** -- critical data is synced to persistent storage
|
||||
2. **Crash resilience** -- recovery from unexpected failures without data loss
|
||||
3. **Consistency** -- file operations maintain a consistent filesystem state
|
||||
4. **Atomic operations** -- file writes appear as all-or-nothing
|
||||
|
||||
---
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
Syncing to disk costs more than buffered writes. OxiCloud mitigates this by:
|
||||
|
||||
1. Applying these measures only to critical operations
|
||||
2. Using timeouts to prevent indefinite blocking
|
||||
3. Implementing parallel processing for large files
|
||||
|
||||
The tradeoff favors safety for critical data while maintaining good performance for most operations.
|
||||
# 03 - Storage Safety
|
||||
|
||||
OxiCloud ensures data integrity and durability through a combination of PostgreSQL transactional guarantees and atomic blob writes. The goal: writes either complete fully or not at all, data reaches persistent storage, and the system recovers from crashes or power loss.
|
||||
|
||||
---
|
||||
|
||||
## Storage Model
|
||||
|
||||
OxiCloud uses a **100% blob storage model**:
|
||||
|
||||
- **Metadata** (file names, folder hierarchy, sizes, MIME types, trash status) lives in **PostgreSQL** — protected by ACID transactions.
|
||||
- **File content** is stored as content-addressed blobs via **DedupService** at `.blobs/{prefix}/{hash}.blob` — protected by atomic writes and fsync.
|
||||
|
||||
---
|
||||
|
||||
## PostgreSQL Safety (Metadata)
|
||||
|
||||
All file and folder metadata operations use PostgreSQL transactions:
|
||||
|
||||
- **Single-row operations** (INSERT, UPDATE, DELETE) are inherently atomic.
|
||||
- **Multi-step operations** (e.g., move file: UPDATE folder_id + UPDATE path) use explicit transactions via `sqlx`.
|
||||
- **Foreign key constraints** prevent orphaned records (e.g., files referencing non-existent folders).
|
||||
- **Unique constraints** prevent duplicate names within the same parent folder.
|
||||
- **Soft-delete** for trash (`is_trashed = TRUE`) preserves data until explicit permanent deletion.
|
||||
|
||||
The `storage.trash_items` VIEW provides a unified read interface over trashed files and folders without duplicating data.
|
||||
|
||||
---
|
||||
|
||||
## Blob Storage Safety (Content)
|
||||
|
||||
### DedupService Atomic Writes
|
||||
|
||||
**File:** `src/infrastructure/services/dedup_service.rs`
|
||||
|
||||
When storing file content, DedupService uses the following pattern:
|
||||
|
||||
1. **Hash computation** — SHA-256 hash of content determines the blob path
|
||||
2. **Deduplication check** — if a blob with the same hash exists, only increment the reference counter (no write needed)
|
||||
3. **Atomic write** — if new content:
|
||||
- Write to a temporary file (`.blob.tmp`)
|
||||
- Call `fsync` to ensure data reaches persistent storage
|
||||
- Atomically rename temp file to final path (`.blobs/{prefix}/{hash}.blob`)
|
||||
4. **Reference counting** — track how many files reference each blob
|
||||
|
||||
This ensures that a blob either fully exists or doesn't — no partial writes.
|
||||
|
||||
### FileSystemUtils
|
||||
|
||||
**File:** `src/infrastructure/services/file_system_utils.rs`
|
||||
|
||||
Low-level utilities used internally by DedupService and other infrastructure services:
|
||||
|
||||
```rust
|
||||
/// Atomic write: temp file → fsync → rename
|
||||
pub async fn atomic_write<P: AsRef<Path>>(path: P, contents: &[u8]) -> Result<(), IoError>
|
||||
|
||||
/// Directory creation with fsync
|
||||
pub async fn create_dir_with_sync<P: AsRef<Path>>(path: P) -> Result<(), IoError>
|
||||
|
||||
/// Rename with directory sync
|
||||
pub async fn rename_with_sync<P, Q>(from: P, to: Q) -> Result<(), IoError>
|
||||
|
||||
/// Delete with directory sync
|
||||
pub async fn remove_file_with_sync<P: AsRef<Path>>(path: P) -> Result<(), IoError>
|
||||
```
|
||||
|
||||
### fsync Guarantees
|
||||
|
||||
- `sync_all()` on written files ensures data and metadata reach the physical storage device
|
||||
- Directory entries are synced after create/rename/delete operations
|
||||
- Prevents data loss during crashes or power failures between OS buffer flush and disk write
|
||||
|
||||
---
|
||||
|
||||
## Transaction Flow: File Upload
|
||||
|
||||
```
|
||||
1. DedupService.store_bytes(content)
|
||||
→ Compute SHA-256 hash
|
||||
→ Check if blob exists (dedup hit → increment ref, return hash)
|
||||
→ Write to .blobs/{prefix}/{hash}.blob.tmp
|
||||
→ fsync + rename → .blobs/{prefix}/{hash}.blob
|
||||
|
||||
2. FileBlobWriteRepository.save_file()
|
||||
→ BEGIN TRANSACTION
|
||||
→ INSERT INTO storage.files (name, folder_id, blob_hash, size, ...)
|
||||
→ COMMIT
|
||||
```
|
||||
|
||||
If step 1 fails, no metadata is written. If step 2 fails, the blob exists but is unreferenced (cleaned up by garbage collection). Data is never in an inconsistent state.
|
||||
|
||||
## Transaction Flow: File Deletion
|
||||
|
||||
```
|
||||
1. FileBlobWriteRepository.delete_file_permanently()
|
||||
→ BEGIN TRANSACTION
|
||||
→ DELETE FROM storage.files WHERE id = $1 (captures blob_hash first)
|
||||
→ COMMIT
|
||||
|
||||
2. DedupService.decrement_ref(blob_hash)
|
||||
→ Decrement reference counter
|
||||
→ If counter reaches 0, delete the blob file
|
||||
```
|
||||
|
||||
If step 2 fails, an unreferenced blob may remain on disk (occupies space but is not a correctness issue). Future garbage collection can clean these up.
|
||||
|
||||
---
|
||||
|
||||
## Benefits
|
||||
|
||||
1. **ACID transactions** — metadata operations are atomic, consistent, isolated, and durable
|
||||
2. **Content-addressable storage** — identical content is stored once, referenced by hash
|
||||
3. **Crash resilience** — atomic blob writes + PostgreSQL WAL ensure recovery
|
||||
4. **No partial writes** — temp file + rename pattern guarantees all-or-nothing
|
||||
5. **Referential integrity** — foreign keys prevent orphaned metadata
|
||||
|
||||
---
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
- PostgreSQL connection pooling (`sqlx::PgPool`) amortizes connection overhead
|
||||
- Dedup hash computation is CPU-bound but avoids unnecessary disk writes for duplicate content
|
||||
- Blob fsync adds latency vs. buffered writes, but ensures durability for critical user data
|
||||
- Content cache (in-memory LRU) serves repeat reads without disk or DB access
|
||||
|
||||
Reference in New Issue
Block a user