feat(jobs): run the thumbnail migration at startup, by default
A migration nobody triggers never finishes. Scheduled ticks deliberately never pass `repair`, so a deployment whose operator never opens the admin panel re-imported the same sidecars forever and never drained the directory — and relying on operators to edit `.env` has the same failure mode one level up. `OXICLOUD_STARTUP_JOBS` dispatches named jobs once, in the background, after the scheduler is ready. Entries use the syntax operators already type at the trigger URL (`name?repair=true`), so the value is literally the request they would otherwise make by hand. It defaults to both migration jobs in repair mode, so an untouched deployment migrates and drains itself. That is a destructive default and a real exception to no-silent-auto-repair, so the guard it rests on had to get stronger: `verify_and_unlink` now compares CONTENT, not length. A blob of the right size and the wrong bytes used to pass — a key-mapping bug handing back another file's preview at the same length would have deleted the original and kept the impostor, and thumbnails cluster tightly enough in size for that to be a real coincidence. The readback streams from the backend with no cache in front, so it proves durability rather than that a write was acknowledged. Deletion of `.thumbnails/` is attempted first and only falls back to renaming it `.thumbnails.migrated` when `remove_dir` refuses because a non-sidecar file is inside (Finder's `.DS_Store`). Either way the directory stops existing, which lets the read-path probe go back to a single `stat` on the root instead of walking the size directories. Validation is fail-fast: an unknown job name or flag panics at boot. A silently dropped `?repare=true` would leave the job in discovery-only mode while the operator believed the tier was draining, surfacing months later as "the migration never finished" with nothing pointing at the config line. Interrupted runs resume. Boot recovery flips abandoned rows to Paused with their cursor, so `run_or_resume` continues rather than rescanning — a long migration completes across however many restarts it takes. That is a scoped exception to "we do not auto-resume": here somebody did ask, in configuration, and not having to ask again is the point. `StartupJob` holds a `JobRunArgs` rather than re-listing its four fields, so a fifth flag cannot be added to the scheduler and silently ignored in configuration. Jobs named here are ordinary registered jobs — visible in the panel, triggerable by hand, same runs and findings. Their rows now carry a `startup` object so an operator can see that a job deletes on every boot rather than only when someone clicks Run. Adds docs/config/thumbnail-migration.md: what runs on first boot, how to snapshot database and storage together beforehand, and how to verify afterwards with satellites_consistency plus backend_consistency ?deep=true. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -2835,6 +2835,108 @@ impl AppServiceFactory {
|
||||
registered
|
||||
);
|
||||
|
||||
// `OXICLOUD_STARTUP_JOBS` — dispatch each named job once, now.
|
||||
//
|
||||
// Exists for the migration jobs. Their scheduled ticks import but
|
||||
// never delete (`repair` defaults false, per no-silent-auto-repair),
|
||||
// so a deployment whose operator never opens the admin panel keeps
|
||||
// importing sidecars it already imported and never drains the
|
||||
// directory. Naming the job in configuration IS the deliberate
|
||||
// consent that rule asks for; it is simply given once, at boot,
|
||||
// rather than per run.
|
||||
//
|
||||
// Validated here, dispatched in the background:
|
||||
//
|
||||
// * Unknown names **panic**. The registry is fully populated at this
|
||||
// point, so a name that does not resolve is a typo or a rename, and
|
||||
// the failure mode of ignoring it is a migration that silently
|
||||
// never runs. Fail at boot, where the operator is watching.
|
||||
// * Dispatch is `tokio::spawn` — readiness must never wait on a job
|
||||
// that walks a filesystem for hours.
|
||||
// * Sequential within the task, not concurrent: these jobs contend
|
||||
// for the same directory and DB, and the exclusivity gate would
|
||||
// turn overlap into a skipped run rather than a queued one.
|
||||
// * Safe on every boot, including a crash loop: each is idempotent
|
||||
// and resumable, and once drained a run is a `read_dir` that
|
||||
// returns nothing.
|
||||
//
|
||||
// **Killed mid-run, this resumes from the cursor.** The boot
|
||||
// recovery sweep runs earlier in this function and flips every row
|
||||
// the dead process abandoned in `Running` to `Paused`, keeping its
|
||||
// cursor. `run_or_resume` then picks Resume over a fresh start, so
|
||||
// a job interrupted by a restart continues where it stopped rather
|
||||
// than rescanning from the beginning — and a long migration
|
||||
// completes across however many restarts it takes.
|
||||
//
|
||||
// That is a deliberate exception to `boot_recovery_sweep`'s "we do
|
||||
// not auto-resume; operators trigger the resume explicitly". The
|
||||
// rule exists so a restart never silently resumes work nobody
|
||||
// asked for. Here somebody did ask, in configuration, and the whole
|
||||
// point of the option is not having to ask again. The exception is
|
||||
// scoped to the named jobs; every other paused run still waits for
|
||||
// an operator.
|
||||
//
|
||||
// The resumed run keeps the flags it started with — `repair` and
|
||||
// `deep` are persisted to the run's `params` on the fresh open and
|
||||
// read back on resume — so editing the config mid-migration does
|
||||
// not retroactively change a run already in flight.
|
||||
if !self.config.startup_jobs.is_empty() {
|
||||
let mut planned = Vec::with_capacity(self.config.startup_jobs.len());
|
||||
for job in &self.config.startup_jobs {
|
||||
if app_state.core.job_registry.get(&job.name).await.is_none() {
|
||||
panic!(
|
||||
"OXICLOUD_STARTUP_JOBS names `{}`, which is not a registered job. \
|
||||
Check the spelling against GET /api/admin/jobs.",
|
||||
job.name
|
||||
);
|
||||
}
|
||||
planned.push(job.clone());
|
||||
}
|
||||
|
||||
let registry = app_state.core.job_registry.clone();
|
||||
tokio::spawn(async move {
|
||||
for job in planned {
|
||||
// Audited, not merely logged: a startup job may delete
|
||||
// files, and "who asked for this" must be answerable
|
||||
// afterwards. The answer is the configuration, which is
|
||||
// exactly what this line records.
|
||||
tracing::info!(
|
||||
target: "audit",
|
||||
event = "job.startup_trigger",
|
||||
job = %job.name,
|
||||
force = job.args.force,
|
||||
deep = job.args.deep,
|
||||
repair = job.args.repair,
|
||||
storage = ?job.args.storage,
|
||||
"👮🏻♂️ dispatching `{}` from OXICLOUD_STARTUP_JOBS",
|
||||
job.name,
|
||||
);
|
||||
match registry.trigger(&job.name, &job.args).await {
|
||||
Some(outcome) => tracing::info!(
|
||||
target: "oxicloud::scheduler",
|
||||
event = "job.startup_completed",
|
||||
job = %job.name,
|
||||
outcome = outcome.kind(),
|
||||
"startup job `{}` finished ({})",
|
||||
job.name,
|
||||
outcome.kind(),
|
||||
),
|
||||
// Unreachable — the name was resolved above, and
|
||||
// nothing unregisters. Logged rather than panicking
|
||||
// because this is a detached task by then.
|
||||
None => tracing::error!(
|
||||
target: "oxicloud::scheduler",
|
||||
event = "job.startup_vanished",
|
||||
job = %job.name,
|
||||
"startup job `{}` disappeared from the registry between \
|
||||
validation and dispatch",
|
||||
job.name,
|
||||
),
|
||||
}
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
Ok(app_state)
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user