feat(jobs): jobs declare their own run parameters

`JobRunArgs` was a fixed struct — `force`, `deep`, `storage`, `repair` —
and six places hardcoded that same list: the engine's persist/restore,
the trigger endpoint's query type, the OXICLOUD_STARTUP_JOBS parser, the
frontend API wrapper, the panel's checkboxes, and `StartupTrigger` on
the wire.

Two costs. Adding a parameter meant editing all six, and forgetting one
dropped it silently — most damagingly in persist/restore, where a
resumed run lost it and a `?repair=true` migration came back as
discovery-only after a restart. And the panel offered the same knobs on
every job: only two jobs read `deep`, six read `repair`, so most of
those controls did nothing with no way to tell which.

Now `JobHandler::parameters()` returns `&'static [JobParam]` — name,
type (boolean/string/number), default, and the job's own description of
what it does. `JobRunArgs` holds a map keyed by those names.

Everything reads the declaration:

* `run_or_resume` iterates it to persist and restore, replacing
  `const FLAGS` plus a `storage` special case. `storage` stops being
  special — it was the one Option<String> among three bools.
* `dispatch` normalises every run against it, which is what makes "a
  handler sees its declared parameters with their declared defaults"
  true rather than usual. The periodic tick passes an empty
  `JobRunArgs::default()`, so a `default: true` parameter would
  otherwise read false on every scheduled run.
* The trigger endpoint takes free-form query params and rejects
  undeclared ones with a 400 naming the real set, instead of ignoring
  them.
* OXICLOUD_STARTUP_JOBS keeps raw pairs (config is parsed before the
  registry exists) and validates at dispatch, where the error can name
  the job's actual parameters. Still a boot panic, same as an unknown
  job name — a typo'd `?repare=true` must not leave a migration
  importing forever in discovery mode.
* `JobSummary.parameters` carries it to the panel, whose `supportsDeep`
  was a hardcoded name allowlist (`consistency_batch ||
  backend_consistency`). A job gaining a deep mode needed a frontend
  release; one losing it left a button that silently did nothing. The
  menu now renders from the declaration, so a newly-declared boolean
  appears with no frontend change.

Three consistency tenants were hand-rolling persist-on-fresh /
restore-on-resume for their own flag, under the same `params` key the
engine already used. Deleted — they read `args.get_bool(…)` now.

Fresh runs also filter to the declaration. `consistency_batch` forwards
its args verbatim to sub-jobs, so a tenant's `params` row could grow
`deep` with no deep mode, and the run-detail view would claim a mode the
job never had.

Two things found while wiring it, both worth knowing:

`RecoverableAdapter` bridges the two traits, and `parameters` has to be
forwarded there or the registry sees `&[]`. Both traits have defaults,
so omitting it compiled cleanly — and the trigger endpoint then rejected
`?repair=true` on the very jobs that declare it, with
OXICLOUD_STARTUP_JOBS panicking at boot. Now covered by
`adapter_forwards_job_metadata_from_inner_handler`.

`TriggerJobQuery` was briefly a newtype over the map. `serde_urlencoded`
cannot deserialize a newtype struct at the top level, so axum's `Query`
rejected EVERY trigger with a 400 — even one with no query string —
before the handler ran. It reads exactly like the new validation
rejecting something, which sent the first diagnosis to the wrong layer.
Now covered by `trigger_query_extracts_from_every_url_shape`.

Wire names are a compatibility surface: `params` rows are keyed by them
and the panel switches on them, so a rename breaks existing run history
the same way renaming a `Mutates` variant does. The JSON shape is pinned
in `snapshot_carries_job_metadata`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Edouard Vanbelle
2026-09-07 12:30:40 +02:00
parent ff286f8159
commit a4101743e0
23 changed files with 1242 additions and 397 deletions
+35 -18
View File
@@ -2922,35 +2922,54 @@ impl AppServiceFactory {
if !self.config.startup_jobs.is_empty() {
let mut planned = Vec::with_capacity(self.config.startup_jobs.len());
for job in &self.config.startup_jobs {
if app_state.core.job_registry.get(&job.name).await.is_none() {
let Some(declared) = app_state.core.job_registry.parameters_of(&job.name).await
else {
panic!(
"OXICLOUD_STARTUP_JOBS names `{}`, which is not a registered job. \
Check the spelling against GET /api/admin/jobs.",
job.name
);
}
planned.push(job.clone());
};
// Same fail-fast rule as the unknown-name panic above, and
// for the same reason: a typo'd `?repare=true` would leave
// a migration importing forever in discovery mode while the
// operator believed the tier was draining. The declaration
// is only reachable here, after the registry is built —
// config parsing kept the pairs untyped.
let args = crate::infrastructure::scheduler::JobRunArgs::from_declared(
declared,
job.raw_params.iter().map(|(k, v)| (k.as_str(), v.as_str())),
)
.unwrap_or_else(|e| {
panic!("OXICLOUD_STARTUP_JOBS entry `{}`: {e}", job.name);
});
planned.push((job.name.clone(), args));
}
let registry = app_state.core.job_registry.clone();
tokio::spawn(async move {
for job in planned {
for (job_name, args) in planned {
// Audited, not merely logged: a startup job may delete
// files, and "who asked for this" must be answerable
// afterwards. The answer is the configuration, which is
// exactly what this line records.
//
// Rendered from the parsed args rather than naming each
// parameter, so a job growing one cannot end up
// dispatched with something the audit trail omits.
let params_desc = args
.iter()
.filter_map(|(k, v)| v.to_param_string().map(|s| format!("{k}={s}")))
.collect::<Vec<_>>()
.join(", ");
tracing::info!(
target: "audit",
event = "job.startup_trigger",
job = %job.name,
force = job.args.force,
deep = job.args.deep,
repair = job.args.repair,
storage = ?job.args.storage,
"👮🏻‍♂️ dispatching `{}` from OXICLOUD_STARTUP_JOBS",
job.name,
job = %job_name,
params = %params_desc,
"👮🏻‍♂️ dispatching `{job_name}` from OXICLOUD_STARTUP_JOBS ({params_desc})",
);
match registry.trigger(&job.name, &job.args).await {
match registry.trigger(&job_name, &args).await {
// Debug, not info. The engine already logs every
// dispatch as `job.run` with the outcome and timing —
// that is the point of routing through `trigger`
@@ -2962,10 +2981,9 @@ impl AppServiceFactory {
Some(outcome) => tracing::debug!(
target: "oxicloud::scheduler",
event = "job.startup_completed",
job = %job.name,
job = %job_name,
outcome = outcome.kind(),
"startup job `{}` finished ({})",
job.name,
"startup job `{job_name}` finished ({})",
outcome.kind(),
),
// Unreachable — the name was resolved above, and
@@ -2974,10 +2992,9 @@ impl AppServiceFactory {
None => tracing::error!(
target: "oxicloud::scheduler",
event = "job.startup_vanished",
job = %job.name,
"startup job `{}` disappeared from the registry between \
job = %job_name,
"startup job `{job_name}` disappeared from the registry between \
validation and dispatch",
job.name,
),
}
}