Signed Experiment Campaigns¶
rle.reusable.campaigns turns a small declarative matrix into immutable,
source-pinned execution units. It is designed for benchmark sweeps where a
missing run must stay missing, two methods must use the same logical seeds, and
an interrupted job must never look complete.
Start with the checked-in examples/reusable/campaigns/smoke.json definition,
then materialize it from a clean checkout:
uv run python scripts/experiment_campaign.py materialize \
examples/reusable/campaigns/smoke.json \
--output exports/campaigns/smoke/resolved.json
The command records the current Git revision and refuses a dirty checkout by
default. --allow-dirty is available for exploratory work and marks a dirty
revision explicitly. A resolved manifest contains one signed run for every
environment, algorithm, and logical-seed combination. Construction, learner,
and evaluation seeds are derived in separate deterministic domains; algorithms
on the same environment and logical seed receive the same domain seeds.
Running One Unit¶
Scheduler array tasks can select a run by ID:
uv run python scripts/experiment_campaign.py list exports/campaigns/smoke/resolved.json
uv run python scripts/experiment_campaign.py show \
exports/campaigns/smoke/resolved.json \
smoke--example_env--ippo--seed-00
The project-owned launcher then uses the lifecycle API:
from rle.reusable.campaigns import (
campaign_suite_config,
finalize_experiment_campaign_run,
prepare_campaign_run,
resolve_campaign_run,
)
from rle.reusable.experiments import run_experiment_suite
run = resolve_campaign_run(manifest, run_id)
state = prepare_campaign_run(run)
if state != "complete":
result = run_experiment_suite(
campaign_suite_config(run),
project_algorithm_specs,
)
finalize_experiment_campaign_run(run, result)
prepare_campaign_run refuses unsigned files or metadata from a different
resolved run. Resolved configuration mappings are deeply immutable, and the
adapter rejects environment, common, algorithm, seed, or identity overrides.
Campaign suites also replace AlgorithmSpec.config defaults with the signed
algorithm configuration instead of merging undeclared defaults into the run.
The run hook receives every declared domain through context.seeds; the
learner-domain value remains available as context.seed.
Finalization hashes the run record, every tracked trainer artifact, and the complete per-run checkpoint directory. Status and reporting re-check the signature, source revision, file size, and SHA-256 digest. Artifact paths cannot escape the run root or traverse symbolic links, and an existing completion closure cannot be replaced with a smaller one.
For an intentional rerun, call prepare_campaign_run(run, restart=True). The
whole completed run directory is renamed atomically beneath
_archived_attempts/<run_id>/attempt_NNNN before a fresh run root is prepared;
restart never removes only complete.json or resumes the old trainer record.
Reporting And Archival¶
Generate a row-complete status bundle with JSON, Markdown, and optional CSV:
uv run python scripts/experiment_campaign.py report resolved.json \
--json-output status.json \
--markdown-output status.md \
--csv-output status.csv \
--require-complete
uv run python scripts/experiment_campaign.py inventory resolved.json \
--output artifact-inventory.json
--require-complete fails if any resolved unit is pending, partial, corrupt, or
has a changed artifact. It never substitutes zero for a missing result. The
artifact inventory is itself signed and is suitable for an archive index.
For paired comparisons, exact_paired_binomial_test covers binary outcomes,
exact_paired_wilcoxon_test covers up to 24 non-zero paired differences, and
holm_adjust_p_values applies family-wise error correction. Callers should
pre-register the outcome, direction, exclusions, and family before using these
generic calculations.
Scheduler Accounting¶
Write campaign_run_id=<signed-run-id> near the start of every retained Slurm
or PBS log. The accounting adapters join that marker to scheduler output rather
than guessing from job order:
uv run python scripts/experiment_campaign.py slurm-accounting resolved.json \
--log-root logs \
--accounting-input sacct.psv \
--site cluster-a \
--json-output accounting.json \
--markdown-output accounting.md
Slurm input is sacct --parsable2 data. PBS input is qstat -f -F json data.
Reports preserve attempt-level state, elapsed time, memory, CPUs, queue or
partition, node list, and signed campaign identity.
Adapting To Your Project¶
Copy the campaigns package, the CLI, and one definition. Keep environment
builders and algorithm registrations in the owning project, then translate a
resolved unit with campaign_suite_config. Add new seed domains when another
independent random stream matters; never reuse the learner seed implicitly.
Treat schema or scientific-config changes as a new campaign_revision, retain
the resolved manifest beside reports, and keep completion validation in the
archive or publication pipeline.
The SHA-256 fields used here are deterministic content identities and corruption checks, not keyed authentication. Load pickle-like trainer payloads only from trusted campaigns.
Feeding a large Slurm campaign¶
A feeder is useful when the campaign exceeds the user queue limit or needs checkpoint-aware retries. It replenishes the pending/running queue; Slurm decides when resources are available and starts jobs. It does not reserve spare CPUs or launch training on a frontend.
First provide a project-owned function your_project.campaign:run(unit).
It receives one validated resolved run, including the signed configuration,
seeds and export directory. It must either complete the run through the
existing campaign lifecycle API or return the artifact paths to finalize.
It may also return ExperimentSuiteResult, which the wrapper finalizes with finalize_experiment_campaign_run.
A successful process exit requires completion validation. The wrapper refuses
source drift, skips already complete units and leaves interrupted units
partial for checkpoint-aware recovery. The adapter must consume the signed
configuration rather than silently using notebook defaults.
This is a minimal artifact-only smoke adapter; substitute your trainer for real experiments:
from pathlib import Path
def run(unit):
artifact = Path(unit.export_root) / "result.json"
artifact.write_text('{"smoke": true}')
return [artifact]
Create submission.json as a JSON array of sbatch arguments, with one
argument per element and no sbatch executable:
[
"--partition=amd48",
"--cpus-per-task=2",
"--mem=8G",
"--output=/homes/oja24/logs/job_%j.out",
"/homes/oja24/py_cpu.sh",
"rl-engine/scripts/run_campaign.py",
"venv=.venv",
"manifest={manifest}",
"run_id={run_id}",
"entrypoint=your_project.campaign:run"
]
Paths, partition and resources are operational settings: adapt them to your
cluster and dedicated campaign checkout. The template accepts {manifest}
(an absolute path), {run_id}, {environment_id}, {algorithm_id},
{logical_seed_id} and {export_root}. Paths containing spaces remain one
argument. Array/cross-cluster submissions and feeder-managed comment flags
are rejected. Existing hyper launchers and notebooks are not automatically
campaign-aware; use the explicit wrapper and an adapter.
Review one replenishment poll, then start the feeder:
uv run python scripts/experiment_campaign.py feed exports/campaigns/resolved.json \
--submission-template submission.json --dry-run
uv run python scripts/experiment_campaign.py feed exports/campaigns/resolved.json \
--submission-template submission.json \
--queue-ceiling 190 --max-in-flight 32 --max-attempts 3
Keep the manifest and feeder state in ignored exports. Run from a checkout
matching the manifest's source revision; the DoC runner pulls its configured
branch, so pin that branch for the campaign. A source check inside
scripts/run_campaign.py rejects an advancing or dirty checkout before
calling the adapter.
The default state is resolved.feeder.json next to the manifest. Restart the
same command with the same state to resume. A flock prevents a second feeder
using that state. Do not feed the same campaign through different state files,
multiple clusters, or an older untagged launcher at the same time. Feeder-owned
jobs carry their signed identity in a Slurm comment, allowing reconciliation
when acceptance succeeded but the feeder crashed before saving the job ID.
Defaults are a 180-second poll and a 60-second retry cooldown. The user ceiling
counts other projects too, including individual array tasks. Rejected
submissions do not consume run retries. Invalid artifacts stop submission;
unknown accounting or uncertain acceptance blocks the affected run. An
uncertain submitting record needs manual reconciliation with scheduler/log
evidence, rather than deletion of attempt history. --once performs one poll;
--dry-run previews current queue room without submitting or creating state.
The feeder exits when every run is complete, or exits nonzero once all remaining
runs exhaust their accepted attempts. Ctrl-C stops the feeder without cancelling
its jobs. Use nohup on a frontend when you want feeding to survive logout.
This feeder supports Slurm. It neither archives signed artifacts nor starts itself for an existing campaign.