Pipeline Directory-Contract Migration¶
This release deliberately has no runtime compatibility shim for the previous pipeline configuration. Keep old output directories intact and create new stage directories for the convention-based workflow.
Required Changes¶
- Give every physical system a semantic
system_id, such asacetateoracetate-contact. IDs must be unique, lowercase, file-safe, and match[a-z0-9][a-z0-9._-]*. Use the optionalsystem_nameonly for display. - Replace
prepare-referenceandevaluate-snapshotswithlabel-snapshots. Provide each system's topology, trajectory, CP2K MD and single-point inputs, and snapshot count explicitly. When isolated-atom energies are enabled, also provide one CP2K input per detected element insingle_atom_inputs; BFF no longer generates these inputs. Regenerate labels; old prepare/evaluate stage layouts are not loaded. - Replace
sample,analyze, andlgpfitwithsample-parameters,build-qoi-datasets, andfit-lgp. The Python names use underscores. - Rename the analysis
samplesection totraining_samples. Move routines out of reference-system entries into the top-levelroutineslist, and give each routine its applicablesystems. - Replace residue shortcuts with complete MDAnalysis selections. RDF routines
require
group_aandgroup_band now emit one curve per atom type ingroup_a. Hydrogen-bond routines requireselectionandwater_selection; donor hydrogens and both interaction directions are inferred automatically. - Remove
loaderfrom custom QoI routines. A callable withinputsis file-based; a callable withoutinputsis trajectory-based. Removerun.gc_collectandrun.maxtasksperchild;run.in_memoryis the only analysis runtime option. - In fit-LGP configs, rename the
lgpfitoptions section tofit. - Replace learning
restartand individual artifact paths withmcmc.resume,output.overwrite, and oneoutput.directory. Learning now owns fixedoutputs/andplots/children below that root. The pluraloutputs/replaces the earlier singular directory. - Regenerate QoI datasets and
.lgpmodels. RDF normalization/PBC handling changed, and model reuse now verifies an exact dataset fingerprint. - Rerun
bff buildto create each system's stablereference/topology.topandreference/coordinates.gro. Configurebuild-qoi-datasetswith those files and the external MLIP trajectory as separate explicit inputs. - Validation may now consume
outputs/posterior.ptdirectly. Posterior mode rejects an externalspecspath and writes the artifact's embedded specifications into the validation campaign. Keepparametersplusspecsonly for explicitly exported YAML samples.
Identity and Paths¶
Training and reference system sets must contain the same unique IDs. YAML list
order is irrelevant: QoI construction pairs records by ID. Build files follow
documented fixed names under systems/<system_id>/; label-snapshots inputs are
explicit. Paths in samples.yaml are relative to that file, while user paths
remain relative to their config.
specs.yaml remains limited to ordered parameter bounds and compiled charge
constraints. It does not carry system, path, scheduler, or learning metadata.
No schema_version or replacement version field was introduced.
Learning Outputs¶
New learning runs reject existing stage-owned artifacts unless
output.overwrite: true is explicit. Resuming requires mcmc.resume: true and
a compatible outputs/mcmc.ckpt plus a matching outputs/specs.yaml; resume
and overwrite cannot be combined.
Only the known learning artifacts are replaced by overwrite, so unrelated
files in the output root are preserved.
Existing posterior files can still be loaded read-only where their contents are compatible, but old checkpoint directories are not migrated in place.