Reference

Artifacts and schemas

We save one directory per episode. For multiple conditions, you can find the episode directories under the run's artifacts directory. The resolved schema version defaults to "1".

Episode files

FileContents
config.yamlResolved configuration used for the condition
manifest.jsonRun and episode IDs, hashes, seed, condition, implementation IDs, selected interventions, and metadata
trajectory.jsonlNewline-delimited episode, transition, intervention, and evaluation entries
interventions.jsonconfigured specs and saved applications
outcomes.jsonEvaluation result, including outcome values, directions, and evidence
snapshots.jsoninitial and final snapshots
blobs/Content-addressed binary images or other portable binary values

When you repeat a command-line run, we archive the previous nonempty output directory before saving new files. ArtifactWriter.write itself requires an empty or absent directory and does not perform that archival step.

Trajectory entry types

typeFields
episodeepisode_id, seed, initial_snapshot
transitionepisode_id, transition
interventionepisode_id, step, application
evaluationepisode_id, evaluation, final_snapshot

The initial episode entry comes first, followed by agent-build intervention details when present. Each transition is followed by its intervention details. The evaluation entry comes last. Interventions also appear inside transitions and in interventions.json; do not sum all copies as separate applications.

Steps are zero-based action indices. Before and after snapshots contain state around each action. The first observation can already contain an observation or setup intervention.

Portable binary values

Image bytes are written to blobs with a hash-based filename. Portable image mappings retain media type and dimensions with an artifact reference and content hash. Replay reads the referenced bytes back from the episode directory.

Keep the directory and its blobs together when moving a run. A snapshot containing an image preserves an observation. The image does not supply enough information to restore an actual application.

An empty visual byte string is stored with artifact: null and sha256: null. The current replay loader assumes an image artifact path and fails on that empty-image representation. Text-only observations with visual: null and nonempty image blobs do not have that issue.

Pair and factorial summaries

paired_effects.json is a mapping from pair ID to paired identity, matching values, hashes, numeric differences, outcome directions, and evidence. Effects are treatment minus control for outcomes that are numeric in both episodes. We treat Boolean values as numeric in these summaries. String choices and missing values remain in episode outcomes but do not get numeric differences.

factorial_contrasts.json is a list of descriptive main effects for two-level factors. Each entry contains factor ID, first and second levels, outcome ID, means, second-minus-first difference, direction, and expected relation. Values are averaged over the other factor combinations and repetitions available in that run.

No file estimates interaction effects, standard errors, confidence intervals, or a population effect. Define those in your analysis if your design requires them.

Field reference

A required field has no default value. With a factory default, each instance has a separate empty mapping or list.

Event

Defined in harness/schema.py.

FieldPython typeDefault
sequenceintRequired
stepintRequired
kindstrRequired
actor_idstr | NoneNone
payloaddict[str, Any]Fresh {}
metadatadict[str, Any]Fresh {}

VisualObservation

Defined in harness/schema.py.

FieldPython typeDefault
databytesRequired
media_typestr"image/png"
widthint | NoneNone
heightint | NoneNone
metadatadict[str, Any]Fresh {}

Observation

Defined in harness/schema.py.

FieldPython typeDefault
stepintRequired
textstr | NoneNone
structureddict[str, Any]Fresh {}
visualVisualObservation | NoneNone
eventslist[Event]Fresh []
metadatadict[str, Any]Fresh {}

Action

Defined in harness/schema.py.

FieldPython typeDefault
kindstrRequired
argumentsdict[str, Any]Fresh {}
textstr | NoneNone
metadatadict[str, Any]Fresh {}

Snapshot

Defined in harness/schema.py.

FieldPython typeDefault
stepintRequired
statedict[str, Any]Required
observationObservation | NoneNone
metadatadict[str, Any]Fresh {}

InterventionApplication

Defined in harness/schema.py.

FieldPython typeDefault
intervention_idstrRequired
factorstrRequired
levelstrRequired
scopestrRequired
targetstrRequired
modalitystrRequired
hookstrRequired
orderintRequired
functionstrRequired
seedint | NoneNone
argumentsdict[str, Any]Fresh {}
selectordict[str, Any]Fresh {}
changed_fieldslist[str]Fresh []
before_hashstrRequired
after_hashstrRequired
held_fixed_hashstr | NoneNone
expected_relationstr | NoneNone
metadatadict[str, Any]Fresh {}

PolicyDecision

Defined in harness/schema.py.

FieldPython typeDefault
allowedboolRequired
rule_idstrRequired
reasonstr""
metadatadict[str, Any]Fresh {}

Transition

Defined in harness/schema.py.

FieldPython typeDefault
stepintRequired
observationObservation | NoneRequired
actionActionRequired
next_observationObservationRequired
before_snapshotSnapshotRequired
after_snapshotSnapshotRequired
rewardfloat | NoneNone
terminatedboolFalse
truncatedboolFalse
infodict[str, Any]Fresh {}
interventionslist[InterventionApplication]Fresh []

RunManifest

Defined in harness/artifacts.py.

FieldPython typeDefault
schema_versionstr"1"
run_idstrRequired
episode_idstrRequired
task_idstrRequired
fixture_versionstrRequired
fixture_hashstrRequired
canonical_initial_state_hashstrRequired
final_state_hashstrRequired
seedintRequired
condition_idstrRequired
pair_idstr | NoneNone
environment_idstrRequired
agent_idstrRequired
evaluator_idstrRequired
interventionslist[dict[str, Any]]Fresh []
metadatadict[str, Any]Fresh {}

PairManifest

Defined in harness/artifacts.py.

FieldPython typeDefault
schema_versionstr"1"
pair_idstrRequired
task_idstrRequired
fixture_hashstrRequired
seedintRequired
control_condition_idstrRequired
treatment_condition_idstrRequired
control_initial_state_hashstrRequired
treatment_initial_state_hashstrRequired
control_intervention_hashstrRequired
treatment_intervention_hashstrRequired
held_fixed_fieldslist[str]Fresh []
effectsdict[str, float]Fresh {}
outcome_directionsdict[str, str]Fresh {}
outcome_evidencedict[str, dict[str, list[dict[str, Any]]]]Fresh {}

OutcomeSpec

Defined in harness/evaluators.py.

FieldPython typeDefault
idstrRequired
directionstr"report"
functionstr | NoneNone
argumentsdict[str, Any]Fresh {}

Outcome

Defined in harness/evaluators.py.

FieldPython typeDefault
idstrRequired
valueAnyNone
directionstr"report"
evidencelist[dict[str, Any]]Fresh []
metadatadict[str, Any]Fresh {}

EvaluationResult

Defined in harness/evaluators.py.

FieldPython typeDefault
outcomeslist[Outcome]Required
metadatadict[str, Any]Fresh {}

FactorialContrast

Defined in harness/evaluators.py.

FieldPython typeDefault
factorstrRequired
low_levelstrRequired
high_levelstrRequired
outcome_idstrRequired
low_meanfloatRequired
high_meanfloatRequired
differencefloatRequired
directionstrRequired
expected_relationstrRequired

ObserverView

Defined in harness/replay.py.

FieldPython typeDefault
parent_episode_idstrRequired
pause_stepintRequired
transitionslist[Transition]Required
snapshotSnapshotRequired
pause_positionstr"after"
metadatadict[str, Any]Fresh {}
interventionslist[InterventionApplication]Fresh []

HandoffLineage

Defined in harness/replay.py.

FieldPython typeDefault
parent_episode_idstrRequired
parent_pause_stepintRequired
parent_snapshot_hashstrRequired
child_episode_idstrRequired

HandoffBranch

Defined in harness/replay.py.

FieldPython typeDefault
observationObservationRequired
lineageHandoffLineageRequired

PauseSpec

Defined in harness/replay.py.

FieldPython typeDefault
idstrRequired
kindstrRequired
stepint | NoneNone
progressfloat | NoneNone

Frame

Defined in harness/environments/driver.py.

FieldPython typeDefault
textstr""
structureddict[str, Any]Fresh {}
screenshotbytes | NoneNone
widthint | NoneNone
heightint | NoneNone
statedict[str, Any]Fresh {}
rewardfloat | NoneNone
terminatedboolFalse
truncatedboolFalse
infodict[str, Any]Fresh {}
Source files for this page

On this page