Reference

Public Python API

Import the types and functions listed here from harness. Import environment implementations from harness.environments; import driver implementations from harness.environments.driver.

Exported functionPurpose
build_agentConstruct the agent selected by an AgentSpec
edit_agent_specApply configured edits before constructing the agent
materialize_designCreate the conditions and repetition assignments
factorial_main_effectsCompute descriptive differences for two-level factors
load_trace_bundleRead episodes and comparisons for the viewer

Observation and action types

ExportPurpose
ActionSubmitted action kind, arguments, text, and metadata
ObservationStep, text, structured fields, image, events, and metadata
VisualObservationImage bytes, media type, dimensions, and metadata
SnapshotState and observation at a step
TransitionAction with before and after observations, snapshots, flags, and interventions
EventOrdered backend event with optional actor and payload
PolicyDecisionAllowed flag, rule ID, reason, and metadata
InterventionApplicationRecorded edit, changed fields, hashes, and context

See Artifacts and schemas for field types and defaults.

Agents and execution

ExportPurpose and supported methods
AgentInterface with reset(seed) and act(observation)
AgentSpecAgent ID, instructions, system prompt, tools, memory, construction, and configuration
LiteLLMAgentModel-backed implementation with reset and act
EnvironmentEpisode environment interface
EpisodeResultInitial and final snapshots, trajectory, evaluation, and metadata
EpisodeRunnerrun(episode_id, seed=0, options=None) for one agent with an action budget
ScheduledEpisodeRunnerrun(episode_id, seed=0, options=None) for configured participant turns
TurnActor ID and participant ID for a scheduled turn
RulePolicyOraclePolicy decision function for configured denied action kinds

EpisodeRunner takes an environment, agent, evaluator, max_steps=50, and progress=true. ScheduledEpisodeRunner takes an environment, participant mapping, turn list, evaluator, and progress=true. Both return an EpisodeResult; neither writes artifacts automatically.

build_agent(spec) instantiates the configured agent. edit_agent_spec(spec, runtime, episode_id, task_id, seed=None, condition=None) returns the edited specification and application details before construction.

Designing and intervention runtime

ExportPurpose
FactorFactor ID, levels, and expected relation
ConditionCondition ID, intervention ID, factor assignments, expected relations, and repetition index
InterventionScopeEnvironment, agent, or observer scope
InterventionTargetObservation, state, message, affordance, observer view, or agent
InterventionModalityLanguage, visual, or structured edit
InterventionHookRuntime application point
InterventionContextEpisode, step, environment, agent, task, seed, and metadata
InterventionResultEdited value, reported changed fields, and metadata
InterventionSpecConfigured intervention definition
CounterfactualRuntimeOrdered application of matching specs

materialize_design(kind, intervention_id, factors=(), explicit_conditions=(), repetitions=1, shuffle=false, seed=0) returns Condition objects. kind is single, paired, or factorial. Factorial designs require factors. Condition IDs cannot contain path separators or equal . or ...

CounterfactualRuntime(specs=(), registry=None) accepts specs and an optional function registry. apply(hook, value, context) returns (edited_value, application_records). reset() resets application counts.

Evaluation

ExportPurpose
OutcomeSpecOutcome ID, direction, optional function, and arguments
OutcomeMeasured value, direction, evidence, and metadata
EvaluationResultOutcome list and metadata; by_id() returns an ID mapping
OutcomeEvaluatorevaluate(trajectory, state) evaluates configured outcomes
FactorialContrastLow and high levels, means, difference, direction, and expected relation

factorial_main_effects(evaluations, factors) expects assignment/evaluation pairs and two-level factors. It returns descriptive FactorialContrast values. paired_effects is available from harness.evaluators, and is not exported by the package root.

Saved data and viewer

ExportPurpose and supported methods
RunManifestIdentity, hashes, seed, condition, implementation IDs, and metadata
PairManifestPaired identities, hashes, differences, directions, and evidence
ArtifactWriterwrite(config, manifest, result) and store_blob(data, extension)

ArtifactWriter(output_dir) requires an absent or empty output directory for write. store_blob returns the blob's relative path, using a content hash in the filename.

load_trace_bundle(run_dir) returns the episodes and comparisons used by the trace viewer. It reads saved artifacts without running an environment or agent.

Replaying and handoff

ExportPurpose
ObserverViewMasked view at a pause, including position and observer interventions
PauseSpecPause ID, kind, optional step, and progress fraction
HandoffLineageParent episode, pause step, canonical snapshot hash, and child episode
HandoffBranchRestored observation and lineage
TrajectoryReplayReplay construction, pause selection, and supported restoration

TrajectoryReplay(episode_id, trajectory, runtime=None) stores an episode trajectory. from_event_log(path, runtime=None, episode_id=None, progress=false) loads one episode and restores portable image data from blobs.

Supported methods are pause(step, metadata=None), pause_before_action(step, metadata=None), materialize_pauses(specs), validate_factor_applications(views, assignments), and restore_handoff(view, environment, child_episode_id).

We use the original saved snapshots for restoration and omit internal state from observer views. See Replaying and handing off for a workflow and Environment support and restoration for limits.

Source files for this page

On this page