Reference

Configuration

We use Hydra to load YAML configuration from conf/ and apply your command-line overrides. A task file begins with # @package _global_ when we set multiple configuration groups together.

Configuration groups

GroupPurpose
taskInstructions, starting data, actions, outcomes, interventions, and default run settings
environmentAdapter, backend, connection, and supported features
agentAgent specification and provider request parameters
designConditions, factor combinations, repetitions, matching, and ordering
experimentOptional named overrides; create conf/experiment/ as needed

We configure the Wikipedia research task and an OpenAI model by default. For that task, we use a paired design instead of the shared single-condition default. Read the composed settings instead of inferring effective values from conf/config.yaml alone.

Run fields

FieldShared value or behavior
seed0; each repetition adds its index
schema_version"1"
max_steps${task.max_steps}
progresstrue
experiment_id${task.id}
run_id${experiment_id}
episode_id${task.id}-${seed}; condition suffixes are added when needed
output_dir${hydra:runtime.output_dir}/artifacts
hydra.run.dirruns/${experiment_id}
hydra.sweep.dirruns/${experiment_id}-sweep
hydra.sweep.subdir${hydra.job.num}
hydra.job.chdirfalse

max_steps is the effective action budget passed to the episode runner. Its direct Python constructor defaults to 50; task configurations usually set 10 or their own budget.

Task fields

FieldMeaning
idTask identifier saved in artifacts
fixture_versionVersion for the starting setup and scoring definition
instructionGoal presented to the agent
fixtureOptions passed to environment reset
environmentTask-specific starting data and environment settings
modalitypruned_html, accessibility_tree, screenshot, text, or a list
max_stepsTask's default action budget
available_actionsAction kinds, argument examples, and descriptions
required_capabilitiesFeatures checked against the instantiated environment
actionsAction sequence for environment checks; does not set the LLM's actions
default_actionFallback action for the testing agent; unused by LLM agents
outcomesOutcome IDs, directions, function paths, and arguments
factorsIDs, levels, and optional expected relations for factorial designs
held_fixedDescriptive notes saved with the run
intervention.idTreatment identifier; control means no selected specs
intervention.specsIntervention specifications

Task fields are composed mappings rather than one strict task model. Referenced values and the constructed objects determine which fields a particular task needs.

Design fields

FieldMeaning
id, kindsingle, paired, or factorial
conditionsExplicit condition IDs; paired defaults to control then the task intervention
repetitionsPositive repetition count; default 1
shuffleShuffle materialized conditions; paired and factorial defaults are true
pair_onConfiguration paths used for matching
cross_factorsFactor definitions; factorial defaults to ${task.factors}
require_held_fixed_hash_matchPaired initial-state check, except for treatment specs using episode_start
require_factor_bindingsRequire configured and observed factor levels; true in the factorial profile
require_intervention_applicationOptional check that every selected spec changed a target; use Hydra + when adding an absent key

The shipped pair matching paths are task.id, task.intervention.id, and seed. Add a case ID path for a task that reuses its definition across cases. Numeric pair summaries require exactly two conditions.

Agent fields

agent.spec contains id, instructions, system_prompt, tools, memory_policy, scaffold, and config. scaffold._target_ identifies the implementation. accepts_spec defaults to true when the agent builder reads it.

Shared model settings select LiteLLMAgent and reference agent.chat_model_args. Shared configuration sets send_seed=false, json_mode=true, vision_detail=high, structured_actions=false, and max_observation_chars=8000. The modality follows task.modality. Shared action history retains six turns.

Provider arguments such as model, timeout, retry count, temperature, and token limits belong under agent.chat_model_args. They are passed through LiteLLM after removing None values. Provider support varies. See Configuring agents and observations.

Environment fields and variables

environment.adapter._target_ identifies the adapter. Nested backend and driver mappings supply construction settings. Browser configuration includes observation selection, setup and state-reader callbacks, and navigation behavior. Desktop profiles supply service endpoints or SDK lifecycle settings.

We check the operations implemented by the selected environment. You must implement an operation before adding it to the supported-feature list.

The shipped .env.example includes OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, AWS_REGION_NAME, ABXLAB_URL, OSWORLD_ENDPOINT, and WIKIPEDIA_URL. Integration generation can add NAME_URL or NAME_ENDPOINT variables. See Installing and configuring.

Source files for this page

On this page