Skip to content

Enabling Guidelines

Guidelines are short, actionable recommendations Evolve extracts from agent conversations ("trajectories") and stores as guideline entities. This guide covers how guideline generation works in full Evolve (MCP server / CLI) and how to choose between the two available generation methods, standard and consistency.

This guide applies to full Evolve. It does not apply to Evolve Lite, where guideline extraction happens entirely inside the host agent's own reasoning (a prompt-driven skill) rather than through the LLM pipeline described here.

Two ways trajectories reach the guideline pipeline

Entry point When it runs
save_trajectory MCP tool Called directly by an MCP client (e.g. Claude Desktop, Claude Code) at the end of a conversation
evolve sync phoenix CLI Pulls previously-traced conversations out of Arize Phoenix in a batch — see Phoenix Sync, which requires traces already flowing in via Low-Code Tracing

Both entry points funnel into the same underlying pipeline and respect the same EVOLVE_GUIDELINES_MODE setting below.

Choosing a guideline generation mode

Set EVOLVE_GUIDELINES_MODE (or pass --guidelines-mode to evolve sync phoenix) to control which pipeline(s) run:

Mode Optimizes for What it does
standard (default) Correctness on a single run Single LLM pass over the trajectory; produces one guideline set.
consistency Reliability across repeated runs Scores agent decision steps for consistency, then a focused LLM pass produces guidelines for the inconsistent ones. Two interchangeable methods compute that score — see Choosing a consistency method below.
all Both Runs both pipelines and stores both sets of guidelines side by side.
# Standard guidelines (default) — no change needed
export EVOLVE_GUIDELINES_MODE=standard

# Consistency guidelines only
export EVOLVE_GUIDELINES_MODE=consistency

# Generate both
export EVOLVE_GUIDELINES_MODE=all

Or for a one-off Phoenix sync, set --guidelines-mode to standard, consistency, or all:

uv run evolve sync phoenix --guidelines-mode consistency

Choosing a consistency method

consistency mode has two interchangeable implementations, controlled by EVOLVE_CONSISTENCY_METHOD (or --consistency-method on evolve sync phoenix). This setting only has an effect when EVOLVE_GUIDELINES_MODE is consistency or all — it's silently ignored in standard mode.

Method How it estimates uncertainty Cost
fast (default) Asks the guideline-generation LLM to judge each step's stability itself, in the same call that produces guidelines — no resampling Same as standard mode: one LLM call per generated subtask, or one for the whole trajectory when it isn't segmented
accurate Resamples each decision step multiple times and measures how much the outcome varies across resamples Several extra LLM calls per trajectory
# Fast (default) — no change needed
export EVOLVE_CONSISTENCY_METHOD=fast

# Accurate — resample each step and measure actual variance
export EVOLVE_CONSISTENCY_METHOD=accurate

Or for a one-off sync:

uv run evolve sync phoenix --guidelines-mode consistency --consistency-method accurate

Use accurate when you want uncertainty estimated by actually observing variance across resamples, not the LLM's own self-assessment of confidence, or when you're comfortable paying for the extra resampling calls in exchange for a measured signal. Use fast (the default) when accurate's per-trajectory resampling cost is too expensive to run at the volume you need.

accurate has further tuning knobs — the single uncertainty threshold that decides what counts as "high" (high_uncertainty_threshold; the most uncertain steps at or above it are flagged, up to a cap of five, and when no step clears it the single most-uncertain non-zero step is flagged as merely "elevated"), and whether to skip generation entirely when nothing looks uncertain (skip_on_no_uncertainty) — defined alongside the resampling config below; they're advanced settings, not something most readers need on a first pass. fast has no equivalent tunables today: it relies entirely on the prompt instructing the LLM to return no guidelines when it judges every step confident, rather than a pre-call numeric skip gate.

Two separate bounds limit which steps can carry a marker, and a long trajectory can hit either one:

  • max_steps (15 by default) caps how many steps are scored at all. Resampling measures only the first max_steps scorable steps — scorable meaning the step can be faithfully resampled, which excludes malformed turns and, on trajectories with no OpenAI tool schema, tool calls with nothing to rebind. An unscored step has no uncertainty value, so it can never be flagged. Raise max_steps to score deeper into a trajectory, at proportionally more resampling calls.
  • Only steps at position 50 or below are rendered, and so only those can be marked. The prompt renders at most the first 50 assistant turns.

The second bound catches trajectories the first does not, because step positions count every assistant turn — including the unscorable ones that are never scored. Positions are therefore sparse rather than consecutive: 50 unscorable turns followed by 3 real ones yields scored steps at positions 51, 52 and 53, which are only the first three scorable steps and so well inside max_steps, yet all fall outside the render window.

When a trajectory's only uncertain steps fall outside either bound, accurate skips it rather than generating against a prompt that explains markers it does not contain — producing no guidelines instead of weakly-grounded ones.

accurate works on providers that can't return several completions at once. The cheap way to get k samples for a step is a single call asking for k completions, because the trajectory prompt is then billed as input once instead of k times. Not every platform/model combination will do that: some reject the request outright (anthropic/*, ollama/*, bedrock/* models addressed directly), some reject it at the API even though LiteLLM advertises support (groq/*), and an OpenAI-compatible gateway will often accept the request and quietly return a single completion — the case you hit when a gateway fronts a Claude deployment, since LiteLLM only sees openai and reports the parameter as supported. In every one of those cases resampling now falls back to fetching the shortfall as separate single-completion calls (capped at 5 samples on the loop route to avoid excessive prompt re-billing). Gateway short returns are cached per model per process so subsequent steps take the fallback route directly, running up to EVOLVE_CONSISTENCY_RESAMPLE_MAX_WORKERS calls in parallel (default 4). Two consequences worth budgeting for: on a fallback provider the trajectory prompt is billed as input once per loop sample rather than once per step, and if some of those calls fail even after retries, the step is scored on the samples that did land (with a warning naming the count) rather than failing the trajectory — below two samples there is no variance to measure, so that does raise.

Samples that record no decision are discarded, not scored. A reasoning model spends its token budget on hidden reasoning before emitting any content, so a resample that runs out returns a successful response carrying an empty content field. Counted as samples, k of those are k identical empties, which scores as perfect consistency under embedding metrics — a step the model could not answer would be read as one it is completely confident about. Such samples are therefore dropped, with a warning naming how many and why; a step left with none is scored as consistency undefined and excluded from the score card, rather than contributing zero uncertainty or false variance. Note that an empty content field is not by itself the signal — a successful tool-call response legitimately has one — so only samples with neither content nor tool calls are dropped.

The resampling behavior (sample count, per-step-type uncertainty metric) and the accurate-only tuning knobs above are all defined in a YAML config file shipped alongside the consistency pipeline (consistency_analyzer/agent_config.yaml); advanced users calling generate_consistency_guidelines() directly from Python can point it at a custom config via config_path=.

Mixed-provider Phoenix syncs and accurate resampling: resampling forwards EVOLVE_CUSTOM_LLM_PROVIDER for every step, regardless of which model the traced step actually used — it's treated as a single deployment-wide routing setting, not per-trace. This is fine when every trajectory in a sync comes from the same provider. If a namespace mixes traces from different providers (e.g. claude-* and gpt-* traces synced from the same Phoenix project) and EVOLVE_CUSTOM_LLM_PROVIDER resolves to openai — which it does by default whenever OPENAI_API_KEY or OPENAI_BASE_URL is set, even if you never set it explicitly — resampling forces every step through the OpenAI provider regardless of the traced model, and non-OpenAI steps fail or get misrouted. For mixed-provider syncs, either set EVOLVE_CUSTOM_LLM_PROVIDER to match the traces you're resampling and run accurate once per provider, or use fast (the default), which never resamples and has no provider-routing step to get wrong.

Verifying output

uv run evolve entities list <namespace> --type guideline

Each guideline's metadata.generation_method is "standard", "consistency" (accurate method), or "consistency-fast" (fast method), so you can tell which pipeline — and, for consistency, which method — produced it when running in all mode. See Guideline Provenance for the full metadata schema, including creation_mode (auto-mcp vs auto-phoenix vs manual).

See also

  • Configuration — model selection (EVOLVE_GUIDELINES_MODEL) and other environment variables
  • Low-Code Tracing — instrumenting your agent so traces reach Phoenix in the first place
  • Phoenix Sync — batch guideline generation from traced trajectories