Evaluating Autonomous Agents Evaluating Autonomous Agents
  • Home
  • Lectures
    • Day 1 · LLM Evaluation Foundations
    • Day 2 · LLM-as-a-Judge, Calibration, and OTel
    • Day 3 · Agent Harness and MCP Mocking
    • Day 4 · AWS CI/CD and Red Teaming
  • Notebooks
    • Day 1 · Deterministic and RAG Evals
    • Day 2 · Judge Calibration and OTel Traces
    • Day 3 · Building a Custom Agent Harness
    • Day 4 · Bedrock Evaluations and CI Gating

    • Day 1 · Solutions
    • Day 2 · Solutions
    • Day 3 · Solutions
    • Day 4 · Solutions
  • Slides
    • Workshop Kickoff
    • Day 1 · Why Your AI Agent Fails in Production
    • Day 2 · When a Model Grades a Model
    • Day 3 · The Model Proposes, the Harness Decides
    • Day 4 · Evals That Block a Merge
  • Teach

Evaluating Autonomous Agents

Systems, Harnesses & AWS Production CI/CD

Genial Labs · a four-day workshop

Evaluating Autonomous Agents

Systems, Harnesses & AWS Production CI/CD

An agent that works in a demo is not evidence that it works. Agents break in the harness and in the environment, and the final answer often still looks fine. This workshop builds the evals that catch those failures, then wires them into a CI gate on AWS. One running example, Stockroom, an inventory and order-support agent; every claim about a failure mode is demonstrated by a test or a notebook cell you can run offline.

Start with Day 1 Kickoff slides Open the notebooks Instructor guide

4 days 4 lectures, 8 notebooks 5 tools, 4 seeded weaknesses 50 golden cases 0 AWS credentials needed

Agent = Model + Harness The model proposes; the harness decides what is allowed to happen. Most production failures live in the harness and the environment, so that is what the eval suite exercises first.

Why single-prompt evals are not enough

A single-prompt eval scores the final answer and nothing else. The Stockroom suite scores tool selection, tool arguments, trajectories, guards, context compaction and injection handling, and only then the answer.

Where failures live The model

Wrong tool, wrong arguments, an answer that is not grounded in what the tools returned. The part single-prompt evals already see.

Where most of them live The harness

Tool-call formatting, loop control, context truncation, state drift, unquarantined tool output. The agent loop you wrote, or the one your framework wrote for you.

Where the rest live The environment

Tools, data and permissions: a misleading tool description, an oversized payload, a shard that is down, a policy document with an instruction planted in it.

flowchart LR
  U[user query] --> H[Harness state machine]
  H -->|Converse API or FakeBedrockClient| M[Model]
  M -->|tool calls| H
  H -->|JSON-schema validation, guards, quarantine| T[Tools: local or MCP server]
  T -->|results| H
  H --> R[RunResult: answer, trajectory, usage, cost, termination reason]
  R --> E[Evals: deterministic metrics, judges, trace analysis]

The harness is src/stockroom/agent/harness.py; its states are PLAN → CALL_MODEL → EXECUTE_TOOLS → OBSERVE → DONE | FAILED, with guards for max steps, a per-run token budget, repeated identical calls and wall-clock time, context compaction, and quarantine of instruction-like text in tool output.

The four days

Each day answers one question, with a 90-minute lecture in the morning and a four-hour lab in two parts. Every lab runs offline in mock mode; live AWS is opt-in.

Day 1

LLM Evaluation Foundations

What does “the agent works” mean, and how would you know?

Lecture Slides Lab

Day 2

LLM-as-a-Judge, Calibration, and OTel

When can a model grade a model, and what did the agent actually do?

Lecture Slides Lab

Day 3

Agent Harness and MCP Mocking

Which failures live in the harness, and how do you build one that refuses them?

Lecture Slides Lab

Day 4

AWS CI/CD and Red Teaming

How does an eval become a gate that blocks a merge?

Lecture Slides Lab

Instructors open with the kickoff slides; the Day 1 deck makes the case for the week, and Days 2, 3 and 4 each have a teaching deck with speaker notes. Each notebook has a matching solutions version.

The path

Each day’s lab uses the metrics of the day before. The same golden set, the same four weakness flags and the same agent come back every day, so each fix is measured, not asserted.

Day 1 · Foundations

L1

LLM Evaluation Foundations: From Vibes to Measured Agents

Vibe-driven versus measured development, a failure taxonomy mapped to the repository, building and versioning a golden dataset, deterministic trajectory metrics with the DeepEval bridge, and RAG evaluation over the policy corpus.

Lab

Deterministic and RAG Evaluations for the Stockroom Agent

Tour the golden set, switch the weakness flags on and watch the taxonomy light up, write DeepEval assertions over a run, score the BM25 policy retriever with RAGAS, then add three golden cases of your own.

Day 2 · Judges and Traces

L2

LLM-as-a-Judge, Calibration, and OpenTelemetry Traces

Grading free-form multi-turn output, error analysis before metrics, calibrating a judge against human labels, probing it for position, verbosity and self-preference bias, and one span tree per run.

Lab

Judge Calibration and OpenTelemetry Traces

Instrument runs with OpenTelemetry (in memory, or Arize Phoenix locally), read loops and context growth off the spans, then take a judge rubric from v1 to v2 against 32 human-labelled items.

Day 3 · The Harness

L3

Building a Custom Agent Harness, Mocking Tools over MCP, and the Managed Alternative

The state machine, five runtime failure classes each with evidence, MCP as the tool boundary, deterministic environments for evals, and Amazon Bedrock AgentCore’s managed harness compared with the hand-built one.

Lab

Building a Custom Agent Harness

Argument validation against schema drift, guards, context compaction and tool-output quarantine; serve the tools over MCP; then switch each seeded weakness on, chart what moves, and fix them one at a time.

Day 4 · The Gate

L4

Production CI/CD for Agents on AWS: Regression Gates, Red Teaming, and Bedrock Evaluations

Offline versus online evaluation, product floors kept apart from measured noise, flaky-eval management, a red-team suite, the CI gate step by step, Bedrock and AgentCore Evaluations, and an OIDC-assumed IAM role with no long-lived keys.

Lab

Bedrock Evaluations, Red Teaming, and CI Gating

Validate synthetic case labels against the data before any agent runs, write a red-team case that catches the unsafe trajectory, set thresholds from run-to-run spread, and make the gate reject the regressions an aggregate-only gate misses, using your Day 1–3 work. Bedrock Evaluations and AgentCore are optional extensions.

Seeded weaknesses

Stockroom ships with every weakness fixed, so the repository passes its own gate. The weaknesses are feature flags: Day 3 switches them on one at a time with STOCKROOM_WEAKNESSES=<flag>[,<flag>] and measures each fix.

ambiguous_tool_desc

search_products claims to cover stock and orders too.

Tool-selection accuracy 1.0 → 0.64 while the answers often still look right.

oversized_payload

search_products returns every field of every product, with no limit.

Context growth, compaction, and runs ended by TOKEN_BUDGET.

naive_retry

The legacy order shard is down, with no bounded retry and no repeated-call guard.

Identical calls until MAX_STEPS; the loop detector fires.

injection_unguarded

Tool output is not quarantined.

An instruction planted in a policy document triggers a 10,000-unit restock and a system-prompt leak.

Quick start

No AWS account is needed. The first make setup downloads packages; everything after that runs offline, with a scripted fake model and a deterministic fake judge.

git clone https://github.com/genial-labs-ai/system_agent_harness_aws.git
cd system_agent_harness_aws
make setup          # uv sync (+ Phoenix extra), npm ci, and a Jupyter kernel
make test           # unit tests + golden-set regression suite
make notebooks      # build student + solution notebooks and execute all eight in mock mode
make ci             # the whole PR gate
STOCKROOM_WEAKNESSES=ambiguous_tool_desc make ci    # watch the gate fail on purpose

Default Mock mode

  • Model: FakeBedrockClient, scripted per golden case plus a deterministic planner.
  • Judge: FakeJudge, deterministic.
  • Credentials: none. Spend: zero.
  • Fixed seeds and chars/4 token estimates, so every run reproduces.

Opt-in with STOCKROOM_MODE=live Live mode on Amazon Bedrock

  • Model: the Bedrock Converse API, AGENT_MODEL_ID from the environment.
  • Judge: a Bedrock judge from a different model family, JUDGE_MODEL_ID.
  • Credentials: GitHub OIDC → IAM role in CI, your AWS profile locally.
  • Spend bounded by MAX_STEPS and TOKEN_BUDGET; an estimated cost is printed per run.

The live workflow is dispatched by hand, never on a schedule, so Bedrock spend happens only when an instructor asks for it. Model access, the S3 bucket and the OIDC role are in the README and in AWS policies.

How each day works

The same skeleton every day, 09:00 to 17:00: a lecture, a lab in two parts, and a review. The instructor guide has the timetable to the minute and the places participants get stuck.

  1. Lecture, 90 min. Objectives, the day’s failure classes, and the code that reproduces each one.
  2. Lab part 1, 90 min. Guided cells: run the metric, read the trajectory, predict before you run.
  3. Lab part 2, 150 min. Exercise cells with a matching solution; fix a weakness and measure the delta.
  4. Review, 60 min. Metric deltas side by side, discussion questions, the mistakes people make.

What you will be able to do

  • Explain why Agent = Model + Harness, and name the failure classes a single-prompt eval cannot see.
  • Build, label, document and version a golden dataset, with a manifest hash and a dataset card.
  • Run deterministic trajectory metrics and DeepEval assertions over a golden set, and read the aggregate a CI gate consumes.
  • Evaluate the RAG half of an agent with RAGAS over its own corpus.
  • Calibrate an LLM judge against human labels, measure agreement, and probe it for bias.
  • Emit one OpenTelemetry span tree per run and read loops and context growth off it.
  • Build a harness with schema validation, guards, context compaction and tool-output quarantine, and put the tools behind MCP.
  • Turn metrics into a PR gate that checks its own evidence, sees per-category and cost regressions, runs a red-team suite, and reaches AWS through GitHub OIDC with no long-lived keys.

Who it is for

Senior and staff AI engineers, MLOps architects and tech leads who already ship LLM features. Nothing here explains what a token or a tool call is. You should be comfortable with:

  • Python: you can read a small package, run pytest, and write a test.
  • LLM applications: you have called a chat model with tools, or built a RAG pipeline.
  • CI: you have edited a GitHub Actions workflow.
  • AWS: optional. Every lab runs without an account; the live sections show what changes with one.

Bring a laptop with Python 3.12, uv, Node 22 (the version CI uses; 20 or newer works) with npm, and make. Or open the repository in GitHub Codespaces and skip the setup.

Back to top

Evaluating Autonomous Agents · © 2026 Genial Labs · MIT licence · Built with Quarto

  • AWS Policies

  • Decisions

  • Contributing

  • Security

  • Edit this page
  • Report an issue