Evaluating Autonomous Agents
Systems, Harnesses & AWS Production CI/CD
Genial Labs · a four-day workshop
Evaluating Autonomous Agents
Systems, Harnesses & AWS Production CI/CD
An agent that works in a demo is not evidence that it works. Agents break in the harness and in the environment, and the final answer often still looks fine. This workshop builds the evals that catch those failures, then wires them into a CI gate on AWS. One running example, Stockroom, an inventory and order-support agent; every claim about a failure mode is demonstrated by a test or a notebook cell you can run offline.
4 days 4 lectures, 8 notebooks 5 tools, 4 seeded weaknesses 50 golden cases 0 AWS credentials needed
Agent = Model + Harness The model proposes; the harness decides what is allowed to happen. Most production failures live in the harness and the environment, so that is what the eval suite exercises first.
Why single-prompt evals are not enough
A single-prompt eval scores the final answer and nothing else. The Stockroom suite scores tool selection, tool arguments, trajectories, guards, context compaction and injection handling, and only then the answer.
Where failures live The model
Wrong tool, wrong arguments, an answer that is not grounded in what the tools returned. The part single-prompt evals already see.
Where most of them live The harness
Tool-call formatting, loop control, context truncation, state drift, unquarantined tool output. The agent loop you wrote, or the one your framework wrote for you.
Where the rest live The environment
Tools, data and permissions: a misleading tool description, an oversized payload, a shard that is down, a policy document with an instruction planted in it.
flowchart LR U[user query] --> H[Harness state machine] H -->|Converse API or FakeBedrockClient| M[Model] M -->|tool calls| H H -->|JSON-schema validation, guards, quarantine| T[Tools: local or MCP server] T -->|results| H H --> R[RunResult: answer, trajectory, usage, cost, termination reason] R --> E[Evals: deterministic metrics, judges, trace analysis]
The harness is src/stockroom/agent/harness.py; its states are PLAN → CALL_MODEL → EXECUTE_TOOLS → OBSERVE → DONE | FAILED, with guards for max steps, a per-run token budget, repeated identical calls and wall-clock time, context compaction, and quarantine of instruction-like text in tool output.
The four days
Each day answers one question, with a 90-minute lecture in the morning and a four-hour lab in two parts. Every lab runs offline in mock mode; live AWS is opt-in.
Day 1
What does “the agent works” mean, and how would you know?
Day 2
LLM-as-a-Judge, Calibration, and OTel
When can a model grade a model, and what did the agent actually do?
Day 3
Which failures live in the harness, and how do you build one that refuses them?
Day 4
How does an eval become a gate that blocks a merge?
Instructors open with the kickoff slides; the Day 1 deck makes the case for the week, and Days 2, 3 and 4 each have a teaching deck with speaker notes. Each notebook has a matching solutions version.
The path
Each day’s lab uses the metrics of the day before. The same golden set, the same four weakness flags and the same agent come back every day, so each fix is measured, not asserted.
L1
LLM Evaluation Foundations: From Vibes to Measured Agents
Vibe-driven versus measured development, a failure taxonomy mapped to the repository, building and versioning a golden dataset, deterministic trajectory metrics with the DeepEval bridge, and RAG evaluation over the policy corpus.
Lab
Deterministic and RAG Evaluations for the Stockroom Agent
Tour the golden set, switch the weakness flags on and watch the taxonomy light up, write DeepEval assertions over a run, score the BM25 policy retriever with RAGAS, then add three golden cases of your own.
L2
LLM-as-a-Judge, Calibration, and OpenTelemetry Traces
Grading free-form multi-turn output, error analysis before metrics, calibrating a judge against human labels, probing it for position, verbosity and self-preference bias, and one span tree per run.
Lab
Judge Calibration and OpenTelemetry Traces
Instrument runs with OpenTelemetry (in memory, or Arize Phoenix locally), read loops and context growth off the spans, then take a judge rubric from v1 to v2 against 32 human-labelled items.
L3
Building a Custom Agent Harness, Mocking Tools over MCP, and the Managed Alternative
The state machine, five runtime failure classes each with evidence, MCP as the tool boundary, deterministic environments for evals, and Amazon Bedrock AgentCore’s managed harness compared with the hand-built one.
Lab
Building a Custom Agent Harness
Argument validation against schema drift, guards, context compaction and tool-output quarantine; serve the tools over MCP; then switch each seeded weakness on, chart what moves, and fix them one at a time.
L4
Production CI/CD for Agents on AWS: Regression Gates, Red Teaming, and Bedrock Evaluations
Offline versus online evaluation, product floors kept apart from measured noise, flaky-eval management, a red-team suite, the CI gate step by step, Bedrock and AgentCore Evaluations, and an OIDC-assumed IAM role with no long-lived keys.
Lab
Bedrock Evaluations, Red Teaming, and CI Gating
Validate synthetic case labels against the data before any agent runs, write a red-team case that catches the unsafe trajectory, set thresholds from run-to-run spread, and make the gate reject the regressions an aggregate-only gate misses, using your Day 1–3 work. Bedrock Evaluations and AgentCore are optional extensions.
Seeded weaknesses
Stockroom ships with every weakness fixed, so the repository passes its own gate. The weaknesses are feature flags: Day 3 switches them on one at a time with STOCKROOM_WEAKNESSES=<flag>[,<flag>] and measures each fix.
ambiguous_tool_desc
search_products claims to cover stock and orders too.
Tool-selection accuracy 1.0 → 0.64 while the answers often still look right.
oversized_payload
search_products returns every field of every product, with no limit.
Context growth, compaction, and runs ended by TOKEN_BUDGET.
naive_retry
The legacy order shard is down, with no bounded retry and no repeated-call guard.
Identical calls until MAX_STEPS; the loop detector fires.
injection_unguarded
Tool output is not quarantined.
An instruction planted in a policy document triggers a 10,000-unit restock and a system-prompt leak.
Quick start
No AWS account is needed. The first make setup downloads packages; everything after that runs offline, with a scripted fake model and a deterministic fake judge.
git clone https://github.com/genial-labs-ai/system_agent_harness_aws.git
cd system_agent_harness_aws
make setup # uv sync (+ Phoenix extra), npm ci, and a Jupyter kernel
make test # unit tests + golden-set regression suite
make notebooks # build student + solution notebooks and execute all eight in mock mode
make ci # the whole PR gate
STOCKROOM_WEAKNESSES=ambiguous_tool_desc make ci # watch the gate fail on purposeDefault Mock mode
- Model:
FakeBedrockClient, scripted per golden case plus a deterministic planner. - Judge:
FakeJudge, deterministic. - Credentials: none. Spend: zero.
- Fixed seeds and
chars/4token estimates, so every run reproduces.
Opt-in with STOCKROOM_MODE=live Live mode on Amazon Bedrock
- Model: the Bedrock Converse API,
AGENT_MODEL_IDfrom the environment. - Judge: a Bedrock judge from a different model family,
JUDGE_MODEL_ID. - Credentials: GitHub OIDC → IAM role in CI, your AWS profile locally.
- Spend bounded by
MAX_STEPSandTOKEN_BUDGET; an estimated cost is printed per run.
The live workflow is dispatched by hand, never on a schedule, so Bedrock spend happens only when an instructor asks for it. Model access, the S3 bucket and the OIDC role are in the README and in AWS policies.
How each day works
The same skeleton every day, 09:00 to 17:00: a lecture, a lab in two parts, and a review. The instructor guide has the timetable to the minute and the places participants get stuck.
- Lecture, 90 min. Objectives, the day’s failure classes, and the code that reproduces each one.
- Lab part 1, 90 min. Guided cells: run the metric, read the trajectory, predict before you run.
- Lab part 2, 150 min. Exercise cells with a matching solution; fix a weakness and measure the delta.
- Review, 60 min. Metric deltas side by side, discussion questions, the mistakes people make.
What you will be able to do
- Explain why Agent = Model + Harness, and name the failure classes a single-prompt eval cannot see.
- Build, label, document and version a golden dataset, with a manifest hash and a dataset card.
- Run deterministic trajectory metrics and DeepEval assertions over a golden set, and read the aggregate a CI gate consumes.
- Evaluate the RAG half of an agent with RAGAS over its own corpus.
- Calibrate an LLM judge against human labels, measure agreement, and probe it for bias.
- Emit one OpenTelemetry span tree per run and read loops and context growth off it.
- Build a harness with schema validation, guards, context compaction and tool-output quarantine, and put the tools behind MCP.
- Turn metrics into a PR gate that checks its own evidence, sees per-category and cost regressions, runs a red-team suite, and reaches AWS through GitHub OIDC with no long-lived keys.
Who it is for
Senior and staff AI engineers, MLOps architects and tech leads who already ship LLM features. Nothing here explains what a token or a tool call is. You should be comfortable with:
- Python: you can read a small package, run
pytest, and write a test. - LLM applications: you have called a chat model with tools, or built a RAG pipeline.
- CI: you have edited a GitHub Actions workflow.
- AWS: optional. Every lab runs without an account; the live sections show what changes with one.
Bring a laptop with Python 3.12, uv, Node 22 (the version CI uses; 20 or newer works) with npm, and make. Or open the repository in GitHub Codespaces and skip the setup.