Manuals / AI Agents & Workflows / Chapter 3

Steady · beginner · ~30 min · Chapter 3 of 3

Evaluate the loop

Score task success, cost, and scary failures — not vibes.

Path progress
100%

Step 1 of 1

Tiny eval set

10 real tasks with expected outcomes. Run weekly when you change prompts/tools.

Tiny eval setDrag stickies · tap for tips
Study mapDrag stickies · tap for tipsKeep it shortdrag · tap →Name the waitdrag · tap →Scope locatorsdrag · tap →Trace when stuckdrag · tap →One browser firstdrag · tap →Isolate statedrag · tap →Assert the UIdrag · tap →Retry wiselydrag · tap →Seed datadrag · tap →Close the loopdrag · tap →Keep it shortdrag · tap →Name the waitdrag · tap →Scope locatorsdrag · tap →Trace when stuckdrag · tap →One browser firstdrag · tap →Isolate statedrag · tap →Pathwise hackdrag · tap →Tiny eval setdrag · tap →Try thisdrag · tap →Follow the dashed drag · tap →

Do this now

Write 5 eval tasks for your agent idea.

Was this step clear?
Chapter learning outcomes
  • Evals
  • Cost

Clear these before you leave