Skip to content
AgentEnv Framework
CookbooksCookbooks

Cookbooks

Complete recipes that take you from a fresh install to a graded run

A cookbook puts the pieces from the other sections together into one task you can run: every file, every command and what each one prints, from an empty folder to a score. Each one explains what happens at every step, and links to the pages that cover it in depth.

Each one is an enterprise task, built to be hard in the ways real work is: the prompt doesn't say what to do, the system enforces its own rules, and the right answer takes several checks the agent has to think of.

  1. Single-Turn Prompt in an RL Env Graded by an Agent Judge: an accounts-payable agent works a week's invoice queue against policy in one prompt, and a judge agent grades every decision it made, with partial credit and a penalty.
  2. Multi-Turn Conversation Between Two Agents in an RL Env Graded by an Agent Judge: a support agent works an enterprise escalation in a conversation with a second agent that plays the customer, and a judge grades every turn.

Last updated on

Ask a question · Report an issue