AgentEnv Framework
Grading

LLM rubric judging

Score a run against written criteria with a judge model

Score a run against written criteria with a judge model.

This page is a stub

Scaffolded as part of the documentation build-out. Content to be written in a later phase — see docs/ASSUMPTIONS.md for the product gaps some pages depend on.