The evaluation specification must exist before inference begins. Option D defines what the evaluation will measure and how results will be segmented. Representative cases measure normal workload performance, edge cases test boundary behavior, and adversarial cases examine safety, security, or robustness under hostile inputs.
Option B then supplies the correctly curated and labeled dataset corresponding to those slices. Labels, expected outcomes, grading rubrics, provenance, and dataset versions must be established before the deployment generates outputs. Otherwise, evaluators risk modifying criteria after seeing results and introducing confirmation bias.
Anthropic defines an evaluation as an input combined with grading logic used to measure success. Its evaluation guidance distinguishes tasks, trials, graders, traces, outcomes, and the harness that runs and aggregates them. Demystifying Evals for AI Agents
Option E occurs after model execution because outputs must exist before they can be scored. Option A follows scoring and aggregation. Option C is an iterative improvement step performed after failure cases become available, although rubric calibration may also be conducted on separate development data.
Study Guide references/topics: Evaluation specification; metrics and slices; representative and adversarial datasets; labeling; scoring; aggregation; release gates.
===============