The verified answer is C. Faithfulness. The company is evaluating foundation models for a summarization task by using an LLM-as-a-judge evaluation job in Amazon Bedrock. AWS documentation states that with a model evaluation job that uses a judge model, Amazon Bedrock uses one LLM to score another model’s responses and provide an explanation of how it scored each prompt and response pair. AWS also states that Amazon Bedrock provides built-in metrics that can be selected for these judge-based evaluation jobs.
Among the answer choices, Faithfulness is the only valid built-in metric for an Amazon Bedrock LLM-as-a-judge model evaluation job. AWS API documentation lists the valid built-in metric names for LLM-as-a-judge evaluation jobs, including Correctness, Completeness, Faithfulness, Helpfulness, Coherence, Relevance, FollowingInstructions, ProfessionalStyleAndTone, and responsible AI metrics such as Harmfulness, Stereotyping, and Refusal. The same documentation identifies Summarization as a valid task type for model evaluation jobs. Therefore, Faithfulness is a valid built-in metric for evaluating whether generated summaries remain grounded in the source input rather than introducing unsupported claims.
Option A. Context relevance and option B. Context coverage are incorrect for this question because AWS lists those as metrics for knowledge base retrieval-only evaluation jobs, not general LLM-as-a-judge model evaluation of foundation model summarization output. These metrics evaluate retrieved context, not the quality of a summarization model’s generated response. Option D. RMSE is also incorrect because root mean square error is a regression metric. It is not a built-in LLM-as-a-judge metric for summarization in Amazon Bedrock.
Because the task is summarization and the evaluation job is LLM-as-a-judge, the correct built-in evaluation metric from the available choices is Faithfulness.