Exam AB-620 Topic 1 Question 55 Discussion
Actual exam question for Microsoft's AB-620 exam
Question #: 55
Topic #: 1
Question #: 55
Topic #: 1
A team is preparing to evaluate an agent in Copilot Studio before expanding access to additional users.
The team must choose an evaluation method that meets the following requirements:
Use a fixed set of prepared interactions.
Determine responses against a predefined baseline.
Support consistent comparison across repeated test runs.
You need to determine the evaluation method that satisfies the evaluation requirements.

The team must choose an evaluation method that meets the following requirements:
Use a fixed set of prepared interactions.
Determine responses against a predefined baseline.
Support consistent comparison across repeated test runs.
You need to determine the evaluation method that satisfies the evaluation requirements.

Suggested Answer:

Explanation:
Fixed prepared interactions # Use a curated prompt list; Predefined baseline # Compare against predefined expected outcomes; Repeatable comparison # Use a consistent scoring evaluation matrix.
Comprehensive and Detailed Explanation From Microsoft AB-620 Study Guide: A curated prompt list creates the controlled input population required for repeatability. Predefined expected outcomes establish what a correct or acceptable response should contain, which lets the evaluator judge performance against a baseline rather than against a reviewer ' s memory. A stable scoring matrix ensures that repeated runs are interpreted under the same thresholds and dimensions. Live conversations and post-rollout feedback can enrich the test inventory, but they cannot replace a fixed regression set because the questions, users, and context change. Publishing status confirms deployment, not response quality. General impressions are similarly unsuitable for consistent comparison. The team should select the scoring method according to the response type: semantic comparison for valid paraphrases, exact match for immutable values, keyword checks for mandatory phrases, and custom criteria for compliance or domain quality. Versioning the prompt list, expected answers, rubric, agent version, and knowledge snapshot is essential; otherwise a changed test set can be mistaken for a change in agent performance. Study Guide alignment: Test and manage agents > Evaluate agent performance > Create a test set; Choose an evaluation method.
by Clark at Aug 26, 2026, 09:04 PM
0
0
0
10
Comments
Upvoting a comment with a selected answer will also increase the vote count towards that answer by one. So if you see a comment that you already agree with, you can upvote it instead of posting a new comment.
Report Comment
Commenting
You can sign-up / login (it's free).