An AI agent can produce a convincing final answer while selecting the wrong tool or performing an action it was never authorized to take. Production evaluation needs to inspect the execution as well as the response.
Begin with a bounded task contract: the input, available tools, allowed actions, review requirements, and completion condition. A vague goal such as handling customer support is difficult to test. Preparing an approved response to a specific type of request gives the evaluation a workable boundary.
Evaluate the task in separate dimensions
Review whether the agent understood the request, obtained the necessary evidence, selected an allowed tool, supplied valid arguments, respected approval, and reported the actual outcome. A single average quality score can conceal a critical failure in any of these dimensions.
My recorded work on Naya and Narravo informs the focus on bounded responsibilities and source access. Those case studies describe engineering scope, not a transferable benchmark score for every agent architecture.
Use a scenario matrix
The following is a fictional evaluation plan for an agent that prepares service-request drafts. It is a template to adapt, not a measured product result.
Scroll sideways to view the full table.
| Scenario | Expected behavior | Evidence to inspect |
|---|---|---|
| Valid request | Prepare the permitted draft | Tool arguments and saved record |
| Missing required detail | Ask for clarification | No fabricated field or write |
| Approval denied | Stop the proposed action | Recorded decision and no execution |
| Tool timeout | Report uncertainty and follow the recovery policy | Attempt history and task state |
| Duplicate delivery | Avoid a duplicated business action | Stable operation identifier and destination state |
| Unauthorized source | Deny access | Retrieval and authorization outcome |
| Instructions inside a document | Treat them as source content | No expansion of tool permissions |
Separate release-blocking failures from quality improvements. A fabricated completion message or unauthorized action deserves a different response from an awkwardly phrased explanation.
Put approvals at the execution boundary
An approval should identify the exact action, relevant target, and important parameters. If those change, the earlier decision should not silently authorize a different action.
LangGraph interrupts support pausing for external input with persisted execution state. The application still needs to authenticate the reviewer and bind their decision to the intended action. A framework feature does not establish your business authorization policy.
Test a declined request, a changed target, and an expired decision. Check the destination system, not just the text shown in the approval interface.
Distinguish retrying from duplicating
A tool timeout does not prove the external action failed. The destination might have accepted it before the response was lost. Retrying a consequential write without checking its state can create a second record.
Define operation identifiers, destination reconciliation, and bounded retry rules around each tool. Where the destination supports idempotency, use it deliberately. Where it does not, design a recovery process that can establish what happened before trying again.
LangGraph's persistence documentation explains checkpointed execution. External side effects still need their own duplication and reconciliation controls.
Keep evaluation cases separate from development examples
Use some cases while developing and reserve others for assessing changes. Otherwise, a prompt can improve on the examples everyone already knows while remaining fragile on unfamiliar requests.
Add production failures to the evaluation set after appropriate redaction. Record model, prompt, tool schema, and source versions so a changed result can be investigated. Keep sensitive task content out of unnecessary logs.
Define a release and rollback decision
Agree who reviews failures, which behaviors block release, and how to disable a problematic capability. Establish operating measurements such as completed tasks, review burden, unresolved failures, and provider usage using your actual application requirements.
The portfolio's agent simulation shows approval, rejection, and escalation with fictional data. For an implementation-specific evaluation plan, explore AI agent development or discuss the workflow.
Prepared with AI assistance using Paul’s documented project work and the linked sources. Examples are illustrative unless identified as project records.