


Technical Evaluation
Technical evaluation measures whether the AI performs the responsibilities and behaviors defined for it. Rather than selecting evaluation methods first, teams begin with the Agent Definition and translate its responsibilities, operating logic, and output requirements into specific evaluation targets. Appropriate methods and success criteria can then be selected for each target based on what is being tested, the degree of judgment involved, and the consequence of failure.
Evaluation is Continuous
Evaluation does not end at launch. Human feedback, technical performance, and business evidence continue to evolve as the product encounters new users, contexts, and conditions. Findings across the three evidence streams should inform ongoing evaluation, product decisions, and changes to the capability itself, creating a continuous loop between expected behavior, observed performance, and realized value.
Business Evaluation
Business evaluation closes the loop between AI performance and the opportunity that justified its investment. It traces the capability's contribution through changes in human behavior and business workflows toward measurable outcomes. Because those outcomes are often influenced by many factors, business evaluation builds a defensible view of AI's contribution without overstating attribution.
Human Evaluation
Human evaluation captures how people experience the AI product before and after launch. Structured testing establishes an early view of whether the capability meets user expectations, while in-product feedback, observed behavior, support issues, and ongoing research reveal how it performs in real-world use. Human feedback should be treated as evidence to investigate, helping teams distinguish AI performance issues from missing context, product experience gaps, operating logic, or differences in human judgment.
AI Evaluation Strategy defines how the capability will be tested, measured, and monitored against its intended behavior and outcomes. It translates the expectations established throughout the framework into measurable criteria for determining whether the AI is performing as intended, producing useful results, and creating the value it was designed to deliver.
The outcome is a clear, repeatable approach for evaluating performance before launch and as the capability evolves in production.
What's Included:
• Human Evaluation
• Technical Evaluation
• Business Evaluation
• Continuous Evaluation
AI Evaluation Strategy brings together human, technical, and business evidence to establish whether an AI product is performing as intended and creating meaningful value. Each evidence stream answers a different question: how people experience and respond to the AI, whether the capability performs according to its defined responsibilities, and whether that performance contributes to the outcomes that justified the investment. No single metric or evaluation method provides the full picture; confidence comes from examining evidence across all three streams and investigating where the findings reinforce or contradict one another. Evaluation continues in production, where new evidence, failure patterns, and user feedback inform ongoing testing and improvement.
From Expectations to Evidence
The AI Product Framework
A capability-first methodology for building products with purpose
Evaluation Strategy
PHASES
AI PRODUCT FRAMEWORK