System owners building responsible, high-stakes agentic AI systems must define explicit intended outcomes, affected populations, and decision context. Claims of effectiveness require expanding the definition of quality beyond accuracy, ensuring fairness across groups, and aiming for a high standard of evidence of desired outcomes.
Clearly Documented Purpose
Builders should clearly document in as much detail as possible the target population, the purpose of the system, the estimated cost, the measurable outcomes it seeks to achieve, and the system context.
Quality Beyond Accuracy
For agentic AI, “quality” must be defined more broadly to include completeness, correctness, explainability and clarity, groundedness, and reliability under variation. AI builders and evaluators should use some form of human evaluation to assess each of these dimensions. Additional automated and LLM-as-judge techniques should be used only to amplify, validate, and scale these approaches.
Fairness Across Groups
Fairness across groups means that the AI agent performs consistently across different user profiles, backgrounds, or personal characteristics. Fairness metrics are necessary (but insufficient) steps to considering discrimination that occurs systemically in institutions that interact with AI. Builders should carefully consider the correct metrics for measuring fairness and evaluate the meaning of each metric based on an expert judgment of risk, potential harms, and user feedback.
Generally, we recommend that builders create a small set of structured scenarios based on the AI agent’s purpose and intended use. They should choose relevant subgroups from the agent’s target population, based on the potential for plausible harm. Evaluators should measure how decisions and outcomes affected by the AI agent differ by subgroup and ask members of each subgroup to test the AI agent to collect differences in responses. They should then measure the decision gap, outcome gap, and bias gap.
A High Standard of Evidence
When developing evidence of desired outcomes, builders should use a documented tiering system, match rigor to risk and fairness, and seek input from target populations.
Evidence Tiering
Builders should document their evidence-tiering approach both before deploying an AI agent and at a regular interval once it is live as part of monitoring activities. All analyses should include a distributional analysis by relevant subgroups where possible. The tiering system should move from weaker evidence (tier 4) to stronger evidence (tier 1 or 2) as the agent matures.
Matching Evidence Rigor to Risk and Fairness
Higher agency on behalf of the agent and higher-stakes outcomes, even within the high-stakes group defined here, should require higher tiers of evidence. If a causal claim is made (e.g., “reduces time spent applying for SNAP by 20 percent”), the evaluation design should support that claim or clearly state the limitations. When presenting higher tiers of evidence, evaluators should also ensure they provide sufficient statistical power to capture important subgroup differences, aligned with the agent’s fairness goals.
Target Population Input as Evidence
“Quality” and “evidence” should be measured by quantitative, qualitative, and community-engaged methods. While it may be tempting for many, especially technical builders, solely to quantify and automate the definitions in this playbook, qualitative and community-informed evidence are also necessary parts of a healthy feedback loop; they gather important intermediate information that can act as leading indicators of key quantitative outcomes down the line and would otherwise be missed.
Part of collecting qualitative evidence is to continuously solicit stakeholder feedback from the target population throughout agent development and implementation. System owners should solicit feedback from a range of users with different backgrounds or relevant experiences, incorporate or accept the risks of not incorporating that feedback, document these decisions, and share the decisions with the user community.
Next section: Principle 2: Provide Oversight and Ownership