AI Implementation / Discovery and Architecture / Evaluation plan
Design acceptance before the prototype.
How acceptance is defined before build: the unit being judged, the test population, the grading method, the thresholds, the release gate and the regression policy.
A convincing demonstration is not evidence that a system can perform real work.
YPAI defines the evaluation structure before implementation begins. The test set represents the real workflow, including ordinary cases, difficult cases, incomplete inputs, conflicting information, tool failures, permission failures and prohibited actions.
The architecture is then measured as a system.
Evaluation plan
- 01 The unit being judged
- 02 The test population
- 03 The grading method
- 04 The acceptance thresholds
- 05 The release gate
- 06 The regression policy
One acceptance plan across product, engineering and operations.
The evaluation plan defines:
The unit being judged
A response, classification, document, case, tool call, workflow or completed business outcome.
The test population
Representative work, difficult cases, known failure modes and scenarios the system must refuse.
The grading method
Deterministic checks, reference-based scoring, model-based evaluation, human review, specialist adjudication or a combined method.
The acceptance thresholds
Required quality, latency, cost, safety, completion and control performance.
The release gate
What must pass before pilot, production release or model change.
The regression policy
Which evaluations rerun after prompt, model, retrieval, tool, data or workflow changes.
The team knows how success will be measured before build decisions become expensive to reverse.
The metric taxonomy
Business outcome metrics
- end-to-end task success
- accepted completion rate
- cycle time
- throughput
- automation rate
- human escalation rate
- correction and rework rate
- cost per accepted outcome
- successful outcomes per unit of time and cost
Knowledge and retrieval metrics
- retrieval precision at K
- retrieval recall at K
- mean reciprocal rank
- normalised discounted cumulative gain
- groundedness
- citation precision
- citation coverage
- response completeness
- source freshness
- unsupported-claim rate
Agent and workflow metrics
- tool-selection accuracy
- tool-argument accuracy
- task completion
- trajectory compliance
- step efficiency
- retry rate
- recovery success
- fallback success
- unauthorised-action rate
- approval-gate compliance
Operational metrics
- time to first response
- end-to-end latency at P50, P95 and P99
- timeout rate
- application and model error rate
- throughput under expected load
- token and compute consumption
- cache utilisation
- cost per request
- cost per completed task
- availability and recovery behaviour
Control metrics
- policy-violation rate
- prompt-injection resistance
- sensitive-data exposure rate
- human override rate
- trace completeness
- evaluation coverage
- regression pass rate
- rollback success
- unresolved incident rate
Not every metric belongs in every system.
The discovery engagement selects the metrics that determine whether the actual workflow succeeds, then attaches thresholds to the approved task set and failure scenarios.