1. Turn requirements into review criteria.

Write down what a satisfactory result must contain and what it must not do. For a document summary, criteria might include preserving exclusions, identifying the source and avoiding facts absent from the supplied text. For extraction, assess required fields and their support in the original. A general impression that an answer reads well is not enough.

Separate different kinds of failure. An incomplete answer, a fabricated statement and an unauthorised action create different risks and need different remedies. Define who can accept each type of residual risk. A high average result must not conceal a failure that is unacceptable regardless of how often it occurs.

2. Assemble representative cases.

Include the ordinary inputs the system is intended to handle, along with awkward cases: missing sections, unusual spelling, conflicting instructions and requests outside scope. Use material you have permission to process. Where personal data is unnecessary, use appropriately constructed test inputs or remove it without assuming that simple redaction makes the records anonymous.

Keep development examples separate from the cases used for a release decision. If prompts are repeatedly adjusted against the same questions, performance on those questions may not indicate broader reliability. Record the input, expected behaviour and reason for including each case so another reviewer can understand the assessment.

3. Assess evidence and variation.

Have reviewers inspect source support, not just wording. A hallucination is generated content that is unsupported or incorrect in the relevant context. Fluent answers can contain subtle mistakes, including reversed conditions or missing qualifications. Reviewers need access to the source and enough subject knowledge to identify these errors.

Automated checks can assess format, required fields and exact matches. Another model can assist with review, but its judgement is also an output that may be wrong. Calibrate any automated evaluator against human decisions. Repeat selected cases where output variation matters, and record the configuration used rather than treating one favourable response as the whole result.

4. Test boundaries around the model.

Check whether untrusted input can cause the system to reveal restricted content, invoke an unintended tool or send data to an unapproved destination. Test permissions in the application, not only the wording of the system prompt. A refusal in one conversation does not establish that the same boundary holds across all inputs.

The OWASP guidance for generative AI security discusses issues including prompt injection and sensitive information disclosure. Use it to identify relevant checks for your architecture. It is a reference for risk analysis, not a certificate that a system is safe or an assurance that every possible attack has been covered.

5. Record trade-offs and release conditions.

The NIST AI Risk Management Framework provides a vocabulary for governing, mapping, measuring and managing AI risk. Connect those activities to actual decisions: who owns the use case, which harms matter, how behaviour is measured and what controls are required. Referencing the framework does not replace those decisions.

Record where the system performs adequately and where it does not. A tool may be acceptable for internal drafting with review but unsuitable for unattended external messages. Set deployment boundaries accordingly. Include a rollback route, an escalation contact and circumstances requiring a fresh assessment, such as a new data source or additional tool access.

6. Keep evaluation connected to operations.

Retain a controlled set of regression cases: inputs used again to detect whether a change has damaged previously acceptable behaviour. Run them when prompts, retrieval settings, connected tools or model configurations change. Review the complete workflow because a model improvement can still break an integration's assumptions.

Operational monitoring should identify failures without retaining unnecessary personal information. Give users a way to report an incorrect result and connect those reports to corrective action. An evaluation engagement should provide criteria, recorded findings and unresolved risks; it should not promise universal accuracy. Pair the assessment with team guidance so reviewers understand what the controls require.