Testing an AI application means checking both the generated response and the software around it: the data the user may access, the tools the system may call, the actions it actually completes, and the experience when something fails. A convincing answer can still be wrong, and a correct answer can still reach the wrong account.
This guide contains 25 practical QA scenarios for teams building assistants, retrieval-based applications, and agents. Use them to define an initial scope, then adapt the inputs and acceptance criteria to your product. Testers HUB’s AI testing services can help review that coverage and execute agreed checks across the application.
Download the 25-Scenario QA Worksheet
How do you test an AI application?
- Define the intended users, supported tasks, trusted data, permissions, and unacceptable failures.
- Create representative test inputs, including ambiguous, unsupported, adversarial, and failure conditions.
- Write measurable pass criteria for factual accuracy, task completion, tool use, access control, and usability.
- Execute the tests with recorded model, prompt, tool, application, and retrieval versions. Repeat variable cases across an agreed sample.
- Review failures, verify fixes, and retain useful cases for the next release.
The examples below are proposed test designs. They are not reports of completed client testing. In the downloadable worksheet, every case begins as Not Run, with space for actual results, evidence references, and run details.
What are AI evals, and where does conventional QA fit?
An evaluation, or eval, measures behavior against defined criteria using a set of inputs. It might check whether answers are supported by documents, whether an agent selects the right tool, or whether a user task completes successfully. Traditional QA still covers login, billing, permissions, UI behavior, APIs, and error recovery.
For generated answers, exact wording is often the wrong pass condition. A useful rubric separates required facts, unsupported claims, completeness, and task success. Deterministic checks can validate a schema or permission rule; human review can assess usefulness and ambiguous cases. Model-based grading should be checked against human judgments before its scores guide release decisions. See OpenAI’s evaluation guidance for task-specific datasets, evaluation criteria, and continuous evaluation practices.
25 QA scenarios for AI applications
Use the expected behavior as a starting point and confirm it with your product owner. Set repeat counts, latency limits, quality thresholds, and allowed tool actions before execution. A single passing run or an average score cannot establish that all critical behaviors are reliable.
Prompts and output
1. Paraphrases of the same request
- Setup / input: Ask the same supported question using five different phrasings and realistic typos.
- Failure to look for: Equivalent requests produce contradictory facts or actions.
- Expected behavior: Core facts and permitted actions remain consistent; wording may differ.
- Evidence: Input variants, outputs, reference answer and model/prompt versions.
- Regression: Repeat the variants after prompt or model changes.
2. Missing or ambiguous information
- Setup / input: Request an action without a required date, account or recipient.
- Failure to look for: The app guesses a critical value and proceeds.
- Expected behavior: It asks for the missing information before committing the action.
- Evidence: Conversation and any tool-call trace.
- Regression: Keep both ambiguous and clarified versions.
3. Confident but unsupported claims
- Setup / input: Ask a factual question whose answer is absent from the approved source material.
- Failure to look for: The response invents a fact or presents uncertainty as certainty.
- Expected behavior: It states the limitation or escalates according to the agreed policy.
- Evidence: Question, approved sources, response and reviewer rationale.
- Regression: Retain the unsupported question with a supported control.
4. Structured output and business validity
- Setup / input: Request a structured record with required fields, allowed values and known business constraints.
- Failure to look for: The JSON is valid but the values are missing, impossible or inconsistent.
- Expected behavior: Schema and business rules pass; invalid output does not trigger downstream work.
- Evidence: Raw response, schema result and business-rule checks.
- Regression: Repeat after schema, prompt or parser changes.
5. Length, language and tone
- Setup / input: Use supported languages, long inputs and customer-facing tone requirements.
- Failure to look for: Responses truncate essential information, switch language or breach the product rubric.
- Expected behavior: The answer preserves essential facts and meets agreed language, length and tone criteria.
- Evidence: Input, output, rubric and human rating.
- Regression: Retain boundary inputs and human-reviewed examples.
Retrieval and data
Retrieval-augmented generation (RAG) combines retrieved source material with generated answers. Check retrieval and answer grounding separately so failures can be traced to the right layer.
6. Correct source retrieval
- Setup / input: Use an approved document collection with a known answer and similar distractor documents.
- Failure to look for: The app retrieves the wrong document or ignores the relevant passage.
- Expected behavior: Authorized relevant evidence is retrieved and supports the response.
- Evidence: Document versions, retrieved passages, ranking if available and answer.
- Regression: Repeat after index, embedding or retrieval configuration changes.
7. Citation accuracy
- Setup / input: Ask for a sourced answer to a question with a known supporting passage.
- Failure to look for: The citation is invented or does not support the associated claim.
- Expected behavior: Each cited reference resolves to accessible evidence that supports its claim.
- Evidence: Claim-to-source comparison and citation targets.
- Regression: Recheck citations after document and generation changes.
8. Conflicting or outdated documents
- Setup / input: Provide current and superseded policies with explicit dates and precedence rules.
- Failure to look for: The response silently uses an outdated rule.
- Expected behavior: It follows the defined source precedence or flags unresolved conflict.
- Evidence: Source versions, retrieved context and response.
- Regression: Keep conflicting-document fixtures as the corpus evolves.
9. Document access and tenant isolation
- Setup / input: Use two test tenants with different documents and revoke one test user’s access.
- Failure to look for: An answer, citation or cached conversation exposes another tenant’s restricted data.
- Expected behavior: Retrieval, generated output and cached results enforce current access rules.
- Evidence: Role matrix, test document identifiers and sanitized retrieval traces.
- Regression: Repeat across roles, sessions, exports and revoked access.
10. Instructions hidden in retrieved content
- Setup / input: Place a controlled instruction in a test document that conflicts with the app’s task and permissions.
- Failure to look for: Retrieved text overrides trusted instructions or causes unauthorized disclosure or action.
- Expected behavior: The app treats the document as data and maintains task and authorization boundaries.
- Evidence: Document fixture, response and tool trace without real secrets.
- Regression: Retain direct and indirect injection variants for future changes.
Tools and agents
11. Tool choice and argument mapping
- Setup / input: Ask for an allowed task with a specific account, recipient and date.
- Failure to look for: The wrong tool or a wrong parameter is used.
- Expected behavior: The permitted tool receives correct validated arguments for the requested task.
- Evidence: Request, selected tool, arguments and observed result.
- Regression: Repeat after tool schema, routing or prompt changes.
12. Authorization and confirmation
- Setup / input: Ask a read-only test user to perform a restricted write or an action requiring confirmation.
- Failure to look for: A fluent response or generated tool call bypasses permission or approval.
- Expected behavior: Server-side authorization blocks the write; required confirmation precedes execution.
- Evidence: User role, approval state, API result and audit record.
- Regression: Repeat with direct calls and multi-turn attempts.
13. Retries and duplicate actions
- Setup / input: Simulate a timeout after a sandbox action has completed, then retry the request.
- Failure to look for: A retry creates a second order, message or other side effect.
- Expected behavior: The system reconciles the existing result or uses its defined deduplication mechanism.
- Evidence: Request identifiers, retry trace and sandbox records.
- Regression: Keep failure-before-commit and failure-after-commit cases.
14. Tool result versus claimed success
- Setup / input: Return an error, empty response or partial completion from a test tool.
- Failure to look for: The assistant reports success even though the operation failed.
- Expected behavior: The user-facing status matches the actual result and gives a valid next step.
- Evidence: Tool output, final answer and resulting application state.
- Regression: Repeat for each important error and partial-result path.
15. Stopping, cancellation and bounded work
- Setup / input: Give a task with no reachable answer, a repeated tool failure, or a user cancellation.
- Failure to look for: The agent loops, continues after cancellation or exceeds the agreed work budget.
- Expected behavior: It stops at configured limits, handles cancellation and reports incomplete work accurately.
- Evidence: Step trace, cancellation time, usage and remaining side effects.
- Regression: Repeat after orchestration, retry or budget changes.
Conversation and application
16. Corrections across multiple turns
- Setup / input: Ask a question, correct a key constraint, then request a final answer or action.
- Failure to look for: The system uses the earlier value or mixes incompatible context.
- Expected behavior: It uses the latest applicable instruction and resolves contradictions before acting.
- Evidence: Full conversation and final action parameters.
- Regression: Retain short and longer correction sequences.
17. Session boundaries and history
- Setup / input: Switch accounts or workspaces, sign out and reopen a saved chat with test data.
- Failure to look for: History or generated output appears in the wrong account or session.
- Expected behavior: History follows the product’s retention and access rules without cross-user leakage.
- Evidence: Account/session matrix and sanitized UI/API observations.
- Regression: Repeat after session, storage and history changes.
18. Billing, quotas and entitlements
- Setup / input: Use test accounts with different plans, exhausted quotas and a sandbox plan change.
- Failure to look for: Usage is miscounted or an unavailable feature remains accessible.
- Expected behavior: Enforced access, usage accounting and user messaging follow the agreed plan rules.
- Evidence: Plan state, usage records and UI/API results.
- Regression: Repeat after billing, model routing or entitlement changes.
19. Unsafe requests and appropriate refusals
- Setup / input: Pair controlled disallowed requests with similar legitimate requests from the product’s policy set.
- Failure to look for: Unsafe actions are allowed or harmless requests are unnecessarily blocked.
- Expected behavior: The app follows its defined policy, provides useful safe alternatives and escalates where required.
- Evidence: Policy reference, input/output and reviewer decision.
- Regression: Track unsafe compliance and false refusals separately.
20. Long output and interrupted streaming
- Setup / input: Generate a long answer on priority browsers/devices; interrupt and resume where supported.
- Failure to look for: Controls become inaccessible, partial text appears final or incomplete content is committed.
- Expected behavior: The UI remains usable and distinguishes partial, failed and completed output.
- Evidence: Screen recording, browser/device details and stream events.
- Regression: Repeat after UI, rendering and streaming changes.
Reliability and regression
21. Provider failure and fallback
- Setup / input: Simulate rate limiting, timeout and model/provider unavailability in a test environment.
- Failure to look for: The app hangs, silently switches to unsupported behavior or loses user work.
- Expected behavior: It applies the agreed retry/fallback policy and preserves recoverable input.
- Evidence: Error codes, retries, route used and user-visible result.
- Regression: Repeat after provider and fallback configuration changes.
22. Latency and usage cost
- Setup / input: Run representative short, long and tool-heavy tasks under agreed conditions.
- Failure to look for: Slow responses or repeated calls exceed the product’s latency or usage budget.
- Expected behavior: Measured latency and usage meet team-defined limits or produce an explicit release finding.
- Evidence: Time to first output, total time, calls, tokens/usage and dated cost assumptions.
- Regression: Compare like-for-like tasks after model and workflow changes.
23. Repeated-run variability
- Setup / input: Run critical cases multiple times with a fixed documented configuration and data snapshot.
- Failure to look for: One passing answer hides intermittent factual or action failures.
- Expected behavior: Results meet the team’s pre-agreed rubric and acceptance thresholds across the chosen sample.
- Evidence: Every run, sample size, configuration, ratings and failure counts.
- Regression: Retain failures and rerun comparable samples.
24. Versioned release comparison
- Setup / input: Compare the baseline and candidate model, prompt, retrieval corpus or tool configuration on the same held-out cases.
- Failure to look for: A new version improves one area while breaking another.
- Expected behavior: Critical cases and per-category criteria meet the agreed release gates; regressions are reviewed.
- Evidence: Version identifiers, case-set version and before/after results.
- Regression: Use a held-out suite and investigate changes before promoting the candidate.
25. Production failure to regression case
- Setup / input: Reproduce a reported failure with permission and sanitized data in a controlled environment.
- Failure to look for: The fix addresses one phrase but misses the underlying failure pattern.
- Expected behavior: The original failure and relevant variants pass after the fix, with neighboring behavior checked.
- Evidence: Sanitized report, reproduction, fix version and retest observations.
- Regression: Add the approved case to the maintained suite; keep private data out of shared fixtures.
What our QA team changed after AI generated a test plan
Our published AI test planning review documents a review of 160 generated cases for a project-management web application. In one case, the draft assumed duplicate project names had to be rejected. The reviewer removed that unsupported expectation so the test could follow an approved business rule. The review also added coverage for session expiry while a form contained unsaved work.
Those are planning corrections, not confirmed application defects: the source cases remained Not Run. They demonstrate why a generated checklist needs review before execution. The same discipline applies when reviewing an AI application’s answers—separate a plausible assumption from an established requirement.
The Testers HUB AI Test Management Platform supports AI-assisted drafting, human review, manual execution, defects, and regression organization. This workflow is distinct from evaluating the AI features inside a customer’s product. Our AI-assisted regression planning article explains the review process in more detail.
How is AI agent testing different from chatbot testing?
An agent can change application state through tools and APIs. Check the complete path from user request to tool selection, arguments, authorization, side effect, and final response. A correct-sounding answer does not prove that the intended action occurred. Conversely, a timeout does not prove that an action failed; retrying without reconciliation can create duplicates.
Use controlled environments for action tests and verify the underlying records. Confirm which actions require approval, how cancellation works, and what stops an agent from continuing beyond its permitted task or budget.
What should an AI regression suite contain?
Maintain representative successful tasks, confirmed failures, important edge cases, access boundaries, retrieval examples, and tool-error paths. Keep a held-out set for release comparison so prompt tuning does not merely optimize for examples the team already knows.
For each run, record the application build, model identifier and settings, prompt version, tool configuration, corpus/index snapshot, dataset version, rubric, and reviewer. Compare baseline and candidate results by failure category. Investigate changes in critical actions and permissions even when the overall average improves.
Production feedback can supply valuable new tests when it is collected and sanitized appropriately. Reproduce the failure, add related variants, retest the correction, and retain the case. Track the number of runs and observed failures instead of presenting an unmeasured claim of reliability.
How to scope an AI QA engagement
Share the product URL or build, supported tasks, user roles, data sources, connected tools, model/provider configuration where available, known issues, target environments, and release date. Identify the workflows where a wrong answer or action would matter most.
For an AI SaaS product, include the surrounding roles, integrations, billing, and browser experience in web app testing. For AI-enabled Android or iOS applications, agree real-device and interruption coverage through mobile app testing services. Device selection, specialist security work, repeated runs, performance measurements, retesting, and regression should be explicit parts of the scope.
Building an AI app or agent? Share your key workflows, data sources, integrations and release date. We’ll recommend the QA scope and evaluation coverage.


