We mapped 23 AI-generated test cases back to a frozen, sanitized production requirement set. Experienced QA found that only 2 of 10 explicit acceptance criteria had complete draft coverage. Five were partial, three were missing, and five additional test groups depended on behavior the source never specified. After review and four targeted additions, 9 of 10 explicit criteria had approved coverage.
Benchmark result
AI draft coverage was 20%; final human-reviewed coverage was 90%
This benchmark measures requirements-to-test traceability, not the number of scenarios an AI can produce. The source was one real onboarding requirement screenshot already analyzed through Testers HUB AI. Client, product, account, email, URL and proprietary workflow identifiers were removed before publication.
Important denominator: the percentages use only the 10 criteria visible or stated in the frozen source. Five useful AI-generated extensions were recorded separately as unsupported or assumed and were not counted as approved requirements.
Method
How the AI requirements traceability experiment was run
- Freeze the evidence: one sanitized onboarding requirement screenshot captured on August 26, 2026.
- Extract source assertions: 10 explicit acceptance criteria and five important requirement gaps or quality questions.
- Preserve the AI draft: 23 saved cases—13 positive and 10 negative—before independent coverage approval.
- Map by criterion: every test was traced to Covered, Partially Covered, Missing or Unsupported/Assumed.
- Review independently: QA accepted, edited, added or held coverage according to the frozen source.
- Calculate only approved coverage: fully approved explicit criteria divided by total explicit criteria.

Traceability rules
Four evidence states kept the review honest
The test directly verifies the criterion and stays inside the documented behavior.
A useful test exists, but it misses a label, state, route, dependency or expected result needed for approval.
No saved AI case maps to the explicit criterion.
The test may describe a sensible risk, but the source does not authorize the behavior as a requirement.
AI draft matrix
What the 23 AI-generated cases actually covered
| Draft state | Explicit criteria | Rate | QA interpretation |
|---|---|---|---|
| Covered | 2 | 20% | Back and Skip controls had direct, source-grounded coverage. |
| Partially Covered | 5 | 50% | Tests existed, but the mapping did not fully prove the documented label, field, preview or later-entry behavior. |
| Missing | 3 | 30% | The optional-step presentation and key preview requirements had no mapped test. |
| Unsupported/Assumed | 5 groups | Excluded | Invite lifecycle, transaction-sync permissions, responsive breakpoints and keyboard rules needed clarification before approval. |
No materially duplicate test titles were found in the 23-case set. The problem was not duplication; it was the difference between plausible risk coverage and approved requirement coverage.
Five reviewed examples
Where AI coverage was correct, partial, missing or assumed
| Requirement evidence | AI mapping | QA decision | Final action |
|---|---|---|---|
| Back returns to the previous onboarding step. | Four Back-navigation cases. | Covered | Accepted the positive, partial-data, repeat-click and presentation coverage. |
| Agent name is optional. | One case submitted email with the name blank. | Partially Covered | Added a field-label, visibility and optionality check without assuming submission behavior. |
| A right-side “Your page, taking shape” preview is visible. | No mapped test. | Missing | Added a preview-presence case covering the documented modules. |
| User can invite an agent later after skipping. | A later-invite case existed. | Partially Covered | Held full approval because the later-entry route and expected state were not supplied. |
| Invite supports transaction coordination. | Cases assumed pending, accepted and declined sync states. | Unsupported/Assumed | Converted the behavior into requirement questions instead of publishing it as approved coverage. |
Human QA review
Four additions raised approved explicit coverage from 20% to 90%
The reviewer did not make the AI score better by accepting assumptions. The improvement came from adding narrow, source-grounded cases and holding anything that needed a product decision.
Verifies the optional label, heading and approved transaction-coordination copy.
Verifies agent-name and agent-email fields without inventing a submit action.
Checks the preview heading, profile, credential, contact and agent-related modules visible in the source.
Confirms placeholder cards remain visibly labelled as sample data.
One explicit criterion remained partial: the promise that a user can invite an agent later. The AI produced a reasonable test, but a tester still needs the owned route, entry control and expected post-skip state.

Requirement quality
The strongest output was a list of requirements that were not test-ready
The source screenshot clearly showed fields, copy, Back, Skip and a page preview. It did not define the invite submission control, success state, duplicate behavior, recipient acceptance, transaction-sync permission boundary, expiry or resend rules.
It also did not specify responsive breakpoints or keyboard acceptance criteria. Those are useful testing risks, but a traceability matrix should identify them as open questions rather than silently promoting them into requirements.
- Where can the user invite an agent after skipping?
- What action sends the invite and what confirmation is expected?
- When does transaction access begin: send, acceptance or another approval?
- How should duplicate, self, expired and resent invitations behave?
- Which mobile and keyboard criteria must the implementation meet?
Downloadable evidence
AI Requirements-to-Test Traceability Workbook
The workbook contains six audit tabs: Source Requirements, AI Draft Coverage, QA Review, Missing Coverage, Final Approved Matrix and a formula-driven Summary. It contains sanitized requirement descriptions and test identifiers only—no client name, product name, credentials, emails, URLs or proprietary screenshots.
See how every criterion was classified and how the 20% and 90% reviewed coverage rates were calculated.
Planning is not execution
Traceability establishes planned coverage—not product quality
A requirements matrix cannot prove that a release works. Approved tests still need execution against user workflows, business rules, forms, states and integrations. That is where functional testing services validate real behavior.
After approval, manual QA executes and explores around the documented criteria. Stable approved cases can then be organized for future regression testing. This benchmark does not claim that AI automatically selects regression tests.
Limitations
What this benchmark does not prove
- One onboarding requirement set does not establish universal AI accuracy.
- Traceability coverage does not measure defect detection or release quality.
- The source was a screenshot, so several workflow rules remained undocumented.
- The 23 cases were reviewed for requirement mapping, not executed against the product.
- The 90% final rate applies only to the 10 explicit source criteria; five assumption groups remained held.
- Results should not be read as a promise of 100% acceptance-criteria coverage.
Related evidence
How this study differs from our AI test-planning review
Our earlier AI Test Planning Human QA Review case study measured what testers accepted, edited, removed or added across 160 generated cases. This benchmark asks a different question: did the final plan map back to the requirement evidence?
Product and QA leads evaluating an AI test management platform can also use our AI test management tool evaluation guide to assess review, execution, defect and regression workflows.
AI-assisted planning with QA approval
Get a Human-Reviewed AI Test Plan
Share a PRD, user story, Figma flow, website, web app or mobile requirement set. Testers HUB can use AI-assisted planning and experienced QA review to build structured, traceable test coverage for your release.
Need QA testing support for a similar release?
Tell us about your app, website, game, platform coverage, and launch timeline. Testers HUB will recommend a practical QA scope and quote.