We mapped 23 AI-generated test cases back to a frozen, sanitized production requirement set. Experienced QA found that only 2 of 10 explicit acceptance criteria had complete draft coverage. Five were partial, three were missing, and five additional test groups depended on behavior the source never specified. After review and four targeted additions, 9 of 10 explicit criteria had approved coverage.

Benchmark result

AI draft coverage was 20%; final human-reviewed coverage was 90%

This benchmark measures requirements-to-test traceability, not the number of scenarios an AI can produce. The source was one real onboarding requirement screenshot already analyzed through Testers HUB AI. Client, product, account, email, URL and proprietary workflow identifiers were removed before publication.

10explicit source criteria
23AI-generated test cases reviewed
20%AI criteria fully approved
90%final criteria fully approved

Important denominator: the percentages use only the 10 criteria visible or stated in the frozen source. Five useful AI-generated extensions were recorded separately as unsupported or assumed and were not counted as approved requirements.

Method

How the AI requirements traceability experiment was run

  1. Freeze the evidence: one sanitized onboarding requirement screenshot captured on August 26, 2026.
  2. Extract source assertions: 10 explicit acceptance criteria and five important requirement gaps or quality questions.
  3. Preserve the AI draft: 23 saved cases—13 positive and 10 negative—before independent coverage approval.
  4. Map by criterion: every test was traced to Covered, Partially Covered, Missing or Unsupported/Assumed.
  5. Review independently: QA accepted, edited, added or held coverage according to the frozen source.
  6. Calculate only approved coverage: fully approved explicit criteria divided by total explicit criteria.
Reviewed coverage rate = fully approved explicit criteria ÷ total explicit source criteria
Testers HUB AI requirement analysis and test coverage planning screen
The platform analyzes product and requirement context before an experienced tester approves coverage.

Traceability rules

Four evidence states kept the review honest

Covered

The test directly verifies the criterion and stays inside the documented behavior.

Partially Covered

A useful test exists, but it misses a label, state, route, dependency or expected result needed for approval.

Missing

No saved AI case maps to the explicit criterion.

Unsupported/Assumed

The test may describe a sensible risk, but the source does not authorize the behavior as a requirement.

AI draft matrix

What the 23 AI-generated cases actually covered

Draft state Explicit criteria Rate QA interpretation
Covered 2 20% Back and Skip controls had direct, source-grounded coverage.
Partially Covered 5 50% Tests existed, but the mapping did not fully prove the documented label, field, preview or later-entry behavior.
Missing 3 30% The optional-step presentation and key preview requirements had no mapped test.
Unsupported/Assumed 5 groups Excluded Invite lifecycle, transaction-sync permissions, responsive breakpoints and keyboard rules needed clarification before approval.

No materially duplicate test titles were found in the 23-case set. The problem was not duplication; it was the difference between plausible risk coverage and approved requirement coverage.

Five reviewed examples

Where AI coverage was correct, partial, missing or assumed

Requirement evidence AI mapping QA decision Final action
Back returns to the previous onboarding step. Four Back-navigation cases. Covered Accepted the positive, partial-data, repeat-click and presentation coverage.
Agent name is optional. One case submitted email with the name blank. Partially Covered Added a field-label, visibility and optionality check without assuming submission behavior.
A right-side “Your page, taking shape” preview is visible. No mapped test. Missing Added a preview-presence case covering the documented modules.
User can invite an agent later after skipping. A later-invite case existed. Partially Covered Held full approval because the later-entry route and expected state were not supplied.
Invite supports transaction coordination. Cases assumed pending, accepted and declined sync states. Unsupported/Assumed Converted the behavior into requirement questions instead of publishing it as approved coverage.

Human QA review

Four additions raised approved explicit coverage from 20% to 90%

The reviewer did not make the AI score better by accepting assumptions. The improvement came from adding narrow, source-grounded cases and holding anything that needed a product decision.

HQA-001: Optional step and purpose copy

Verifies the optional label, heading and approved transaction-coordination copy.

HQA-002: Field labels and optionality

Verifies agent-name and agent-email fields without inventing a submit action.

HQA-003: Preview presence and modules

Checks the preview heading, profile, credential, contact and agent-related modules visible in the source.

HQA-004: Sample versus live content

Confirms placeholder cards remain visibly labelled as sample data.

One explicit criterion remained partial: the promise that a user can invite an agent later. The AI produced a reasonable test, but a tester still needs the owned route, entry control and expected post-skip state.

Testers HUB AI human-reviewed manual execution and QA summary screen
Approved coverage can move into manual execution, defects, retesting and release reporting without treating the AI draft as final.

Requirement quality

The strongest output was a list of requirements that were not test-ready

The source screenshot clearly showed fields, copy, Back, Skip and a page preview. It did not define the invite submission control, success state, duplicate behavior, recipient acceptance, transaction-sync permission boundary, expiry or resend rules.

It also did not specify responsive breakpoints or keyboard acceptance criteria. Those are useful testing risks, but a traceability matrix should identify them as open questions rather than silently promoting them into requirements.

  • Where can the user invite an agent after skipping?
  • What action sends the invite and what confirmation is expected?
  • When does transaction access begin: send, acceptance or another approval?
  • How should duplicate, self, expired and resent invitations behave?
  • Which mobile and keyboard criteria must the implementation meet?

Downloadable evidence

AI Requirements-to-Test Traceability Workbook

The workbook contains six audit tabs: Source Requirements, AI Draft Coverage, QA Review, Missing Coverage, Final Approved Matrix and a formula-driven Summary. It contains sanitized requirement descriptions and test identifiers only—no client name, product name, credentials, emails, URLs or proprietary screenshots.

Review the complete mapping

See how every criterion was classified and how the 20% and 90% reviewed coverage rates were calculated.

Download the Traceability Workbook

Planning is not execution

Traceability establishes planned coverage—not product quality

A requirements matrix cannot prove that a release works. Approved tests still need execution against user workflows, business rules, forms, states and integrations. That is where functional testing services validate real behavior.

After approval, manual QA executes and explores around the documented criteria. Stable approved cases can then be organized for future regression testing. This benchmark does not claim that AI automatically selects regression tests.

Limitations

What this benchmark does not prove

  • One onboarding requirement set does not establish universal AI accuracy.
  • Traceability coverage does not measure defect detection or release quality.
  • The source was a screenshot, so several workflow rules remained undocumented.
  • The 23 cases were reviewed for requirement mapping, not executed against the product.
  • The 90% final rate applies only to the 10 explicit source criteria; five assumption groups remained held.
  • Results should not be read as a promise of 100% acceptance-criteria coverage.

AI-assisted planning with QA approval

Get a Human-Reviewed AI Test Plan

Share a PRD, user story, Figma flow, website, web app or mobile requirement set. Testers HUB can use AI-assisted planning and experienced QA review to build structured, traceable test coverage for your release.

Get a Human-Reviewed AI Test Plan

Need QA testing support for a similar release?

Tell us about your app, website, game, platform coverage, and launch timeline. Testers HUB will recommend a practical QA scope and quote.