← AI trust center

Validation & testing

Last updated: 14 July 2026

Testing program adopted 14 July 2026 — first cycle scheduled

Validation starts with being precise about the claim. Most assessment tools fail diligence not because their testing is weak but because their claims are bigger than any testing could support. This page states what Stori claims, what it deliberately does not claim, and how the claims it does make are checked.

The claims — and the non-claims

We claim

The interview elicits and organizes role-relevant evidence from the candidate's own experience; scores express the presence, specificity, and checkability of that evidence, with stated confidence and abstention where evidence is thin; and the two lenses (interview evidence and self-report) are surfaced separately so convergence and divergence can be weighed by a human.

We do not claim

That any score predicts job performance (no predictive-validity claim is made anywhere); that Stori identifies “the best candidate”; that personality measures ability; or that the AI's interpretation should be trusted without checking it against its source.

Why the smaller claim

Performance depends on a future no assessment can see, and “performance” itself is rarely measured well enough to validate against — the criterion problem. A tool that claims prediction owes evidence it cannot honestly produce. We claim evidence organization, and we can demonstrate exactly that.

The basis for each instrument

  • The interview (evidence lens): a content-based measure — it samples what the candidate has actually done, in their own words, against criteria the employer selects for the role. Job-relatedness is anchored to the employer's stated criteria, not to a universal template of a “good candidate.”
  • The self-report (context lens): a forced-choice instrument built on the Big Five — the most widely validated and empirically supported personality framework across cultures. It claims to describe self-reported disposition, nothing more, and never contributes to ranking, prioritization, filtering, or exclusion.

The testing program

Adopted 14 July 2026; the tests below are designed to catch the specific ways an interview-based instrument could actually fail, and are run on every material change to the evaluation prompts or model versions, and at least quarterly. Results are logged with date, scope, model version, and the response taken.

  • Delivery-invariance: the same facts expressed with added disfluencies, simplified vocabulary, or non-native syntax must score the same. Speech style must earn nothing and cost nothing.
  • Text-vs-voice parity: the same story typed or spoken must score the same — the precondition for an alternative-format interview being a true equivalent.
  • Abstention behavior: “insufficient signal” must track the thinness of evidence, not the style of the speaker.
  • Length-robustness: verbosity itself must earn nothing once evidence is held constant.
  • Input audit: every input that reaches scoring is enumerated and checked against protected characteristics and their proxies.

Status, stated plainly

The methodology is adopted and published in this form; the first full test cycle is scheduled and its dated results will be reflected on this page. We state testing exactly as it stands — what has run, when, and what was found — and nothing beyond it. Group-level statistics on protected classes require demographic data Stori does not currently collect; the approach to that, including why we refuse to infer demographics, is described in the bias risk assessment.

This document describes how Stori is designed. It is not legal advice and does not determine whether any specific employer deployment is lawful. Questions from legal or procurement teams: support@onestori.com.