Known limitations
Last updated: 14 July 2026
Every assessment instrument has limits, and the trustworthy ones state theirs. These are Stori's — the things the system cannot see, cannot verify, and deliberately does not attempt — together with how the design responds to each.
The limits
Evidence is not truth
Stori weighs how specific and checkable a candidate's claims are; it does not verify that they are true. “Cut costs 18% over two quarters” scores well because it is checkable — the checking is the employer's (references, records, the interview they run themselves). We say “information,” not “proof,” for exactly this reason.
What isn't said isn't scored
The interview reads the stories a candidate chooses to tell. Real capability that never surfaces in the narrative goes unmeasured. The design response is abstention: an unmeasured facet is reported as “insufficient signal,” never scored down — but employers should know a gap in the grid means “not observed,” not “absent.”
Facet coverage is uneven by design
A comfortable, authentic story is worth more than full trait coverage bought with contrived questions, so some facets surface rarely. Confidence levels and abstentions make the unevenness visible instead of hiding it.
No forecast of performance
Job performance depends on a future no assessment can see — the reorg, the manager change, the life event. Stori makes no predictive claim, and any use of its output as a performance forecast is use outside the claim.
The self-report is a self-report
The Big Five instrument captures how a candidate describes their own disposition. It is ipsative (scores are relative to the person's own other traits), it does not measure ability, and it can be shaped by self-perception. That is why it is context and self-insight only, and never ranks, filters, or gates anyone.
AI interpretation can be wrong
The AI's reading of a transcript is an interpretation, and interpretations err. This is the reason receipts exist: every interpretation links to its source so a human can catch a wrong reading and overrule it. You should not trust the interpretation alone — the design assumes you won't.
Interview performance varies
People differ in how comfortably they narrate their experience, and an interview samples one sitting of one person. Adjustments — extra time, restarts, an alternative format — are available on request before the interview, and evaluation never scores delivery, only substance.
Language coverage
Evaluation quality is strongest in English. Non-native phrasing carries no penalty by design — delivery is never scored — but nuance in other languages may surface less evidence than the same story told natively. Employers assessing multilingual pools should weigh this.
Why we publish this
A limitations page is not a disclaimer; it is calibration. Several regimes now require developers to disclose known limitations to the organizations deploying their systems, and any employer relying on an output deserves to know how much weight it can bear. The short version of this entire page: Stori improves the evidence a hiring decision rests on — it does not remove the uncertainty that makes it a decision.
This document describes how Stori is designed. It is not legal advice and does not determine whether any specific employer deployment is lawful. Questions from legal or procurement teams: support@onestori.com.