Methods
How this works, and what it cannot tell you.
Every number this app produces is an inference about a person, so the reasoning behind it is written down here rather than kept in the code. Read this before you use a readiness band to make a decision about someone.
Who it is written for
Early-career individuals — new to the workforce or returning after a significant break with limited experience.
This is the first thing that decides whether a result means anything. Every level, item, and entry bar below is calibrated for this population — applied to experienced workers the bar reads as low, and applied to people well outside it the bar reads as unfair. Neither is the instrument working.
01Content comes from the occupations, not from a template
Module purposes, learning outcomes, rubric anchors, and the creditable option in every judgment item are generated from leveled behavior statements written for the specific occupations inside a sector. The scenario's setting, tools, and the people in the room come from O*NET work context for those same occupations.
The alternative — a generic item with the sector name pasted in — produces construct-irrelevant variance: you end up measuring how well someone reads generic business English. [Messick, S.]
02Distractors are built to be tempting
Each judgment item draws its options from four places: the target behavior at the level being measured, the same skill one level down (partial credit), an adjacent cluster (plausible, off-target), and a hand-written anti-pattern — the shortcut that feels efficient and creates the follow-on problem.
An item with three obviously wrong options measures reading speed. [Motowidlo, S. J., Dunnette, M. D., & Carter, G. W.] [McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P.]
03Levels are attained, not averaged
Evidence is level-anchored: an item written at Proficient tells you about Proficient and nothing else. A level is reached by earning 70% of that level's evidence with every sufficiently-evidenced level beneath it at 60% or better. A strong answer high up does not paper over a failed one lower down.
A level carrying under 2.0 weighted evidence points cannot be claimed and cannot block the levels above it — one unlucky item should not sink a cluster that holds everywhere else. Constructs resting on thin evidence are labeled provisional in the report rather than reported as a confident finding. [AERA, APA, & NCME]
04Design follows the content
Procedural work with one right answer is taught and tested differently from judgment work with a defensible range. Twelve patterns, each matched to what it can actually reach:
- Sector orientation. Setting, employers, entry roles, and the conditions the work happens under.
- Guided practice. Procedural work with a correct method — documentation, setup, handoffs, tools.
- Task simulation. Core job tasks where sequence, timing, and decision points matter.
- Case analysis. Judgment content — root cause, tradeoffs, competing priorities, risk.
- Role-play. Interaction-heavy skills — customer service, escalation, difficult news, handoffs.
- Judgment lab. Durable skills where the right move depends on reading the situation.
- AI-in-the-loop studio. The AI work already happening in this sector's occupations.
- Anchored reflection. Self-management and character content, where self-report needs a check.
- Retrieval drill. Safety rules, terminology, tolerances, codes — content that must be recalled cold.
- Team sprint. Creativity, leadership, and collaboration, which only show up under a shared deadline.
- Portfolio build. Capstone evidence — the thing that proves the rest of the program happened.
- Interview simulation. The last mile — translating what you can do into what you can say you can do.
Time is split across five strands:
- 30% Sector core — The work itself — tasks, tools, settings, and standards of this sector's entry roles.
- 33% Durable skills — The ten domains, weighted by what this sector's occupations actually demand.
- 17% AI at work — The AI use already documented in these occupations, and the judgment it requires.
- 11% Work readiness — Safety, compliance, workplace norms, and turning evidence into a hire.
- 9% Capstone — One artifact and one interview that carry the weight of the whole program.
05Every method has a limit, and the limits are published
Each question in an assessment carries this alongside its scoring key, where the person assigning it can read it:
sjt
Low-fidelity simulation. A written work situation with response options that differ in effectiveness rather than in correctness.
Cannot tell you: Measures knowledge of effective action, which is not the same as taking it under real pressure. Pair with observed performance before a high-stakes decision.
best-worst
Forced-choice discrimination. The respondent identifies the strongest and weakest of several defensible-looking behaviors.
Cannot tell you: Ranking ability can outrun performing ability. Reads discrimination, not execution.
mcq
Retrieval of sector knowledge, cued by a work context rather than a definition.
Cannot tell you: Recognition among options is easier than free recall on the job. Treat as a floor, not a ceiling.
sequence
Ordering task. The respondent reconstructs the order of a procedure.
Cannot tell you: Some real procedures allow more than one defensible order. Where they do, the scoring key over-constrains and should be reviewed.
triage
Prioritisation under a stated constraint. Ranked by consequence of delay.
Cannot tell you: Removing time pressure removes most of the difficulty. The paper version is easier than the shift.
verify
Multi-select audit with a penalty for wrong selections, so guessing everything scores worse than selecting carefully.
Cannot tell you: Selecting the right checks is not running them. Pair with an observed AI-assisted task.
constructed
Short constructed response scored against a published checklist of what the response must do.
Cannot tell you: Self-marked in this build, which inflates scores. The scorer weights self-marked evidence below observed evidence; an instructor override should replace it in a live cohort.
rubric
Behaviorally anchored rating scale. Each criterion's four levels are anchored in occupation-specific behavior statements, not in adjectives.
Cannot tell you: Rater-dependent. Anchors reduce drift but do not remove it; inter-rater agreement should be checked before the scores drive decisions.
artifact
Work sample. A finished product judged against criteria published before the work began.
Cannot tell you: Authorship is hard to verify outside a supervised setting, and one artifact is one occasion.
selfcheck
Anchored self-rating. Reported beside the measured level, never combined with it.
Cannot tell you: Self-assessment correlates weakly with performance and worst among the lowest performers. It carries zero weight in every score this instrument produces.
06Self-ratings are reported, never scored
Learners rate themselves against the same four levels. Those ratings carry zero weight in every score the app produces, and appear beside the measured level so the gap is visible. Self-assessment correlates weakly with measured performance, and worst among the people furthest from the bar. [Dunning, D., Heath, C., & Suls, J. M.] The gap is worth collecting because it is the coaching conversation, not because it is evidence.
07Prior credentials shorten the program only under three rules
Turning a credential into shortened seat time is a transfer inference. The credential says someone did something, somewhere, at some point; the claim being made is that they can do a related thing, in this sector, now. That gap needs its own warrant. [Kane, M. T.]
- Only signed, leveled credentials shorten anything. A résumé line or an unsigned badge personalises the material and buys no time.
- Four constructs can never be credited away — core, character, communication, output-verification. These are the readiness gates. A program that lets someone skip its gates on the strength of a prior badge is not measuring readiness.
- Total shortening is capped at 35%. Past a point, what is left is not an assessment of this sector's work.
A shortened module becomes a challenge check rather than disappearing: the two highest-level items remain, so the transferred claim is still tested here.
08What the instruction covers, and what it leaves to you
Durable skills and AI at work are taught here in full: the concept, the four levels as a ladder, a worked example with the reasoning shown, and the tempting mistake named. Those are teachable in prose, and a participant working alone learns them from this app.
Technical technique is not. The competency data describes what good looks like at four levels and contains no procedures — a statement reading “follow manufacturer protocols and facility checklists” points at a procedure without containing it, and searching a sector's entire bank for procedural markers returns nothing. That is not a gap in the data; occupational frameworks describe work, not curricula.
So the sector core strand teaches the judgment around a technical task — the standard it must meet, what gets verified, what the handoff needs, where it goes wrong — and measures whether someone has the technique. Building the technique happens in a shop, a lab, or a placement. Every core module says so, and an organization can attach its own training to a sector so participants are pointed at it rather than left guessing.
A program adopting this instrument expecting technical training would find the gap halfway through a cohort, which is the worst moment to find it. [Messick, S.]
09Time is a target, not a dose
The fifteen-to-twenty-hour figure is a planning target: what a learner with no prior credentialed evidence usually needs to produce enough evidence to be placed at a level. It is not a duration anyone is required to sit through, and sitting through it does not on its own guarantee a decision can be made.
It covers this program only. Technical training runs alongside it and is not counted in the figure, so a program scheduling a cohort should add its own shop, lab, or placement hours on top rather than treating fifteen to twenty hours as the whole intervention.
It is also time in this program only. Of a typical eighteen hours, about eleven are self-contained — durable skills, AI at work, work readiness — and about six build the judgment around a sector's technical work and assess it. The technique itself is learned in a shop, a lab, or a placement, and those hours are additional. A program that read the target as “eighteen hours and my welders are ready” would be reading it wrong.
What is fixed is the standard of evidence. What varies is the time to reach it. Three mechanisms make that real rather than rhetorical:
- Prior credentials shorten the plan before it starts, under the three rules in section 07.
- Settled constructs stop being asked about. Once a construct's evidence is clear of the attainment threshold by a margin — not sitting on it — further items on it cannot change the result, and the learner is told they can skip them.
- Unsettled constructs get extension items. Where evidence is thin or sits inside the band where one more item could flip the level, the program adds items and runs long. Stopping at the target hour there would mean reporting a level the evidence cannot carry.
The rule underneath is sequential: stop when the decision is stable, continue while it is not. [Kane, M. T.] [AERA, APA, & NCME] After two extension rounds a construct is flagged for an observed task rather than extended again — at that point more items are not the missing ingredient.
A learner may submit with areas unsettled. The report says which, and marks those levels provisional rather than quietly presenting them as findings.
10Reporting is de-identified, and suppression is visible
The reporting view carries no names, no email addresses, and no per-person rows. Any figure resting on fewer than five submitted attempts is suppressed, because a distribution over three people in a cohort of three is a name with extra steps. Suppressed cells are labeled as suppressed, so an empty chart is never read as a finding.
Timing is wall-clock between answers, capped per item. It measures pace, not effort. Read it for pacing problems; do not read it as how hard someone tried.
11Why the language is the workplace's, and what that costs
The scenarios, the rubric anchors, and the level descriptions are written in the language of the occupations they came from. A statement reads “evaluate scope condition and reprocessing outcomes for non-routine instruments” because that is how the work is described where the work happens.
This is a deliberate choice, not an oversight. Knowledge is bound to the context it is learned in, and stripping a task of its working context produces knowledge that does not travel back into the work. [Brown, J. S., Collins, A., & Duguid, P.] Learning an occupation includes learning how its people talk — the language is part of the practice rather than packaging around it. [Lave, J., & Wenger, E.] And transfer is most reliable when the learning context resembles the application context. [Barnett, S. M., & Ceci, S. J.] A learner who has only ever met a simplified version of a handover meets the real one for the first time on their first shift.
What it costs, stated plainly
Measured on exactly what a participant receives — not on anything reserved for instructors — a generated program reads as follows. Adult education targets grade 6 to 8, and a large share of adults read below the level their work documents require. [Kirsch, I. S., Jungeblut, A., Jenkins, L., & Kolstad, A.]
| What the participant reads | Grade | Written by |
|---|---|---|
| Instructional prose | 8.9 | this app |
| Item questions | 10.8 | this app, from source statements |
| Answer options | 16.0 | this app, from source statements |
| Rubric anchors | 20.0 | the source competency data |
| Level descriptions on the ladder | 21.0 | the source competency data |
| Everything a participant reads | 12.0 | — |
The split is the finding. What this app writes lands close to target. What it inherits from the occupational data does not, and a participant reads that too — the level descriptions they place themselves against, and the anchors they score themselves on, are the hardest text in the program.
Those anchors are the instrument. A learner who cannot read one is being measured on reading rather than on the competency — construct-irrelevant variance, and measurement error that falls hardest on the people this exists to serve. [Messick, S.]
What was tried, and what was kept
An automated plain-language rewrite was built and removed. Word substitution moved the anchors from grade 20.4 to 18.9; aggressive sentence splitting reached 17.2. It fails because most of the long words are the technical vocabulary — reprocessing, accreditation, competency frameworks are the names of the things, and replacing them makes the statement wrong. A companion labelled “in plain words” that still reads at graduate level promises an accessibility it does not deliver.
What is in place instead is scaffolding rather than simplification:
Every passage can be heard rather than read. Reading comprehension is decoding multiplied by language comprehension, so where decoding is the limiting factor — which it is for a large share of adult learners — hearing the words removes the barrier almost entirely. [Gough, P. B., & Tunmer, W. E.] The level descriptions, the rubric anchors, the questions and their options, and every instructional passage carry a Listen control.
What it does not fix: a grade-20 sentence heard aloud is still a grade-20 sentence. This removes decoding load, not syntactic complexity. For a learner whose limitation is language comprehension rather than decoding it does very little, and it is not a substitute for rewriting the source statements.
- Technical terms are defined at the point of reading, on tap or by keyboard. Looking one up is never recorded and never affects a result.
- Every item reports its reading level to the instructor, flagged when above target, with what to do — read it aloud, gloss the terms, or observe the behaviour directly.
- Each module opens with what the learner will be able to do, in ordinary language, before any source wording appears.
Whether that scaffolding is sufficient is an open question, and not one this page can settle. It needs a cohort that includes learners at the lower literacy levels. Until then the honest position is that this instrument uses the workplace's language on purpose, that the choice has a documented cost, and that the cost is carried by scaffolding rather than denied.
12What holds it to a standard
Subject matter experts validate this app in production. Sector content, the anti-pattern library, the crosswalks, and every question format are reviewed by people who do the work, and that review continues against live use rather than stopping at launch.
Alongside it: every question carries its scoring key, its provenance, and its published limits to the person assigning it; levels resting on thin evidence are marked provisional rather than reported as findings; and the measurement model is set out in full above rather than held as method.
13What this instrument cannot do
- Item statistics need real responses. Difficulty, discrimination, and distractor performance cannot be established from generated data — a simulation rediscovers whatever response model it was given. These come from a live cohort.
- Inter-rater agreement is unmeasured. Anchors reduce rater drift. They do not remove it. [Smith, P. C., & Kendall, L. M.]
- Self-marked written responses inflate. The scorer weights them below observed evidence; an instructor override should replace them in a live cohort.
- Knowing the right action is not taking it under pressure. Pair judgment items with observed performance before any decision that affects someone. [Roth, P. L., Bobko, P., & McFarland, L. A.]
- The sector crosswalks are unofficial. Prefix rules mapping occupations to sectors, written for this app.
References
Standard sources in personnel assessment and learning science.
- Gough, P. B., & Tunmer, W. E. (1986). Decoding, reading, and reading disability. Remedial and Special Education, 7(1).Reading comprehension is the product of decoding and language comprehension. Where decoding is the limiting factor, hearing the text removes the barrier; where language comprehension is, it does not.
- Brown, J. S., Collins, A., & Duguid, P. (1989). Situated cognition and the culture of learning. Educational Researcher, 18(1).Knowledge is bound to the activity and context in which it is learned. Stripping a task of its working context produces knowledge that does not travel back into the work.
- Lave, J., & Wenger, E. (1991). Situated Learning: Legitimate Peripheral Participation. Cambridge University Press.Becoming competent in an occupation includes learning how its people talk. Language is part of the practice, not packaging around it.
- Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4).Transfer is more reliable the closer the learning context is to the application context. Distance in physical context, functional context, and modality all reduce it.
- Kirsch, I. S., Jungeblut, A., Jenkins, L., & Kolstad, A. (2002). Adult Literacy in America: A First Look at the Findings of the National Adult Literacy Survey (3rd ed.). US Department of Education, NCES.A large share of the adult population reads below the level of the documents their work requires. Text pitched above a reader's level measures reading rather than the thing it intends to measure.
- Motowidlo, S. J., Dunnette, M. D., & Carter, G. W. (1990). An alternative selection procedure: The low-fidelity simulation. Journal of Applied Psychology, 75(6).Written descriptions of work situations, with response options rated for effectiveness, predict job performance without the cost of a full simulation.
- McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P. (2001). Use of situational judgment tests to predict job performance: A clarification of the literature. Journal of Applied Psychology, 86(4).Situational judgment tests show useful criterion validity and add incremental prediction over cognitive ability and personality measures.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3).Retrieving information from memory produces better long-term retention than restudying the same material for the same time.
- Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In Psychology and the Real World.Spacing and effortful retrieval slow apparent progress during practice while improving durable performance.
- Sweller, J., van Merriënboer, J. J. G., & Paas, F. (1998). Cognitive architecture and instructional design. Educational Psychology Review, 10(3).Studying worked examples before independent problem-solving reduces extraneous load and improves acquisition for novices.
- Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2).Rating scales anchored in concrete observed behaviors produce more consistent judgments than scales anchored in adjectives.
- Roth, P. L., Bobko, P., & McFarland, L. A. (2005). A meta-analysis of work sample test validity. Personnel Psychology, 58(4).Work samples — performing an actual task from the job — are among the stronger predictors of job performance.
- Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3).Structured interviews — consistent questions, anchored rating scales, multiple raters — substantially outperform unstructured ones.
- Dunning, D., Heath, C., & Suls, J. M. (2004). Flawed self-assessment: Implications for health, education, and the workplace. Psychological Science in the Public Interest, 5(3).Self-assessments of skill correlate weakly with measured performance, and least well among the lowest performers.
- Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1).Feedback about the task and the process behind it changes performance; feedback about the person generally does not.
- Messick, S. (1995). Validity of psychological assessment. American Psychologist, 50(9).Validity is a property of score interpretation, not of a test. Construct under-representation and construct-irrelevant variance are the two central threats.
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1).A score's use has to be justified as an argument: from observation, to generalisation, to extrapolation, to the decision being made.
- Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3).Improvement depends on practice targeted at a specific weakness with immediate informative feedback, not on time on task.
- AERA, APA, & NCME (2014). Standards for Educational and Psychological Testing.Score reports should state what evidence supports each claim, and should not report a result the evidence cannot carry.
- America Succeeds. Pathsmith™ Durable Skills Framework — ten clusters, four performance levels. Cluster-level content in content/rubric.clusters.json.Defines the durable skill clusters and the four-level performance scale this instrument reports against.
- NSX Competency Framework (2026-07-10) — leveled core competency, durable skill, and AI-at-work statements for 1,016 O*NET occupations.Supplies the occupation-specific behavioral statements each item's creditable options and rubric anchors are written from.
- O*NET 25.1 occupation data (onetcenter.org) — tasks, skills, knowledge, technology, work context, and Job Zones.Supplies the work setting, tools, and task content that make each scenario specific to the sector rather than generic.
The shorter version of this page · The durable skills framework