800+ assessments.
Which one is right?
Google "assessments" and you get 800+ of them, growing every year. Most were never built for hiring. This is a fair, fully sourced comparison of eleven tools you'll actually be pitched — grouped into four tiers by what the data, and each publisher, actually says.
The short answer: the assessment is one ingredient of five. Trueseat is the methodology that plugs into any of them.
Two questions every
assessment must answer.
Before you pick a tool for your next hire, run it through this test. It rules out about half the assessments on the market — and it is the reason the eleven below fall into four very different tiers.
The bar every hiring tool has to clear.
Every assessment either predicts real-world job performance or it does not. These two questions separate the tools built for hiring from the tools built for coaching, development, team dynamics, or self-awareness — all legitimate, none of them selection.
Predictive validity
If a candidate scores high, does that score correlate with actual job performance? Measured as a correlation coefficient (r) between 0 and 1.
Above 0.30 is meaningful. Below 0.20 is closer to noise than signal.Test-retest reliability
If the same person takes the assessment twice, do they get similar results? Measured as a correlation between the two scores.
Above 0.70 is acceptable. Above 0.90 is exceptional.The 2026 validity hierarchy
From Schmidt & Hunter (1998) and the Sackett et al. (2022) recalibration — the numbers peer-reviewed I/O psychology accepts as the current state of the science.
| Method | Validity (r) | What it predicts |
|---|---|---|
| Structured behavioral interview | ≈0.42 | The single best individual predictor. |
| Job knowledge tests | ≈0.40 | Role-specific mastery. |
| General mental ability (cognitive) | 0.30–0.40 | Learning speed and problem-solving. |
| Work samples | 0.30–0.35 | Actual job performance in miniature. |
| Personality / behavioral assessment | 0.15–0.30 | Behavioral fit to role demands. |
| Unstructured interview | ≈0.20 | A coin flip in a nice shirt. |
| Years of experience | ≈0.18 | Surprisingly weak on its own. |
| Educational credentials | ≈0.10 | Near-noise for most roles. |
Scroll the table sideways to see every column.
The most-validated person-side predictor in the science.
The single most important variable in most hiring conversations is whether the assessment includes a measure of general mental ability. This is where a lot of otherwise-good tools fall short.
General mental ability (GMA) — sometimes called g-factor or cognitive ability — measures how fast a person learns, how well they handle ambiguity, and how effectively they solve problems they have not seen before. It is measured by tests like the Wonderlic (used by the NFL Combine), the Criteria CCAT, and the PI Cognitive Assessment.
Why it matters for hiring. Sixty years of meta-analytic research (Schmidt & Hunter 1998 onward) has consistently found GMA to be the strongest individual predictor of job performance for complex roles. Its validity does not collapse across cultures, industries, or role types the way personality-based measures sometimes do. For roles above roughly $75K base — foreman, sales manager, operations lead, technical specialist — cognitive ability adds meaningful predictive lift on top of any behavioral assessment. For simpler or highly routinized roles, it matters less.
Only three of the eleven include cognitive ability
Of every assessment in this guide, only Predictive Index (built-in, PI Cognitive), Bryq (built-in), and Hogan (sold separately) include a cognitive measure. Aptive explicitly avoids it; Kolbe, DISC, Culture Index, CliftonStrengths, MBTI, Working Genius, and Enneagram do not measure it at all. Standalone tools like CCAT and Wonderlic add it on the side. For roles above roughly $75K base — where cognitive predicts performance most strongly — whether cognitive is in the stack is one of the two or three biggest variables.
Assessment is one of five. It is not the system.
A great hire is the product of five distinct ingredients, each with its own contribution to composite validity. The assessment is just one. Trueseat handles the methodology; the assessment plugs into ingredient two — and any of the eleven can fill that slot.
| Ingredient | What it contributes | Standalone validity |
|---|---|---|
| 1. Role definition (Job Target) | Aligns stakeholders on what the role actually demands. | Foundation — required for the rest. |
| 2. Assessment (behavioral + cognitive) | Measures who the candidate is wired to be and how they think. | r ≈0.30–0.40 |
| 3. Structured behavioral interview | Captures actual behavioral evidence against a rubric. | r ≈0.42 (highest single) |
| 4. Panel debrief with values veto | Removes single-interviewer bias. | Multiplier on interview validity. |
| 5. First 90 with manager coaching | Retention layer. Protects the hire from ambiguous onboarding. | 50%+ reduction in 90-day turnover. |
"Any of the eleven assessments in this guide can fill ingredient two. What determines the outcome is whether the other four ingredients are present."
Eleven assessments,
side by side.
Read the scientific rigor first, the capabilities second. Both tables are grouped by tier — validated for hiring at the top, publisher-acknowledged limitations and unvalidated tools below.
Scientific rigor at a glance.
The core scientific attributes of each assessment. Where a cell reads "not published," "not for selection," or "publisher advises against," the tool may still be genuinely useful for other purposes — it simply does not clear the bar for defensible hiring decisions.
| Assessment | Category | Hiring validity | Peer-reviewed base | EEOC defense |
|---|---|---|---|---|
| Predictive Index | Behavioral + Cognitive | 0.30–0.40+ stack | 350+ studies, EFPA-certified | Strongest |
| Hogan (HPI/HDS/MVPI) | Personality (3 layers) | 0.25–0.30 + derailment | 250+ criterion studies | Strong |
| Bryq | 16PF / Big Five + Cognitive | Framework-backed | 16PF · Big Five · Holland (decades) | Adverse-impact tested |
| Aptive Index | 8-attribute behavioral | Not published | 1 study, 400 participants (2024) | Adequate |
| Kolbe A | Conation (execution style) | Non-standard metric | Internal validation extensive | Supportable |
| Myers-Briggs (MBTI) | 16-type typology | Publisher: not for selection | Not built for selection | Publisher advises against |
| CliftonStrengths | Top themes (5 / 34) | Publisher: development, not selection | Gallup: development tool | Not for selection |
| Working Genius | 6-type working energy | Team tool, not selection | Team-dynamics tool | Not for selection |
| Culture Index | Adjective checklist (7 categories) | Not published | None — founder confirmed to NPR | Not documented |
| DISC | Behavioral preferences (4) | ≈0.20 | Not built to SIOP for selection | Not designed for it |
| Enneagram | 9-type typology | Not published | No hiring-relevant research | Not validated |
Scroll the table sideways to see every column.
What each platform actually does.
Beyond the science, assessments differ in what they can do inside a business. Cognitive ability, team-dynamics coverage, leadership reporting depth, and AI-enabled insights are the four capability dimensions that matter most in hiring conversations.
| Assessment | Cognitive test | Team dynamics | Leadership reports | AI insights |
|---|---|---|---|---|
| Predictive Index | Yes (PI Cognitive) | Yes (Inspire) | Extensive | Yes (2026 roadmap) |
| Hogan | Separate purchase | Limited | Extensive | Growing |
| Bryq | Yes (built-in) | Yes | Yes | Yes (bias-audited) |
| Aptive Index | No | Yes (integrated) | Yes | Yes (Aria coach) |
| Kolbe A | No | Yes (Kolbe C) | Yes | Limited |
| Myers-Briggs (MBTI) | No | Yes | Development | Limited |
| CliftonStrengths | No | Yes | Extensive | Growing |
| Working Genius | No | Yes (core) | Team-focused | Limited |
| Culture Index | No | Yes | Moderate | Limited |
| DISC | No | Yes | Basic (vendor-varies) | Varies by vendor |
| Enneagram | No | Yes (self-awareness) | Coaching-focused | Limited |
Scroll the table sideways to see every column.
Tier rules are separated by the heavier lines above. Bryq's capability set is self-reported by the vendor; its scientific placement is discussed in full below.
Four tiers, read
straight off the data.
Not all assessments are hiring assessments. Some were built for coaching, development, team dynamics, or self-awareness — and several publishers say so themselves. Each tool below gets its legitimate use named, whatever tier it lands in.
Built and validated for defensible hiring.
Three assessments meet the bar for defensible hiring decisions at the peer-reviewed level: Predictive Index, Hogan, and Bryq.
Predictive Index (PI Behavioral + PI Cognitive)
Hiring-gradeBehavioral drives (Dominance, Extraversion, Patience, Formality) plus a general mental ability test built to SIOP standards. 350+ criterion validity studies, EFPA-certified 2018. Behavioral r ≈0.25–0.30, cognitive r ≈0.30–0.40; stacked with a Job Target, composite reaches 0.60+. The deepest published criterion-validity base in the category, and the only major hiring instrument that ships cognitive as core, not add-on.
Hogan (HPI / HDS / MVPI)
Hiring-gradeThree instruments measuring bright-side personality (HPI), dark-side / derailment risk (HDS), and motives-values-preferences (MVPI). 250+ peer-reviewed criterion studies. Uniquely catches derailment patterns under stress that other tools miss — the reason it dominates executive selection. No cognitive test in the core suite (Hogan sells cognitive separately). Premium price point; typically $200–$500 per candidate.
Bryq
Hiring-gradeA behavioral assessment on the 16 Personality Factors (16PF) framework aligned with the Big Five (OCEAN) model, plus cognitive ability (verbal, numerical, logical reasoning, attention to detail), skills assessments, and an AI-proficiency assessment across five competency dimensions. It uses Holland Codes (RIASEC) — the same framework the U.S. Department of Labor and O*NET use — and is reviewed regularly by in-house I/O psychologists. Bryq reports SOC 2 Type II and ISO 27001:2022 certification, GDPR compliance, annual third-party AI-bias audits, and documented adverse-impact testing. It ships pre-built role libraries spanning the trades (electricians, HVAC, plumbers, carpenters) and white-collar functions, integrates with 14+ ATS systems including Workday, Greenhouse, SAP, and Lever, supports 12 languages, and lists an enterprise client roster including Deloitte, EY, Samsung, Mercedes-Benz, and Equifax.
In the hiring conversation for specific reasons.
Two assessments earn a place in the conversation on their own terms: Aptive Index, a legitimately-built newer platform, and Kolbe, the only commercial assessment that measures conation.
Aptive Index
Modern entrantEight-attribute behavioral assessment built on modern psychometric methodology. Factor analysis (KMO 0.78–0.89), test-retest above 0.90 on a 200-participant sample, Cronbach's alpha 0.71–0.86. Rigorous construct validity. Criterion validity coefficients not yet published. No cognitive assessment (explicitly avoided). Modern UI, built-in Aria AI coach. Founded ~2023. Mid-tier scientific depth, best-in-class product experience.
Kolbe A Index
SpecializedMeasures conation — the instinctive way a person takes action when striving. Four Action Modes: Fact Finder, Follow Thru, Quick Start, Implementor. Not a personality test and not a cognitive test. Highest test-retest reliability of any commercial assessment (90%+ stability over 20+ years). Internal validation research is extensive; independent peer-reviewed criterion validity is thinner. Kolbe RightFit pairs Kolbe A with Kolbe C (job requirements) for structured hiring analysis.
The publishers themselves say: not for hiring.
Three widely-used tools whose creators state directly that they aren't built for selection. Each has legitimate uses — self-awareness, coaching, team dynamics, post-hire role placement — that the publishers stand behind. Selection is not one of them.
Myers-Briggs (MBTI)
Publisher: not for hiring16-type personality typology developed in the 1940s by Isabel Myers and Katharine Cook Briggs, based on Carl Jung's theory. Roughly two million people take it each year. The publisher's own ethical principle is explicit: "MBTI results should not be used to select or limit anyone on the basis of type" (The Myers-Briggs Company, Ethical Principles). Allen Hammer, a psychologist and former chair of the Myers & Briggs Foundation, put it plainly on NPR: "I don't think the MBTI should ever be used to either select people into an occupation or to promote them" (Hidden Brain, "The Sorting Hat," 2019).
CliftonStrengths (StrengthsFinder)
Publisher: developmentTop five (or top 34) talent themes from a list of 34. Gallup, the publisher, says the assessment was "calibrated for development" and "doesn't encourage or support or defend its use in a selection context" (Gallup). Theme reliability is moderate; predictive validity for job performance is not established, because the tool was not designed for that purpose. Excellent for understanding what work engages people.
Working Genius
Team toolSix-type working-energy framework (Wonder, Invention, Discernment, Galvanizing, Enablement, Tenacity) developed by Patrick Lencioni and The Table Group, launched 2020. It measures where individuals gain energy versus where they get drained, and Lencioni positions it primarily as a team-dynamics tool. His own guidance on hiring is a place-after-hire framing, not a selection one: "Hire them because they are a core value fit, but find out right away what their genius is and put them there" (Lencioni, The 6 Types of Working Genius).
Legitimate for coaching — without hiring validation.
Three tools with real uses in coaching and personal development, but without independent peer-reviewed criterion validity for hiring: Culture Index, DISC, and Enneagram.
Culture Index
Coaching toolA proprietary behavioral survey launched in 2004 by Gary Walstrom, developed in consultation with a psychology professor. It uses an adjective checklist methodology — the same methodology family as Predictive Index, not DISC — and measures seven categories: autonomy, social ability, pace, conformity, logic, ingenuity, and energy units. Founder Gary Walstrom confirmed on the record that the Culture Index has not been through a scientific peer review process (NPR, Hidden Brain, "The Sorting Hat," April 2019). No independent peer-reviewed criterion-validity studies exist at the scale PI (350+) or Hogan (250+) have accumulated; test-retest reliability is not independently verified in the peer-reviewed literature. EEOC defensibility is not documented at the PI or Hogan level.
DISC
Communication frameworkFour behavioral preferences (Dominance, Influence, Steadiness, Compliance). One of the oldest and most widely-deployed frameworks in business training. Standalone predictive validity for job performance is around r ≈0.20 — comparable to unstructured interviews. Not built to SIOP standards for selection. Excellent as a communication and team-building framework; not a hiring instrument.
Enneagram
Self-awarenessNine-type personality typology with ancient roots. No published test-retest reliability suitable for hiring. Type assignments are unstable across retakes in independent research. Widely used in coaching, spiritual formation, and pastoral contexts for self-awareness — where its lack of hiring validation is not a concern.
Where the real
leverage lives.
No assessment alone gets you above r = 0.40. The methodology around it is what pushes composite validity to 0.60 and above — and that methodology is assessment-agnostic.
Any assessment, plus the four other ingredients, equals a defensible hire.
The methodology Trueseat productizes — role definition, structured interviewing, panel debrief, and First 90 onboarding — is assessment-agnostic. The value is in the system, not in the tool.
What the multiplier looks like in practice
| Configuration | Composite validity | Expected hit rate |
|---|---|---|
| Assessment alone | 0.25–0.40 | 50–55% |
| Assessment + structured interview | 0.45–0.55 | 60–70% |
| Full 5-ingredient stack | 0.60+ | 70–80% |
| Gut-feel hiring (no system) | ≈0.15–0.25 | 40–50% |
Trueseat is assessment-agnostic
The Trueseat methodology plugs into whatever assessment you already use or prefer. On PI, Trueseat works with PI. On Hogan for the executive tier and PI everywhere else, Trueseat works with both. Running Bryq for the trades, or a lighter-tier tool at a smaller shop, Trueseat still runs the other four ingredients around it. Any of the eleven can fill ingredient two. The system is where the multiplier lives. The assessment is the ingredient it multiplies.
"The best assessment run inside a bad process still produces bad hires. A moderate assessment run inside the full stack produces defensible ones."
Matching the assessment
to your hire.
A short decision framework for picking the right ingredient two for the seat in front of you.
By company type, by budget, by defensibility need.
Assessment selection is a fit conversation, not a rank-ordering exercise. The best assessment is the one that matches your budget, hiring volume, defensibility exposure, and any assessment you already have in place.
| Your situation | Recommended primary assessment | Why |
|---|---|---|
| 25+ employees hiring 6+ per year | Predictive Index | Unlimited-use subscription pays for itself. Cognitive included. Deepest EEOC defensibility. |
| Executive / VP-level searches | PI + Hogan (both) | PI for baseline, Hogan HDS for derailment risk under stress. |
| Critical single hire (execution risk) | PI + Kolbe (both) | Add Kolbe for conation when the role demands a specific execution style. |
| Trades / blue-collar, higher-volume | Bryq | Pre-built role libraries for the trades, cognitive built in, adverse-impact tested. |
| Roles where AI proficiency matters | Bryq | Built-in AI-proficiency assessment across five competency dimensions. |
| Sub-25 employees, low-volume hiring | Aptive Index | Mid-tier science, modern platform, all-inclusive pricing fits the budget. |
| Already invested in DISC, MBTI, or Culture Index | Keep it; add the stack | Trueseat multiplies whatever assessment is in place. Don't replace it unless you want to. |
| Coaching or development engagement only | CliftonStrengths / Working Genius / Kolbe | Use development-grade tools for their intended purpose. |
| Faith-integrated leadership formation | Enneagram optional | Legitimate for self-awareness and formation. Not for selection. |
Scroll the table sideways to see every column.
Validity figures reflect published I/O psychology (Schmidt & Hunter 1998 onward, with the Sackett et al. 2022 recalibration); they describe the methods, not a measured outcome on any one platform. Publisher positions are quoted from primary sources and linked inline. Ready Aim Climb is a certified Predictive Index consultant; PI is recommended for the largest number of client scenarios. Every other assessment's data speaks for itself. Assessment names and vendor claims are the property of their respective owners.
Not sure which one fits?
That's a sixty-minute conversation.
Whichever of the eleven you use — or none yet — Trueseat runs the other four ingredients around it. Book a working call, bring the seat you need to fill next, and leave knowing exactly which read your decision actually needs.
No funnel, no sequence. A real person, on the phone, working the problem with you.