← Composa

How Composa works

Methodology, validation status, and the ROI math · Last updated 2026-07-18 · Ousios LLC

This page exists because you shouldn't have to take our word for anything. It covers what the assessment measures, how scoring works, what we consider validated versus provisional, and exactly how the dollar figures are computed — including the discounts we apply so they stay honest.

What the assessment measures

The 12-question assessment measures five behavioral axes — how you initiate work, respond to change, relate to structure, engage with people, and orient toward outcomes. Questions describe concrete work situations ("a new project lands on the team — your instinct is to…"), not self-image adjectives. Your position across the five axes maps to one of 11 working-style archetypes, grouped into families of adjacent styles.

Working styles describe how you tend to operate right now — they are contextual, they can shift with role and season, and the strategy re-computes whenever your team's mix changes. They are not personality types, not aptitude scores, and never a basis for evaluating anyone.

Scoring is deterministic — and says so when it's unsure

The same answers always produce the same result: scoring is a fixed computation over the five axes, with no AI model and no randomness in the loop. When your top two archetypes are close, the result says so — you'll see a confidence indicator and a family-level result rather than false precision.

Where Composa sits among frameworks

What we build on — the non-proprietary, academic tradition: the five continuous behavioral axes follow the five-factor lineage that made the Big Five the best-validated general model of individual differences; our validation study uses the public-domain Mini-IPIP as its reference measure; the team-health score uses the published Lau & Murnighan (1998) faultline measure; and the pairing model rewards complementarity because that is what team-composition research keeps finding matters.

The concession, up front: if what you want is the best-validated measurement of personality, use the Big Five — instruments like the BFI-2 and Mini-IPIP are free and carry decades of peer-reviewed evidence. Composa will not out-validate them and doesn't try. Composa's job is different: it outputs a working plan — AI workflows matched to how you already operate, and a team-composition conversation — not a percentile. You leave with something to run on Monday, which is the thing a trait score doesn't give you.

What we are not: Composa is a different product from the established commercial type and role instruments — MBTI®, DiSC®, Belbin® Team Roles — and is not affiliated with, endorsed by, or derived from any of them. Where results show an MBTI or Enneagram label, it is decorative narrative context only (see the ladder below): flagged in-product as not validated, with zero role in scoring or recommendations.

The honest mirror: Composa also names types — so isn't that the same move the type instruments get criticized for? Partly, yes: any typology trades nuance for usability. Our mitigations are structural: the continuous axes are always shown, the family layer is the load-bearing claim (the archetype is labeled directional), and every result states its own confidence — including "close call." The type is the conversational handle; the axes are the measurement.

MBTI® and Myers-Briggs® are trademarks of The Myers & Briggs Foundation. DiSC® is a trademark of John Wiley & Sons, Inc. Belbin® is a trademark of Belbin Associates. All are named here for identification and comparison only.

Validation status — the honest ladder

Different layers of the result have different evidential standing. We label them rather than blur them:

Most reliable · family layer

The working-style family (groups of adjacent archetypes) is the layer we lead with. In our internal consistency analysis of assessment responses, family-level assignment is recoverable from answer patterns roughly 84% of the time — versus ~51% at the finer 11-archetype level. That's why results emphasize the family and show confidence for the specific archetype.

Directional · archetype layer

The specific archetype within a family is directional: useful for conversation and workflow matching, less stable than the family. An external validation study is pre-registered below — this page will link the results when they're in.

Decorative · narrative extras

Optional personality-framework and symbolic overlays (MBTI, Enneagram, and similar) are shown as narrative context only, clearly flagged in-product as "not psychology and not validated." They play zero role in scoring or in any recommendation.

Every AI-workflow recommendation is computed exclusively from the five behavioral axes — nothing decorative feeds the strategy.

The validation study — gates published before data

We are validating the instrument on an external sample (N = 250, recruited via Prolific), and we froze the pass/fail bars before collecting any data, so the result cannot quietly bend the standard it's judged by. The frozen gates:

Family recoveryFamily-level assignment recoverable from answer patterns in ≥ 80% of responses. Our internal figure is ~84% — on a synthetic ceiling. External recovery is expected to come in lower, which is why this gate is a genuine risk and not a formality.
Axis reliabilityCronbach's α ≥ .70 on each of the five axis scales.
Convergent / discriminantThe Composure overlay must correlate with the established measure it should track (r ≤ −.50 vs Mini-IPIP Emotional Stability) and must not correlate with the ones it shouldn't (|r| < .35).

The pledge, signed 2026-07-18: we will publish the results against these gates whether they pass, fail, or split — per gate, with every exclusion counted, the internal figures shown beside the external ones, and an independent reviewer's assessment linked alongside. A mixed result gets reported as mixed; a failed gate means the claim it supports is withdrawn from this page, dated. The framings for all three outcomes are already written — before the data — so the words can't bend to the result.

The full pre-registration (sample and stop rule, exclusion rules, the mixed-case adjudication, the pre-written outcome framings) is a fixed document that will be externally timestamped on OSF before recruiting begins — a page we host can be edited; a timestamped registration can't. The link to the timestamped copy will appear here. One more cap, stated now: a single N=250 sample can make the family-layer claim externally sampled; it will not make Composa "a validated instrument," and we won't use that phrase on the strength of this study alone.

Team composition health, exactly

The team score is a weighted blend of three published components — health = 0.6 × coverage + 0.4 × (1 − faultline):

The score deliberately rewards complementarity, never similarity — a team of identical high performers scores low on purpose, because that's what the composition research says to worry about. It is a conversation instrument, not a grade.

The ROI math, exactly

The dollar estimate is deliberately discounted — it models realistically recoverable value, not a theoretical ceiling:

value / yr = hours saved per week × team size × hourly rate × 40% capture × adoption rate
Hours savedPer-workflow weekly estimates, reconciled against each workflow's own worked scenario (e.g. "a 90-minute Monday report becomes 15 minutes" = 1.25 hrs).
Working weeks47 per year — not 52.
SalaryYou set it (default $85k). Hourly rate derives from a 2,080-hour paid year.
40% captureNominal time saved never converts fully into recovered output — context-switching, verifying AI output, partial reallocation. We keep only 40%.
Adoption rateEach workflow carries its own adoption assumption (55–85%) — not everyone on a team picks up every workflow.

What this is: a structured planning reference for prioritizing where AI helps first. What this is not: a forecast, a benchmark, or a guarantee. Actual results depend on implementation quality and team adoption — the two things a calculator cannot know.

What Composa is not for

Questions we haven't answered

Start with What Composa won't tell you — the published list of known limits and failure modes. If something still isn't covered, ask: hello@composa.team. Material changes to the methodology will be noted on this page with dates.