A pre-registered test of whether multi-agent AI deliberation improves decision hygiene over a single strong AI context.
Kriterion runs the same evidence-first investment case through four conditions: one continuous AI context, five independent one-shot views, a full five-seat committee with structured challenge and preserved dissent, and a heterogeneous-model variant of that committee. The pre-registered headline comparison below covers the first three (A, B, C); Treatment D is a later, exploratory stretch condition outside that criterion. Every deterministic number (NPV, payback, tornado sensitivity) is computed by pure code before any model call; the models judge the case, they never compute it.
| Condition | Contexts | Protocol | Models |
|---|---|---|---|
| A: Baseline | 1 | Full protocol, collapsed into one context | qwen2.5:14b-instruct |
| B: Baseline | 5 | Independent one-shot views, deterministic aggregation, no interaction | qwen2.5:14b-instruct |
| C: Treatment | 5 | Full committee: independent assessment → structured challenge → revision → synthesis with preserved dissent | qwen2.5:14b-instruct |
| D: Treatment | 5 | Same as C | 5 distinct models, one per seat |
Case C (the negative-NPV case) is provisional: the perturbation-pair runs that produce the third sub-metric were only ever run on Case A, per the pre-registered design, and only once per perturbation type, at a single seed, not repeated across seeds. Neither case's result implements the full pre-registered margin (must beat by >2× seed-to-seed standard deviation) or the inference-cost condition (must cost >2× Baseline A); both are disclosed gaps in the comparison output itself, not hidden in this summary. What's reported above as "True" is this weaker, bare-aggregate check, not the full pre-registered criterion: adding the missing margin could only reduce how many sub-metrics count as a win for C, but the missing cost measurement means the full criterion's actual verdict remains genuinely unevaluated, not just unconfirmed.
Framing changed the single-context baseline's decision. It did not change the committee's.
The same evidence, reworded in a more optimistic tone with every number and fact held identical,
was run against Baseline A and Treatment C on the same case. Baseline A's recommendation flipped from
DEFER to PILOT: a funding-adjacent decision, changed by tone alone.
Treatment C's recommendation did not move. Across all four perturbation types tested (optimistic
reframing, a stated CEO endorsement, a misleading headline number, and evidence reordering), the
single-context baseline drifted on 1 of 4; the five-context committee drifted on 0 of 4. Each
perturbation type was run once, at a single seed: a concrete illustration of drift under one input
change, not a result replicated across seeds.
Heterogeneity didn't fix dissent-preservation. It made one model's unreliability measurable.
Under Condition D, each of the five committee seats ran a different model (3B to 24B parameters; only the CFO seat kept Baseline A/B/C's default 14B model). Across five seeds, the CISO seat (a 7B model, not the smallest of the five) held its genuine minority position twice and folded fully into the majority three times, and in every fold, it cited only evidence that was already available at the start, never anything new. That's a real, quantifiable reliability gap for that specific model/seat pairing under this protocol, not an argument that mixing models generally is either the problem or the fix.
Single continuous context, all five charters collapsed.
Five independent one-shot views, no interaction.
The full committee: challenge, revision, preserved dissent.
The same committee, five different models: this run (seed 0) is one where the CISO seat held; three of five seeds folded (see finding above).
Full comparison: Case A (all 3 sub-metrics)
Full comparison: Case C (provisional, 2 of 3)
Every real bug found while building this, every design decision and why it was made, and the full task-by-task build log live in the project's decision log. The pre-registered criteria themselves (written before any comparison batch ran) are in the experiment plan. Nothing here supports a claim about decision quality: these are authored fixture cases measuring trap detection and evidence discipline, not real-world outcomes.