AUTHORED FIXTURE CASE · SYNTHETIC DATA

Enterprise coding-agent rollout

1. The decision

Should we fund a staged rollout of an enterprise coding agent to 5,000 engineers?

Ask
staged_funding: £4,200,000 over 24 months
Deadline
2026-12-01
Sponsor
CTO
Decision owner
cio
Top sensitivity assumption
attribution_factor (±£16,922,975 NPV swing)
Evidence strength (this run's ledger)
6 HIGH, 11 MEDIUM, 9 LOW

2. Evidence map

SUPPORTED

  • ev-001 MEASURED AUTHORED
    Pilot cohort (140 engineers, 12 weeks) accepted 61% of agent-suggested changes without material edits.
  • ev-002 MEASURED AUTHORED
    Median time-to-merge for agent-assisted routine changes: 4.1 hours, vs. 9.6 hours cohort baseline.
  • ev-003 MEASURED AUTHORED
    Post-merge defect rate on agent-assisted changes: 3.2%, vs. 2.9% cohort baseline (not statistically distinguishable at pilot's n).
  • ev-004 MEASURED AUTHORED
    Pilot engineers logged an average of 6.4 agent invocations per working day.
  • ev-005 MEASURED AUTHORED
    Pilot support desk logged 23 agent-related tickets over 12 weeks; 4 concerned suggested changes that would have introduced a regression.
  • ev-006 MEASURED AUTHORED
    Pilot infrastructure cost (inference + tooling): GBP38,400 over 12 weeks for 140 seats.
  • ev-007 MEASURED AUTHORED
    Pilot exit survey: 71% of participating engineers want continued access; 9% want to opt out entirely.
  • ev-008 EXTERNAL_REFERENCE AUTHORED
    Internal security review of the pilot found cost per accepted change of GBP0.72 in review overhead; secrets-handling controls were reviewed and passed; a prompt-injection concern via untrusted code comments was raised and left open pending a follow-up review.
  • ev-009 EXTERNAL_REFERENCE AUTHORED
    Industry survey of 400 engineering orgs: median reported coding-agent adoption at comparable scale is 34% of eligible engineers after 12 months, with wide variance by language/stack mix.
  • ev-010 EXTERNAL_REFERENCE AUTHORED
    Vendor-published benchmark claims 55% cycle-time reduction on greenfield tasks; benchmark methodology does not disclose task selection criteria.
  • ev-011 EXTERNAL_REFERENCE AUTHORED
    Analyst note: enterprise coding-agent contracts in this band typically include a 15-20% annual list-price escalator after the first renewal.
  • ev-012 EXTERNAL_REFERENCE AUTHORED
    Regulatory guidance published this quarter recommends (not yet mandates) documented human review of AI-suggested code changes to regulated systems before merge.
  • ev-013 EXPERT_JUDGMENT AUTHORED
    Principal architect assessment: agent-suggested changes are strongest on well-tested, well-typed codebases and weakest on the legacy monolith's untyped modules, which cover roughly 30% of engineering headcount's day-to-day work.
  • ev-014 EXPERT_JUDGMENT AUTHORED
    Engineering manager judgment across 6 teams: onboarding new engineers with agent access reduced time-to-first-merged-change, but managers could not isolate the agent's contribution from a concurrent onboarding-process change.
  • ev-015 EXPERT_JUDGMENT AUTHORED
    CISO delegate judgment: current pilot scope did not exercise the agent against any system handling regulated customer data, so the security review's findings do not generalise to a full rollout without a further review.

ASSUMED

  • ev-016 FORECAST AUTHORED
    Finance projects fully-loaded support and platform-engineering cost to sustain a 5,000-seat rollout at GBP1.1m-1.6m annually, driven mainly by internal tooling integration work, not licence cost.
  • ev-017 FORECAST AUTHORED
    Adoption-curve forecast: at the pilot's observed opt-out rate, a targeted rollout to 1,200 engineers is projected to reach 65-75% active usage by month 6; an unrestricted enterprise rollout is projected to reach only 40-55% by month 6 due to weaker onboarding support per engineer.
  • ev-018 ASSUMPTION AUTHORED
    Assumed engineering productivity uplift from broad agent access: 18%, range 5-30%, extrapolated from the pilot's cycle-time reduction under an assumption that pilot conditions generalise to the wider engineering population.
  • ev-019 ASSUMPTION AUTHORED
    Assumed attrition-risk reduction from improved engineer experience: 1-2 percentage points off the current 14% annual voluntary attrition rate, based on the pilot exit survey's stated preference for continued access.
  • ev-020 ASSUMPTION AUTHORED
    Assumed training and change-management cost per engineer onboarded: GBP340, based on the pilot's per-seat onboarding spend.
  • ev-021 ASSUMPTION AUTHORED
    Assumed licence unit cost holds flat for 24 months before any renewal escalator applies, despite ev-011's analyst note on typical escalator timing.
    contradicts: ev-011
  • ev-022 INFERENCE AUTHORED
    Inferred from ev-002 and ev-004 together: engineers who invoke the agent more than 5 times per day show a larger median cycle-time reduction than low-frequency users, suggesting usage intensity, not mere access, drives most of the benefit.
  • ev-023 INFERENCE AUTHORED
    Inferred from ev-009 (industry adoption survey) and ev-017 (internal adoption forecast): this organisation's projected adoption trajectory sits below the reported industry median, consistent with its higher share of legacy/untyped codebase per ev-013.

UNKNOWN

  • ev-024 UNKNOWN AUTHORED
    No agreed methodology exists yet for attributing observed cycle-time or defect-rate changes at enterprise scale specifically to agent usage versus concurrent process changes (see ev-014).
  • ev-025 UNKNOWN AUTHORED
    Long-run behaviour of agent-suggested code quality as the underlying model provider updates its model version is not established; the pilot ran on a single fixed model version throughout.
  • ev-026 UNKNOWN AUTHORED
    Real-world incident rate for the open prompt-injection concern (ev-008) at full rollout scale is unknown; the pilot recorded zero exploited incidents, but pilot scope and duration may be too limited to observe a low-frequency event.

3. Independent positions

cfo

DEFER LOW
  • The NPV range is too wide, indicating high uncertainty about the financial viability of this project. [ev-018, ev-019]
  • There are significant unknowns and assumptions that need further investigation before proceeding with a full-scale rollout. [ev-024, ev-025]
Blocking unknowns:
  • Long-run behaviour of agent-suggested code quality as the underlying model provider updates its model version is not established.
  • No agreed methodology exists yet for attributing observed cycle-time or defect-rate changes at enterprise scale specifically to agent usage versus concurrent process changes.

cto

DEFER MEDIUM
  • The pilot's codebase and team mix may not fully represent the organization's engineering estate, particularly in terms of legacy/untyped codebases which cover a significant portion of daily work. [ev-013, ev-024]
  • The projected adoption trajectory for an unrestricted enterprise rollout is below industry median due to weaker onboarding support per engineer, suggesting potential lower-than-expected benefits from full-scale deployment. [ev-009, ev-017]
  • There are significant unknowns regarding the long-term behavior of agent-suggested code quality and security concerns that have not been fully addressed in the pilot scope, particularly concerning systems handling regulated customer data. [ev-025, ev-015]
Blocking unknowns:
  • Long-run behaviour of agent-suggested code quality as model versions update
  • Real-world incident rate for the open prompt-injection concern at full rollout scale

ciso

DEFER MEDIUM
  • The pilot's security review raised a prompt-injection concern that remains unresolved, and the review did not cover systems handling regulated customer data. [ev-008, ev-012]
  • There is no established methodology to attribute observed cycle-time or defect-rate changes at enterprise scale specifically to agent usage versus concurrent process changes. [ev-024]
Blocking unknowns:
  • Real-world incident rate for the open prompt-injection concern at full rollout scale is unknown; the pilot recorded zero exploited incidents, but pilot scope and duration may be too limited to observe a low-frequency event.
  • Long-run behaviour of agent-suggested code quality as the underlying model provider updates its model version is not established.

cro_compliance

DEFER MEDIUM
  • The pilot results show a positive impact on cycle time and engineer satisfaction, but the evidence does not fully address concerns around security for regulated systems and the potential long-term behavior of agent-suggested code quality. [ev-002, ev-015, ev-025]
  • There is a regulatory guidance recommending documented human review of AI-suggested changes to regulated systems, which needs to be addressed before scaling the rollout. [ev-012]
  • The adoption forecast suggests that a full enterprise rollout may not achieve the same level of active usage as observed in the pilot due to weaker onboarding support per engineer. [ev-017]
Blocking unknowns:
  • Long-run behavior of agent-suggested code quality with model version updates
  • Real-world incident rate for open prompt-injection concern at full rollout scale

business_executive

DEFER MEDIUM
  • The pilot's cycle-time reduction and usage metrics are promising, but they do not fully account for the complexity of a full-scale rollout. The projected adoption rate is below industry norms, suggesting potential challenges in scaling. [ev-002, ev-017]
  • The security review's findings are limited to the pilot scope and do not cover systems handling regulated customer data. A full-scale rollout would require a more comprehensive security assessment, especially given the open prompt-injection concern. [ev-008, ev-015]
  • The attrition-risk reduction benefit is based on a low-strength assumption and should not be treated as a certain financial benefit. The pilot exit survey's stated preference for continued access does not necessarily translate to productivity gains at scale. [ev-019, ev-007]
Blocking unknowns:
  • The long-run behavior of agent-suggested code quality as the underlying model provider updates its model version is unknown (ev-025).
  • The real-world incident rate for the open prompt-injection concern at full rollout scale is unknown, despite no incidents observed in the pilot (ev-026)
NPV lowNPV midNPV highPaybackPeak funding
-£5,351,240 £5,085,455 £38,138,347 0.0y £0
Tornado (assumption swing on NPV, most sensitive first)
attribution_factor£16,922,975
uplift£15,669,421
fully_loaded_cost_gbp£3,760,661
training_cost_per_engineer_gbp-£507,769
annual_support_cost_gbp-£413,223
Staged-funding ladder:
  1. Discovery: £50,000
  2. Pilot: £420,000
  3. Targeted scale: £1,600,000

4. What changed minds

No belief updates recorded for this run (no revision phase in this condition).

5. Decision record

Baseline B has no synthesized recommendation: deterministic aggregation only, no chair/narrative step.

DEFER (5/5 members)
Unioned blocking unknowns:
  • Long-run behaviour of agent-suggested code quality as the underlying model provider updates its model version is not established.
  • No agreed methodology exists yet for attributing observed cycle-time or defect-rate changes at enterprise scale specifically to agent usage versus concurrent process changes.
  • Long-run behaviour of agent-suggested code quality as model versions update
  • Real-world incident rate for the open prompt-injection concern at full rollout scale
  • Real-world incident rate for the open prompt-injection concern at full rollout scale is unknown; the pilot recorded zero exploited incidents, but pilot scope and duration may be too limited to observe a low-frequency event.
  • Long-run behavior of agent-suggested code quality with model version updates
  • Real-world incident rate for open prompt-injection concern at full rollout scale
  • The long-run behavior of agent-suggested code quality as the underlying model provider updates its model version is unknown (ev-025).
  • The real-world incident rate for the open prompt-injection concern at full rollout scale is unknown, despite no incidents observed in the pilot (ev-026)
HUMAN DECISION

No human decision recorded yet for this run.