AUTHORED FIXTURE CASE · SYNTHETIC DATA

Enterprise coding-agent rollout

1. The decision

Should we fund a staged rollout of an enterprise coding agent to 5,000 engineers?

Ask
staged_funding: £4,200,000 over 24 months
Deadline
2026-12-01
Sponsor
CTO
Decision owner
cio
Top sensitivity assumption
attribution_factor (±£16,922,975 NPV swing)
Evidence strength (this run's ledger)
6 HIGH, 11 MEDIUM, 9 LOW

2. Evidence map

SUPPORTED

  • ev-001 MEASURED AUTHORED
    Pilot cohort (140 engineers, 12 weeks) accepted 61% of agent-suggested changes without material edits.
  • ev-002 MEASURED AUTHORED
    Median time-to-merge for agent-assisted routine changes: 4.1 hours, vs. 9.6 hours cohort baseline.
  • ev-003 MEASURED AUTHORED
    Post-merge defect rate on agent-assisted changes: 3.2%, vs. 2.9% cohort baseline (not statistically distinguishable at pilot's n).
  • ev-004 MEASURED AUTHORED
    Pilot engineers logged an average of 6.4 agent invocations per working day.
  • ev-005 MEASURED AUTHORED
    Pilot support desk logged 23 agent-related tickets over 12 weeks; 4 concerned suggested changes that would have introduced a regression.
  • ev-006 MEASURED AUTHORED
    Pilot infrastructure cost (inference + tooling): GBP38,400 over 12 weeks for 140 seats.
  • ev-007 MEASURED AUTHORED
    Pilot exit survey: 71% of participating engineers want continued access; 9% want to opt out entirely.
  • ev-008 EXTERNAL_REFERENCE AUTHORED
    Internal security review of the pilot found cost per accepted change of GBP0.72 in review overhead; secrets-handling controls were reviewed and passed; a prompt-injection concern via untrusted code comments was raised and left open pending a follow-up review.
  • ev-009 EXTERNAL_REFERENCE AUTHORED
    Industry survey of 400 engineering orgs: median reported coding-agent adoption at comparable scale is 34% of eligible engineers after 12 months, with wide variance by language/stack mix.
  • ev-010 EXTERNAL_REFERENCE AUTHORED
    Vendor-published benchmark claims 55% cycle-time reduction on greenfield tasks; benchmark methodology does not disclose task selection criteria.
  • ev-011 EXTERNAL_REFERENCE AUTHORED
    Analyst note: enterprise coding-agent contracts in this band typically include a 15-20% annual list-price escalator after the first renewal.
  • ev-012 EXTERNAL_REFERENCE AUTHORED
    Regulatory guidance published this quarter recommends (not yet mandates) documented human review of AI-suggested code changes to regulated systems before merge.
  • ev-013 EXPERT_JUDGMENT AUTHORED
    Principal architect assessment: agent-suggested changes are strongest on well-tested, well-typed codebases and weakest on the legacy monolith's untyped modules, which cover roughly 30% of engineering headcount's day-to-day work.
  • ev-014 EXPERT_JUDGMENT AUTHORED
    Engineering manager judgment across 6 teams: onboarding new engineers with agent access reduced time-to-first-merged-change, but managers could not isolate the agent's contribution from a concurrent onboarding-process change.
  • ev-015 EXPERT_JUDGMENT AUTHORED
    CISO delegate judgment: current pilot scope did not exercise the agent against any system handling regulated customer data, so the security review's findings do not generalise to a full rollout without a further review.

ASSUMED

  • ev-016 FORECAST AUTHORED
    Finance projects fully-loaded support and platform-engineering cost to sustain a 5,000-seat rollout at GBP1.1m-1.6m annually, driven mainly by internal tooling integration work, not licence cost.
  • ev-017 FORECAST AUTHORED
    Adoption-curve forecast: at the pilot's observed opt-out rate, a targeted rollout to 1,200 engineers is projected to reach 65-75% active usage by month 6; an unrestricted enterprise rollout is projected to reach only 40-55% by month 6 due to weaker onboarding support per engineer.
  • ev-018 ASSUMPTION AUTHORED
    Assumed engineering productivity uplift from broad agent access: 18%, range 5-30%, extrapolated from the pilot's cycle-time reduction under an assumption that pilot conditions generalise to the wider engineering population.
  • ev-019 ASSUMPTION AUTHORED
    Assumed attrition-risk reduction from improved engineer experience: 1-2 percentage points off the current 14% annual voluntary attrition rate, based on the pilot exit survey's stated preference for continued access.
  • ev-020 ASSUMPTION AUTHORED
    Assumed training and change-management cost per engineer onboarded: GBP340, based on the pilot's per-seat onboarding spend.
  • ev-021 ASSUMPTION AUTHORED
    Assumed licence unit cost holds flat for 24 months before any renewal escalator applies, despite ev-011's analyst note on typical escalator timing.
    contradicts: ev-011
  • ev-022 INFERENCE AUTHORED
    Inferred from ev-002 and ev-004 together: engineers who invoke the agent more than 5 times per day show a larger median cycle-time reduction than low-frequency users, suggesting usage intensity, not mere access, drives most of the benefit.
  • ev-023 INFERENCE AUTHORED
    Inferred from ev-009 (industry adoption survey) and ev-017 (internal adoption forecast): this organisation's projected adoption trajectory sits below the reported industry median, consistent with its higher share of legacy/untyped codebase per ev-013.

UNKNOWN

  • ev-024 UNKNOWN AUTHORED
    No agreed methodology exists yet for attributing observed cycle-time or defect-rate changes at enterprise scale specifically to agent usage versus concurrent process changes (see ev-014).
  • ev-025 UNKNOWN AUTHORED
    Long-run behaviour of agent-suggested code quality as the underlying model provider updates its model version is not established; the pilot ran on a single fixed model version throughout.
  • ev-026 UNKNOWN AUTHORED
    Real-world incident rate for the open prompt-injection concern (ev-008) at full rollout scale is unknown; the pilot recorded zero exploited incidents, but pilot scope and duration may be too limited to observe a low-frequency event.

3. Independent positions

cfo

DEFER LOW
  • The NPV range is too wide, indicating high uncertainty about the financial viability of this project. [ev-018, ev-019]
  • There are significant unknowns and assumptions that need further investigation before proceeding with a full-scale rollout. [ev-024, ev-025]
Blocking unknowns:
  • Long-run behaviour of agent-suggested code quality as the underlying model provider updates its model version is not established.
  • No agreed methodology exists yet for attributing observed cycle-time or defect-rate changes at enterprise scale specifically to agent usage versus concurrent process changes.

cto

DEFER MEDIUM
  • Pilot conditions (140 engineers, 12 weeks) may not generalise to the full 5,000-engineer estate due to legacy/untyped codebase gap (30% of work) and weaker onboarding support at scale. Adoption projections (ev-017) suggest targeted rollout is more realistic than enterprise-wide adoption. [ev-013, ev-017]
  • Support and integration cost at 5x pilot scale (GBP1.1m-1.6m annually, ev-016) is significant and not fully offset by productivity gains. Engineering capacity consumed by rollout support displaces other technical priorities. [ev-016, ev-022]
  • Security and regulatory gaps (ev-008, ev-015) remain unresolved for full rollout, including prompt-injection risks and regulated systems review. Long-term model quality stability is unknown (ev-025). [ev-008, ev-015, ev-025]
Blocking unknowns:
  • Real-world incident rate for prompt-injection concern at full scale (ev-026)
  • Methodology to attribute cycle-time/defect-rate changes to agent usage vs. concurrent process changes (ev-024)
  • Long-run code quality stability as model provider updates its version (ev-025)

ciso

REQUEST_EVIDENCE LOW
  • The open prompt-injection concern (ev-008) has not been resolved, leaving unaddressed a potential security risk in the full rollout. [ev-008]
  • The pilot did not exercise the agent against any system handling regulated customer data (ev-015), so it's unclear if the security review's findings generalise to a full rollout. [ev-015]
  • The real-world incident rate for the open prompt-injection concern at full rollout scale is unknown (ev-026), which could impact security posture. [ev-026]
Blocking unknowns:
  • Long-run behaviour of agent-suggested code quality as the underlying model provider updates its model version is not established (ev-025)
  • Real-world incident rate for the open prompt-injection concern at full rollout scale is unknown (ev-026)

cro_compliance

DEFER MEDIUM
  • The proposal lacks a documented human-review process for AI-suggested changes specifically for regulated systems, which is critical given the non-mandatory but significant signal of new regulatory guidance (ev-012). [ev-012]
  • Current security review findings do not generalize to a full rollout because the pilot did not include systems handling regulated customer data (ev-015). [ev-015]
  • A known prompt-injection vulnerability remains open and unaddressed at scale, posing an enterprise risk. [ev-008, ev-026]
Blocking unknowns:
  • Whether the current general code-review process is sufficient to satisfy the 'documented human review' signal in ev-012.
  • The specific impact of prompt injection at scale (ev-026).

business_executive

DEFER MEDIUM
  • Concern about change-management and training cost realism at 35x the pilot's scale. [ev-020]
  • Uncertainty around attrition-risk reduction benefit being treated as a measured saving when it is an explicitly low-strength assumption. [ev-019]
  • Lack of clear understanding on how to attribute observed cycle-time or defect-rate changes at enterprise scale specifically to agent usage versus concurrent process changes. [ev-024]
Blocking unknowns:
  • How to accurately measure the impact of agent-suggested code quality updates on long-run behaviour
  • The real-world incident rate for the open prompt-injection concern at full rollout scale
NPV lowNPV midNPV highPaybackPeak funding
-£5,351,240 £5,085,455 £38,138,347 0.0y £0
Tornado (assumption swing on NPV, most sensitive first)
attribution_factor£16,922,975
uplift£15,669,421
fully_loaded_cost_gbp£3,760,661
training_cost_per_engineer_gbp-£507,769
annual_support_cost_gbp-£413,223
Staged-funding ladder:
  1. Discovery: £50,000
  2. Pilot: £420,000
  3. Targeted scale: £1,600,000

4. What changed minds

cfo

evidence_driven retrofit

DEFER → DEFER [ev-015, ev-026]

The staged rollout proposal presents a measured approach to expanding an enterprise coding agent to 5,000 engineers based on promising pilot results. However, unresolved security concerns and the potential for significant technical challenges in scaling up to legacy codebases remain critical uncertainties that could undermine the viability of this investment.

cto

no_change

held position: whether the agent's performance generalizes to legacy/untyped codebases at scale

ciso

evidence_driven

REQUEST_EVIDENCE → REQUEST_EVIDENCE [ev-008, ev-012, ev-015]

The unresolved security concerns, particularly the open prompt-injection vulnerability and lack of documented human review processes for AI-suggested changes in regulated systems, pose a significant risk to the organization's security posture. These issues need to be adequately addressed before proceeding with the proposed rollout.

business_executive

evidence_driven

DEFER → DEFER [ev-015, ev-026]

The staged rollout proposal fails to adequately address unresolved security concerns and uncertainties around full-scale adoption rates, which pose unacceptable risks to the organization's security posture.

5. Decision record

SYNTHETIC RECOMMENDATION · NOT A DECISION
DEFER £4,200,000 over 24 months MEDIUM
Conditions:
  • whether the security review covers regulated-data systems
  • the real-world incident rate for open prompt-injection vulnerabilities
  • whether the agent's performance generalizes to legacy/untyped codebases at scale
  • real-world incident rate for prompt-injection vulnerabilities in production
  • generalizability of security review findings to regulated-data systems
  • long-term stability of model quality as the underlying provider updates its version
  • attribution of productivity gains to the agent versus concurrent process changes
  • whether the agent performs well on untyped codebases
  • whether the pilot's success metrics generalize to the full estate
  • whether the adoption forecast is accurate
  • whether the productivity uplift assumptions are validated at scale
  • the impact of prompt-injection vulnerabilities on enterprise scale
  • the long-term stability of agent model quality
STRONGEST DISSENT

ciso: REQUEST_EVIDENCE — The open prompt-injection concern (ev-008) has not been resolved, leaving unaddressed a potential security risk in the full rollout.; The pilot did not exercise the agent against any system handling regulated customer data (ev-015), so it's unclear if the security review's findings generalise to a full rollout.; The real-world incident rate for the open prompt-injection concern at full rollout scale is unknown (ev-026), which could impact security posture. [ev-008, ev-015, ev-026]

Unresolved unknowns:
  • whether the security review covers regulated-data systems
  • the real-world incident rate for open prompt-injection vulnerabilities
  • whether the agent's performance generalizes to legacy/untyped codebases at scale
  • real-world incident rate for prompt-injection vulnerabilities in production
  • generalizability of security review findings to regulated-data systems
  • long-term stability of model quality as the underlying provider updates its version
  • attribution of productivity gains to the agent versus concurrent process changes
  • whether the agent performs well on untyped codebases
  • whether the pilot's success metrics generalize to the full estate
  • whether the adoption forecast is accurate
  • whether the productivity uplift assumptions are validated at scale
  • the impact of prompt-injection vulnerabilities on enterprise scale
  • the long-term stability of agent model quality
HUMAN DECISION

No human decision recorded yet for this run.