Agentic Architecture for Autonomous Issue Resolution — A Reference Framework
Fullstory
PROOF CASE  ·  MARRIOTT INTERNATIONAL
Reference Framework  ·  Agentic Architecture for Autonomous Issue Resolution

From Signal to Shipped Fix,
Without Guessing

A resilient agentic architecture for autonomous issue detection, root cause analysis, and remediation — built so that every conclusion the system reaches is traceable to observed guest behavior, and every line of code it proposes passes through a human.

✓ Orchestrator + 4 scoped specialists 2 mandatory human gates per loop Baseline to beat: 48h to containment
Purpose
A reference architecture and methodology for enterprise agentic build-outs: autonomous issue detection, root cause analysis, and remediation
Audience
Enterprise customers and systems integrator partners building practices around this pattern
Proof case
Marriott International — Dynatrace and Qualtrics remediation loops
Engagement shape
A 60-day Phase 4a pilot, two named prerequisites, and an Agent Review Board

Build a routing orchestrator over narrow specialists. Keep humans on the merge button.

Behavioral evidence proves that a failure happened and to whom. It does not localize the defect to a line of code. The consequence is that server-side correlation is essential: Fullstory proves impact and sizes the money, Dynatrace and Splunk localize the defect, and neither source alone authorizes a code change.

The topology that follows is a central orchestrator routing to four tool-scoped specialist agents, strict hub-and-spoke. This document supplies the methodology, the reasoning behind each control, and the infrastructure build-out — a framework that systems integrator partners can deploy, with Marriott as the proof case.

The engagement shape is a 60-day Phase 4a pilot with two named prerequisites and an Agent Review Board: a human in every step, no code generation, and a target time-to-evidence well under the 48-hour baseline.

Who this is for, and what the proof case proves

This is a reference architecture and methodology for enterprise agentic build-outs, not a proposal awaiting a decision. It is written for two audiences: enterprise customers planning an autonomous issue-resolution capability, and systems integrator partners building practices around deploying one on a customer's behalf. The framework is designed so that those partners can take it to market and build engagements around it.

Marriott is the proof case, and the claim is a specific one. The Qualtrics-to-revenue loop was demonstrated in a prototype using non-Marriott data: a negative verbatim became a confirmed behavioral finding, a 2.73-point conversion delta, and a $200M annualized exposure — the ceiling the prototype showed is reachable. The 48-hour human containment baseline is the target to beat. The design pattern itself has a track record across multiple APM platforms in partnership with Fullstory. What the Marriott architecture adds is what makes that pattern production-grade: the methodology, the controls, and the infrastructure.
One constraint shapes everything below

Before the architecture: a limit worth stating plainly, because designing around it is the difference between a system leadership trusts and a system that gets switched off after its first confident wrong answer.

Behavioral evidence proves that a failure happened and to whom. It does not localize the defect to a line of code. A session replay can show a guest clicking Book Now three times against a "Correction needed" message, the DOM state at that instant, the failing network call, and the console error. That is decisive proof of impact. It is not proof of cause. Any agent that generates a pull request from session evidence alone will confidently invent a root cause.

The architectural consequence: server-side correlation is essential, not supplementary. The Dynatrace trace and the Splunk log are what localize the defect; Fullstory is what proves the defect mattered and sizes the money. Neither source alone authorizes a code change. This is why the recommended topology separates observation from inference, why the orchestrator requires two independent sources to agree before a case advances, and why autonomous PR generation begins on front-end presentation defects — the stay-date display class of bug, where the client-side evidence is the localization — and expands outward only as measured precision earns it.

Agentic topology

Recommendation: a central Triage Orchestrator routing to four tool-scoped specialist agents, strict hub-and-spoke, no peer-to-peer calls. Not a monolith, and not a swarm.

Control plane

Triage Orchestrator

Owns the case file and the state machine. Decides which specialist runs next, enforces the two-source reconciliation rule, enforces step and cost budgets, and posts to the human gates. Deliberately has no analytical tools and no write access to any system of record except the case store and the notification channel — it cannot investigate, and it cannot ship.

case_store dispatch_agent notify (Teams / Slack) jira_create
Cannot: call Fullstory, Dynatrace, Splunk, or GitHub directly. Cannot assert a root cause. Cannot advance a case whose claims lack provenance tokens.
Specialist  ·  observe only

Session Evidence Agent

The sole holder of the Fullstory session lease. Opens the exact session and timestamp, reads DOM and accessibility state, captures the visible failure, diffs the interaction, closes the lease. Emits observations with timestamps — never causes.

session_open session_view session_diff session_screenshot session_get_a11y_tree session_close get_session_events
Cannot: use causal language. Cannot see the code repository. Cannot compute cohort metrics. System prompt rejects any output containing "because", "caused by", or a proposed fix.
Specialist  ·  quantify

Cohort Quantification Agent

Sizes the cohort and the money. Confirms what a funnel or segment actually measures before computing it, then applies the customer's already-ratified formula: conversion delta between sessions with and without the friction, annualized, times AOV. Returns sample size with every figure.

build_segment get_segment build_funnel get_funnel compute_funnel compute_metric get_opportunities snowflake_query (parameterized)
Cannot: write free-form SQL. Cannot publish a dollar figure below the 100-session floor. Cannot reference an object ID that does not exist in the org.
Specialist  ·  localize

Ops Correlation Agent

Pulls the failing trace and the surrounding logs for the same correlation ID and time window, and produces the server-side failure signature that actually localizes the defect. This is the agent whose output authorizes engineering work.

dynatrace_problem_detail dynatrace_purepath splunk_search (saved)
Cannot: read Fullstory. Cannot write to any system. Cannot run unbounded Splunk queries — saved searches with bound time windows only.
Specialist  ·  propose code

Remediation Engineering Agent

Runs only after a human approves the evidence package. Works in an ephemeral sandboxed worktree with no production network egress. Must produce a regression test that fails before its change and passes after — a branch without one is rejected automatically. Opens a draft pull request and stops.

repo_read branch_create commit pr_open_draft test_run
Cannot: merge, force-push, or touch protected branches. Cannot modify CI config, infrastructure-as-code, secrets, database migrations, or auth and payment modules. Cannot see raw guest survey text. Cannot exceed the diff cap.

Why not one monolithic agent

Four of these reasons are standard enterprise architecture. The first is specific to this toolchain and is, on its own, decisive.

Why not a peer-to-peer multi-agent swarm

Closed-loop workflows

Two trigger loops and one always-on monitor. Gold rows are mandatory human checkpoints. Red rows are stop conditions where the system is designed to decline rather than proceed.

Loop A  ·  Infrastructure / Ops (Dynatrace-triggered)

Target: ops signal to approved evidence package in minutes; signal to merged fix within the same business day. The X-FullStory-URL request attribute on the failing PurePath is the bridge that makes the whole loop possible — and its absence is the loop's first stop condition.

StepOwnerAction & data inputTool callsHand-off
A0 Intake service
deterministic
Davis AI problem event arrives by webhook. Verify HMAC signature, extract problem ID, affected service, error signature, impacted request count, time window, and the X-FullStory-URL attribute. Hash the error signature and dedupe against open cases. webhook_verifysignature_hashdedupe New case, or increment an existing one
A0•STOP Intake service No session link, no agentic RCA. If the failing trace carries no Fullstory URL attribute, the case is routed to human triage with the ops context only. No model is invoked. At the proof case's 4.7M visits a day, a single popular failure can emit thousands of events — dedupe by signature means one case per signature per window, not one case per event. — Human triage queue
A1 Orchestrator Open an append-only case file. Assign case ID. Thread the correlation ID: Dynatrace problem → case → Jira key → PR → release. Set step, token, and wall-clock budgets. case_store.create Dispatch A2 and A3 in parallel
A2 Ops Correlation Retrieve the failing PurePath and the surrounding logs for the same correlation ID and window. Produce the server-side failure signature: service, method, status, exception class, first-seen, rate. dynatrace_problem_detaildynatrace_purepathsplunk_search Technical finding → case file
A3 Session Evidence Parse device and session IDs from the bridge URL. Open the session at the event timestamp. Capture the accessibility tree for DOM state, a screenshot of the visible failure, and a diff across the failing interaction. Close the lease. session_opensession_viewsession_get_a11y_treesession_screenshotsession_diffsession_close Observation record → case file
A4•STOP Orchestrator Reconciliation gate. Does the client-side observation corroborate the server-side signature inside the same time window? If the two do not agree, the case is marked unreconciled and escalated to a human. It does not proceed to quantification and it certainly does not proceed to code. reconcile Escalate, or advance to A5
A5 Cohort Quantification Build the segment of sessions exhibiting the signature. Read the checkout funnel definition, confirm what it measures, then compute with and without the signature. Apply the ratified formula. Return affected users, conversion delta, annualized exposure, and sample size. build_segmentget_funnelcompute_funnelcompute_metricsnowflake_query Quantified impact → case file
A6•HITL 1 Named engineering owner Evidence review. The complete package — technical signature, behavioral observations with session links and timestamps, screenshot, dollar exposure with sample size, and the methodology used — posts to Teams or Slack and files a Jira issue. Approve, reject, or request more evidence. Nothing reaches the repository before this. notifyjira_create Approval unlocks A7
A7 Remediation Engineering Clone to an ephemeral sandboxed worktree. Locate the defect from the server-side signature. Write the minimal fix plus a regression test that reproduces the failing condition. Run the existing suite. Open a draft PR whose body carries the full evidence package and the agent-generated label. repo_readbranch_createtest_runpr_open_draft Draft PR → CODEOWNERS
A8•HITL 2 CODEOWNERS reviewer Code review and merge. Normal review process, normal release train. The agent has no merge permission at the GitHub App level — this gate is enforced by platform configuration, not by instruction. — Merge → release
A9 Cohort Quantification Post-release verification. Re-compute the identical funnel and segment on the post-deploy window and re-check the Dynatrace problem state. Confirm resolution, or reopen the case automatically. compute_funnelcompute_metricdynatrace_problem_detail Close, or reopen at A1

Loop B  ·  Voice of Customer (Qualtrics-triggered)

This loop differs from Loop A in one structural way, and the difference matters more than anything else in it: a single negative survey is an anecdote. Loop B will not run root cause analysis on a sample of one.

StepOwnerAction & data inputTool callsHand-off
B0 Intake service
deterministic
Qualtrics response webhook on a CSAT or NPS score below threshold, or a negative verbatim. Extract the embedded Fullstory session link, survey metadata, brand, property, loyalty tier, and device. Treat the verbatim as untrusted data from this moment forward. webhook_verifyextract_session_link Theme classifier
B1 Classifier
small model
Map the verbatim to a fixed theme taxonomy — date handling, points redemption, sign-in, rate display, reservation change, payment. Closed vocabulary, not free-form generation. Output is a label and a confidence, nothing else. classify_theme Theme + confidence → cluster store
B2•STOP Cluster service Cluster threshold. A case opens only when at least K responses share a theme on the same page or flow inside a rolling window, or a single response correlates to a live Dynatrace problem. Everything else accumulates as aggregate VoC signal. Responses with no session link follow the aggregate-only path and never trigger RCA. cluster_eval Aggregate dashboard, or open a case
B3 Orchestrator Open the case. Select up to five representative sessions across the cluster — not one — spanning device, brand, and loyalty tier. Dispatch three independent investigations in parallel. case_store.createget_sessions Dispatch B4, B5, B6
B4 Classifier Subjective. Structured summary of what guests say is wrong, across the cluster: theme, frequency, affected tiers, verbatim exemplars quoted and attributed, never paraphrased into a cause. qualtrics_responses Subjective finding
B5 Ops Correlation Technical. Were there errors, latency excursions, or failing traces on the implicated flow during the cluster window? A negative result here is a finding, not a failure — it points the case toward a UX or content defect rather than an engineering one. dynatrace_problem_detailsplunk_search Technical finding
B6 Session Evidence Behavioral. Walk each representative session in turn — lease held by one agent, one session at a time. Capture what the guest actually saw and did: DOM state, visible messaging, dead and rage interactions, the point of abandonment. session_opensession_viewsession_get_a11y_treesession_screenshotsession_close Observation records
B7•STOP Orchestrator Two-of-three triangulation. At least two of the three independent findings must agree on the same defect before the case is declared validated. One source alone — however vivid the verbatim — does not advance the case. reconcile Escalate, or advance to B8
B8 Cohort Quantification Size the exposure beyond the survey respondents — the respondents are a sample, and a biased one. Build the behavioral segment matching the validated defect across all traffic, compute the conversion delta, annualize against AOV. build_segmentcompute_funnelcompute_metricsnowflake_query Quantified impact
B9•HITL 1 CX + Product owner Validated issue review. Jira issue with the triangulated evidence, the cohort exposure, and the recommended disposition. This is the queue the proof case's VoC initiative already calls for: a visible, quantified pipeline of guest pain that product squads can prioritize. jira_createnotify Approval unlocks B10
B10 Remediation Engineering Fix plus regression test, deployed to a lower environment only. The agent receives the structured finding — never the raw guest text. Production deployment is not in this agent's capability set at any phase. branch_createtest_rundeploy_lowerpr_open_draft Lower-env build
B11 Session Evidence + Quantification Agentic validation in the lower environment. Synthetic traffic drives the repaired flow; the agent inspects the resulting sessions and confirms the defect no longer reproduces. This step requires Fullstory instrumentation in lower environments — a named prerequisite, not an assumption. get_sessionssession_opensession_viewsession_closecompute_funnel Validation verdict
B12•HITL 2 Release management Production release approval. Normal release train, normal change control. The validation verdict is an input to the human decision, not a substitute for it. — Production release
B13 Orchestrator Close on theme decay, not on merge. Monitoring continues; the case closes when the Qualtrics theme rate returns to baseline and the behavioral segment shrinks. A merged PR is not a resolved guest experience. cluster_evalcompute_metric Close, or reopen at B3

Loop C  ·  Always-on funnel vigilance

Loops A and B are both reactive — they wait for infrastructure to notice or a guest to complain. The highest-value failures are the silent ones that do neither.

A scheduled job recomputes the checkout funnels hourly against statistical control limits rather than fixed thresholds, so normal daily and weekly seasonality does not generate noise. A step-level drop outside the control band opens a case at step A1 of Loop A, bypassing the Dynatrace trigger entirely. This is the pattern that catches the failure class that never throws a server error and never produces a survey response: the rate card that renders empty, the date that silently reverts, the points balance that displays as unavailable. Early cases from this loop should run evidence-only, because without a server-side signature there is nothing to localize — which means Loop C feeds human investigation first, and earns code-generation rights last.

Resiliency & hallucination defense

Ten controls. The first three are the ones that matter most, and none of them is a prompt instruction — they are enforced by deterministic code between the agents.

The failure mode

One model gathers and explains

Root-cause hallucination has a specific origin. A single context window takes in partial evidence, feels pressure to produce an answer, and generates a plausible causal story that reads exactly like a real finding. It cites the session it looked at. It sounds confident. It is wrong in a way no reviewer can detect without redoing the work.

  • Confident causes from a single source
  • Dollar figures with no sample size
  • Object IDs that do not exist in the org
  • Behavioral explanations for tracking gaps
The control

Observation and inference are different agents

The Session Evidence Agent is permitted to say what it saw and when. It is architecturally prevented from saying why. Inference happens only in the orchestrator's reconciliation step, which cannot run on one source. The model never gets to hold an uncited claim across a hand-off.

  • Every claim carries a provenance token
  • Two independent sources or the case stops
  • Named-object allowlist, no invented IDs
  • Abstention is a first-class outcome

1  ·  The citation contract

Every assertion written into a case file must carry a provenance token: a session ID plus timestamp, a Dynatrace problem ID, a saved-search hash, or a metric ID plus the computed value. A deterministic validator sits between every agent hand-off and strips any claim without one. The stripped claim is logged — a rising strip rate is an early warning that an agent is drifting, and it is one of the metrics the review board watches.

// case-file assertion — rejected without provenance { "claim": "Guests saw 'Correction needed' when committing a 14-night reservation", "provenance": [ { "type": "session_observation", "session": "<id>", "ts": "20:52:14.330Z", "artifact": "a11y_tree + screenshot" }, { "type": "dynatrace_problem", "id": "<problem>", "signature": "POST /reservation/commit 500" } ], "inference_allowed": true // two independent sources present }

2  ·  How Fullstory MCP grounds the reasoning

"Use MCP as ground truth" is a slogan until it is four specific mechanisms. These are the four.

3  ·  Instrumentation fragility is a first-class outcome

The proof case's configuration audit flags two gaps that bound agent reliability: inconsistent data attributes and CSS selector standards across the site, and iOS and Android SDKs not yet at full parity with web. Element-level agent reasoning inherits that fragility directly. An agent that cannot tell "the guest did not click it" from "we stopped capturing it" will report a behavior change that is actually a selector change.

The control is a pre-flight instrumentation health check on every case: are the elements and pages this case depends on being captured stably across the window under examination? If not, the agent returns a tracking cause, not a behavioral one, and routes to the capture-governance queue. This must be a rewarded outcome in the agent's evaluation, not a failed run. A system that can only ever find product defects will find product defects in instrumentation gaps.

4  ·  Confidence floors and abstention

5  ·  Destructive-action prevention

Capability-scoped at the platform layer. A prompt that says "do not merge" is a suggestion; a GitHub App without merge permission is a guarantee.

ControlEnforcement
No merge, everDedicated GitHub App with contents:write on unprotected branches and pull_requests:write only. No merge scope, no admin scope. Branch protection and CODEOWNERS enforced server-side.
Path denylistCI configuration, infrastructure-as-code, secrets, database migrations, and authentication and payment modules are blocked at the pre-commit hook. Touching them requires an explicit human-set flag on the case.
Diff capChanges above roughly 200 lines or 5 files are refused. Above the cap the agent files an analysis issue with its findings instead of a branch — a large diff is a design decision, not a hotfix.
Sandboxed executionEphemeral worktree, no production network egress, no production credentials in the environment, destroyed at case close.
Provenance on every PRagent-generated label plus the full evidence package in the body. No silent authorship anywhere in the repository.
IdempotencyEvery external write carries an idempotency key derived from the case ID and step. A retry cannot double-file a Jira issue or double-open a PR.
Circuit breakerRepeated failures on the same error signature trip the breaker and route to humans. Per-case step, token, and wall-clock budgets terminate runaway loops.
Session lease safetySingle-holder lease with TTL and a reaper process. Close precedes open, always. Only the Session Evidence Agent holds session tools at all.

6  ·  Prompt injection through guest-supplied text

Loop B ingests free text written by members of the public. A guest can type instructions into a survey comment box. Session DOM content can contain attacker-controlled strings. Any of it can reach a model context that also holds repository write capability.

7  ·  Human-in-the-loop checkpoints, stated plainly

GateOwnerWhat cannot happen before it
Evidence review (A6 / B9)Named engineering owner; CX and Product for Loop BNo repository access of any kind. No branch, no commit, no PR. The agent's investigation is complete and the human decides whether it is right.
Code review and merge (A8 / B12)CODEOWNERS; Release managementNo merge to any protected branch. No production deployment. Enforced by platform permissions, not policy.
Scope expansionAgent Review BoardNo new repository, path, or capability enters an agent's allowlist without a board decision backed by measured precision on the current scope.
Methodology changeExperimentation & Analytics ownerNo change to the quantification formula or the cohort definitions behind a published figure. Dollar figures must stay comparable quarter over quarter.
Production deployment is never an agent capability. Not in Phase 4a, and not at full maturity. The ceiling on autonomy here is a validated draft PR and a lower-environment deployment. Every production change goes through the customer's existing release train, reviewed and merged by a person.

8  ·  Observability of the agents themselves

9  ·  The evaluation harness the proof case already owns

At the proof case, twelve quantified frictions were identified in one year and eight have been found and fixed, each with a human-ratified root cause and a human-ratified dollar figure. That is a golden regression set. Before any agent is trusted with a live case, it replays those twelve incidents from their original signals and is scored on two things: did it reach the same root cause, and did it land inside a tolerance band of the ratified number. This converts "do we trust the agent" from a matter of opinion into a measurement, and it is the single highest-leverage thing to build in the first sixty days.

10  ·  Graceful degradation

Hosting & enterprise infrastructure

Agents run inside the customer's cloud perimeter. Models are consumed as managed in-region services. Fullstory MCP is reached outbound-only. No inbound internet path to the agent runtime.

Reference architecture

LayerRecommendationRationale
Ingress & control
Trigger ingressAPI gateway with mTLS and HMAC signature verification, fronting a durable queueDynatrace and Qualtrics webhooks are the only inbound path. The queue absorbs alert bursts, survives downstream outages, and makes replay possible for incident forensics.
Agent runtimeContainerized on the customer's existing Kubernetes platform, private VPC, no inbound internet, scale-to-zero workersThis workload is bursty, not steady. Event-driven workers with reserved capacity only for the always-on funnel job matches spend to incidents.
Egress controlExplicit allowlist: Fullstory API, Dynatrace, Qualtrics, GitHub Enterprise, the model endpoint. Everything else denied.The primary technical mitigation for prompt injection. An injected instruction has no reachable destination.
Reasoning
Primary model layerManaged in-region multimodal endpoint with customer-managed encryption keys, no-training and no-retention terms, and VPC service controlsMultimodal reasoning over session screenshots and DOM state is a hard requirement, and code generation quality is the gating factor on Loop A's value. This is where the capability actually lives.
Model abstractionProvider-agnostic routing layer; per-agent model selection; versions pinned per agent and promoted deliberatelyThe right reasoning model will change inside twelve months. Pinning per agent means a model upgrade is a measurable change against the golden regression set, not a surprise in production.
Self-hosted tierSmall open-weight classifier for theme classification and signature dedupe onlyHigh volume, low judgment, no data egress, trivially cheap. This is where self-hosting genuinely pays. We do not recommend self-hosting the reasoning or code-generation tier: open-weight quality at realistic internal scale is not there, and GPU capacity planning for bursty incident traffic is poor use of capital.
Data & evidence
Fullstory MCP accessOutbound-only, scoped service credential per agent role, read-only for every agent except the authoring scope the quantification agent needsThe Session Evidence Agent's credential cannot compute metrics. The quantification agent's credential cannot open sessions. Credential scope mirrors the topology.
Snowflake — groundingParameterized, version-controlled queries over the clickstream already landing from Phase 3. The agent calls a reviewed query; it does not write SQL.This is the strongest grounding move available. The dollar figure comes from a reviewed, versioned data model that analytics owns — not from a language model's arithmetic. It also makes every published figure reproducible by a human running the same query.
Snowflake — evaluationCase files, agent decisions, and outcomes persisted as the longitudinal evaluation datasetPrecision trends, dollar-estimate error bands, and PR acceptance rates are queries against this table. Autonomy decisions are made from it.
SplunkAudit sink for all agent activity, and the read source for the Ops Correlation Agent via saved searches with bound windowsOne correlation ID chain end to end is what makes the loop auditable to risk and audit functions. Saved searches prevent unbounded query cost.
DynatraceTrigger source via Davis AI, and APM for the agent services themselvesAlready the detection layer. Extending it over the agent runtime means agent latency and error rates are visible in the same place as everything else.
Security & privacy
Identity & secretsPer-agent workload identity, short-lived tokens from the existing enterprise vault, separate GitHub App per agent role. No long-lived personal access tokens anywhere.Separation of duties is only real if the credentials are actually separate. One compromised agent identity reaches one system with one scope.
Guest data in contextTokenized identifiers only. Loyalty tier carried as a cohort label, never as identity. Screenshots retained in customer-controlled storage on a short TTL.Fullstory's private-by-default capture and element-level exclusion mean payment and personal fields are excluded at the source, so the agent context does not contain them to begin with. Tokenization closes the remainder.
Prompt & response logsClassified sensitive, retained in Splunk under existing policy, access-controlled to the review board and securityThese logs are the audit record. They are also, unavoidably, the place where guest context would surface if a control failed — so they are governed accordingly.
Lower environmentsFullstory instrumentation in staging on a separate org, driven by synthetic trafficNamed prerequisite. Loop B's agentic validation step cannot exist without it. Not currently in place at the proof case.
Two prerequisites carry the whole program, and both are concrete and assignable. First, the X-FullStory-URL request attribute must be present on every service in the booking and authentication paths — where it is missing, Loop A is blind, and the system will correctly refuse to run. Second, lower-environment Fullstory instrumentation, without which Loop B's closed validation step is aspiration rather than architecture.
Phasing, governance & engagement shape

Autonomy expands as a function of measured precision, never by calendar date. Each phase must clear its metrics on the golden regression set before the next unlocks.

4a
0–60d

Human in every step. No code generation at all.

Complete the bridge attribute across booking-path services. Stand up the case file, the correlation ID chain, and the Splunk audit sink. Grant governed read-only MCP access. Build the golden regression set from the twelve quantified frictions and score a human-supervised agent against it. Loop A runs to the evidence package and stops there.

Unlock criteria: root-cause precision of at least 80% on the golden regression set (12 incidents), dollar estimates within ±15% of the ratified figure, and median time-to-evidence of 24 hours or less — materially under the 48-hour human containment baseline. These are starting values for board ratification, not production commitments.
4b
60–120d

Agentic investigation, human decision.

Both loops run end to end through evidence and quantification. Jira issues auto-file with full provenance. Loop C begins in observation mode. Still no pull requests. This phase is where a VoC pipeline of the kind the proof case has already scoped becomes continuous rather than episodic.

Unlock criteria: stable precision across both loops, abstention rate inside the healthy band, and zero unreconciled cases advanced in error.
4c
120–180d

Supervised remediation on an allowlisted scope.

Draft pull requests enabled on a small set of low-risk repositories and paths — front-end presentation defects first, the stay-date display class of bug, where client-side evidence is also the localization. Two human gates on every case. Every merge performed by a person.

Unlock criteria: PR acceptance rate above the agreed threshold, zero escaped defects attributable to an agent-authored change, and no denylist violations.
4d
180d+

Closed-loop validation, scope widened by evidence.

Lower-environment deployment and agentic validation go live for Loop B. The path allowlist widens one scope at a time, each expansion a board decision backed by measured precision on the current scope. Production deployment remains outside agent capability permanently.

Steady state: recovered revenue attributed to agent-originated cases becomes a reported line in the quarterly value readout.

The metrics that govern the program

MetricWhy it governs
Time to containmentThe headline. Current human baseline is 48 hours from identification to deployed temporary fix. This is the number the program is accountable to.
Root-cause precisionAgent-proposed cause versus human-ratified cause on the golden set and, later, on live cases. The primary trust metric.
Dollar-estimate error bandAgent figure versus the eventually ratified figure. Credibility with leadership depends more on this being tight and stable than on it being large.
PR acceptance rateAgent-authored draft PRs merged without substantive rework. The gate on expanding the path allowlist.
Abstention rateCases the system declined to conclude. Watched for being too low, which indicates hallucination, as much as for being too high.
Provenance strip rateClaims rejected by the citation validator. A rising rate is the earliest available signal of agent drift.
Tracking-cause shareCases resolved as instrumentation gaps rather than product defects. Both a health signal for capture governance and proof the agents are not forcing behavioral explanations.
Escaped defect rateProduction issues traceable to an agent-authored change. Target is zero, and a single occurrence pauses the phase.

Governance

We recommend a standing Agent Review Board with representation from technology, digital innovation, experimentation & analytics, capture and instrumentation governance, and enterprise security and architecture. It owns four decisions and only four: scope expansion, methodology change, phase advancement, and incident review on any agent-attributable defect. It meets on the existing cadence rather than creating a new one, and it is the natural first charter for a Center of Excellence.

What a Phase 4a engagement looks like

Nothing in the first sixty days generates code or touches a repository. The engagement stands up the evidence chain and proves precision against incidents the customer has already solved by hand.

Part 1

A 60-day Phase 4a pilot on Loop A, evidence-only

Dynatrace trigger through to a human-reviewed evidence package. No code generation in scope. Scored against the twelve already-quantified frictions.

Part 2

Two prerequisites, each with a named owner

Bridge attribute coverage across booking and authentication path services, and Fullstory instrumentation in lower environments. Both are engineering tasks with clear completion criteria, and both gate program value.

Part 3

Hosting and privacy posture ratified with security and architecture

In-region managed model endpoint with customer-managed keys and no-retention terms, private VPC runtime, allowlisted egress, tokenized identifiers, Splunk audit retention. Confirm the pattern before the build rather than during review.

Part 4

An Agent Review Board, with unlock thresholds agreed in advance

Specifically: the precision figure and the dollar tolerance band that unlock Phase 4b, agreed in advance. Setting the bar before the first result is what keeps the program honest.

Open risks, stated plainly