This is a reference architecture and methodology for enterprise agentic build-outs, not a proposal awaiting a decision. It is written for two audiences: enterprise customers planning an autonomous issue-resolution capability, and systems integrator partners building practices around deploying one on a customer's behalf. The framework is designed so that those partners can take it to market and build engagements around it.
Before the architecture: a limit worth stating plainly, because designing around it is the difference between a system leadership trusts and a system that gets switched off after its first confident wrong answer.
The architectural consequence: server-side correlation is essential, not supplementary. The Dynatrace trace and the Splunk log are what localize the defect; Fullstory is what proves the defect mattered and sizes the money. Neither source alone authorizes a code change. This is why the recommended topology separates observation from inference, why the orchestrator requires two independent sources to agree before a case advances, and why autonomous PR generation begins on front-end presentation defects — the stay-date display class of bug, where the client-side evidence is the localization — and expands outward only as measured precision earns it.
Recommendation: a central Triage Orchestrator routing to four tool-scoped specialist agents, strict hub-and-spoke, no peer-to-peer calls. Not a monolith, and not a swarm.
Triage Orchestrator
Owns the case file and the state machine. Decides which specialist runs next, enforces the two-source reconciliation rule, enforces step and cost budgets, and posts to the human gates. Deliberately has no analytical tools and no write access to any system of record except the case store and the notification channel — it cannot investigate, and it cannot ship.
Session Evidence Agent
The sole holder of the Fullstory session lease. Opens the exact session and timestamp, reads DOM and accessibility state, captures the visible failure, diffs the interaction, closes the lease. Emits observations with timestamps — never causes.
Cohort Quantification Agent
Sizes the cohort and the money. Confirms what a funnel or segment actually measures before computing it, then applies the customer's already-ratified formula: conversion delta between sessions with and without the friction, annualized, times AOV. Returns sample size with every figure.
Ops Correlation Agent
Pulls the failing trace and the surrounding logs for the same correlation ID and time window, and produces the server-side failure signature that actually localizes the defect. This is the agent whose output authorizes engineering work.
Remediation Engineering Agent
Runs only after a human approves the evidence package. Works in an ephemeral sandboxed worktree with no production network egress. Must produce a regression test that fails before its change and passes after — a branch without one is rejected automatically. Opens a draft pull request and stops.
Why not one monolithic agent
Four of these reasons are standard enterprise architecture. The first is specific to this toolchain and is, on its own, decisive.
- The session lease is a single-holder resource. Fullstory's visual inspection tools are stateful: a session is opened, inspected, and must be closed before another is opened in the same run. A monolithic agent that fans out across several sessions in parallel will leak or deadlock leases. Confining the lease to one single-threaded specialist with a TTL and a reaper makes close-safety an architectural property rather than a prompt instruction the model may forget under pressure.
- Tool-selection accuracy degrades with surface area. Fullstory MCP alone exposes 35 tools. Add Dynatrace, Splunk, Qualtrics, GitHub, Jira, Snowflake and Teams and a single agent is choosing among well over a hundred — in an incident context, under time pressure, where the wrong tool choice is a wrong answer. Each specialist here sees between three and eight.
- Separation of duties limits blast radius. The agent that can write to the repository holds no Fullstory, Dynatrace, or Splunk credential. The agent that gathers evidence cannot see code. No single compromised or drifting context can both fabricate evidence and act on it.
- Narrow contracts are testable. "Did the evidence agent correctly report the console error present in this session?" is a regression test. "Did the agent handle the incident well?" is not. Specialists can be evaluated and improved independently, which is what makes staged autonomy expansion possible.
- Model routing and cost. Multimodal session inspection is the expensive step; routing and classification are nearly free. Per-agent model selection lets the customer spend reasoning budget where judgment actually lives, and swap models without rewriting the system.
Why not a peer-to-peer multi-agent swarm
- No agent calls another agent. All control flows through the orchestrator. This keeps the audit trail linear — one ordered case file per incident — and makes emergent loops structurally impossible rather than merely discouraged.
- Intake and routing are deterministic code, not inference. Webhook parsing, signature verification, dedupe, and threshold checks run as ordinary services. The model is invoked only where judgment is genuinely required, which shrinks both the hallucination surface and the bill.
- Specialist count stays fixed. Capability grows by widening a specialist's allowlist, never by spawning new agents at runtime.
Two trigger loops and one always-on monitor. Gold rows are mandatory human checkpoints. Red rows are stop conditions where the system is designed to decline rather than proceed.
Loop A · Infrastructure / Ops (Dynatrace-triggered)
Target: ops signal to approved evidence package in minutes; signal to merged fix within the same business day. The X-FullStory-URL request attribute on the failing PurePath is the bridge that makes the whole loop possible — and its absence is the loop's first stop condition.
| Step | Owner | Action & data input | Tool calls | Hand-off |
|---|---|---|---|---|
| A0 | Intake service deterministic |
Davis AI problem event arrives by webhook. Verify HMAC signature, extract problem ID, affected service, error signature, impacted request count, time window, and the X-FullStory-URL attribute. Hash the error signature and dedupe against open cases. | webhook_verifysignature_hashdedupe | New case, or increment an existing one |
| A0•STOP | Intake service | No session link, no agentic RCA. If the failing trace carries no Fullstory URL attribute, the case is routed to human triage with the ops context only. No model is invoked. At the proof case's 4.7M visits a day, a single popular failure can emit thousands of events — dedupe by signature means one case per signature per window, not one case per event. | — | Human triage queue |
| A1 | Orchestrator | Open an append-only case file. Assign case ID. Thread the correlation ID: Dynatrace problem → case → Jira key → PR → release. Set step, token, and wall-clock budgets. | case_store.create | Dispatch A2 and A3 in parallel |
| A2 | Ops Correlation | Retrieve the failing PurePath and the surrounding logs for the same correlation ID and window. Produce the server-side failure signature: service, method, status, exception class, first-seen, rate. | dynatrace_problem_detaildynatrace_purepathsplunk_search | Technical finding → case file |
| A3 | Session Evidence | Parse device and session IDs from the bridge URL. Open the session at the event timestamp. Capture the accessibility tree for DOM state, a screenshot of the visible failure, and a diff across the failing interaction. Close the lease. | session_opensession_viewsession_get_a11y_treesession_screenshotsession_diffsession_close | Observation record → case file |
| A4•STOP | Orchestrator | Reconciliation gate. Does the client-side observation corroborate the server-side signature inside the same time window? If the two do not agree, the case is marked unreconciled and escalated to a human. It does not proceed to quantification and it certainly does not proceed to code. | reconcile | Escalate, or advance to A5 |
| A5 | Cohort Quantification | Build the segment of sessions exhibiting the signature. Read the checkout funnel definition, confirm what it measures, then compute with and without the signature. Apply the ratified formula. Return affected users, conversion delta, annualized exposure, and sample size. | build_segmentget_funnelcompute_funnelcompute_metricsnowflake_query | Quantified impact → case file |
| A6•HITL 1 | Named engineering owner | Evidence review. The complete package — technical signature, behavioral observations with session links and timestamps, screenshot, dollar exposure with sample size, and the methodology used — posts to Teams or Slack and files a Jira issue. Approve, reject, or request more evidence. Nothing reaches the repository before this. | notifyjira_create | Approval unlocks A7 |
| A7 | Remediation Engineering | Clone to an ephemeral sandboxed worktree. Locate the defect from the server-side signature. Write the minimal fix plus a regression test that reproduces the failing condition. Run the existing suite. Open a draft PR whose body carries the full evidence package and the agent-generated label. | repo_readbranch_createtest_runpr_open_draft | Draft PR → CODEOWNERS |
| A8•HITL 2 | CODEOWNERS reviewer | Code review and merge. Normal review process, normal release train. The agent has no merge permission at the GitHub App level — this gate is enforced by platform configuration, not by instruction. | — | Merge → release |
| A9 | Cohort Quantification | Post-release verification. Re-compute the identical funnel and segment on the post-deploy window and re-check the Dynatrace problem state. Confirm resolution, or reopen the case automatically. | compute_funnelcompute_metricdynatrace_problem_detail | Close, or reopen at A1 |
Loop B · Voice of Customer (Qualtrics-triggered)
This loop differs from Loop A in one structural way, and the difference matters more than anything else in it: a single negative survey is an anecdote. Loop B will not run root cause analysis on a sample of one.
| Step | Owner | Action & data input | Tool calls | Hand-off |
|---|---|---|---|---|
| B0 | Intake service deterministic |
Qualtrics response webhook on a CSAT or NPS score below threshold, or a negative verbatim. Extract the embedded Fullstory session link, survey metadata, brand, property, loyalty tier, and device. Treat the verbatim as untrusted data from this moment forward. | webhook_verifyextract_session_link | Theme classifier |
| B1 | Classifier small model |
Map the verbatim to a fixed theme taxonomy — date handling, points redemption, sign-in, rate display, reservation change, payment. Closed vocabulary, not free-form generation. Output is a label and a confidence, nothing else. | classify_theme | Theme + confidence → cluster store |
| B2•STOP | Cluster service | Cluster threshold. A case opens only when at least K responses share a theme on the same page or flow inside a rolling window, or a single response correlates to a live Dynatrace problem. Everything else accumulates as aggregate VoC signal. Responses with no session link follow the aggregate-only path and never trigger RCA. | cluster_eval | Aggregate dashboard, or open a case |
| B3 | Orchestrator | Open the case. Select up to five representative sessions across the cluster — not one — spanning device, brand, and loyalty tier. Dispatch three independent investigations in parallel. | case_store.createget_sessions | Dispatch B4, B5, B6 |
| B4 | Classifier | Subjective. Structured summary of what guests say is wrong, across the cluster: theme, frequency, affected tiers, verbatim exemplars quoted and attributed, never paraphrased into a cause. | qualtrics_responses | Subjective finding |
| B5 | Ops Correlation | Technical. Were there errors, latency excursions, or failing traces on the implicated flow during the cluster window? A negative result here is a finding, not a failure — it points the case toward a UX or content defect rather than an engineering one. | dynatrace_problem_detailsplunk_search | Technical finding |
| B6 | Session Evidence | Behavioral. Walk each representative session in turn — lease held by one agent, one session at a time. Capture what the guest actually saw and did: DOM state, visible messaging, dead and rage interactions, the point of abandonment. | session_opensession_viewsession_get_a11y_treesession_screenshotsession_close | Observation records |
| B7•STOP | Orchestrator | Two-of-three triangulation. At least two of the three independent findings must agree on the same defect before the case is declared validated. One source alone — however vivid the verbatim — does not advance the case. | reconcile | Escalate, or advance to B8 |
| B8 | Cohort Quantification | Size the exposure beyond the survey respondents — the respondents are a sample, and a biased one. Build the behavioral segment matching the validated defect across all traffic, compute the conversion delta, annualize against AOV. | build_segmentcompute_funnelcompute_metricsnowflake_query | Quantified impact |
| B9•HITL 1 | CX + Product owner | Validated issue review. Jira issue with the triangulated evidence, the cohort exposure, and the recommended disposition. This is the queue the proof case's VoC initiative already calls for: a visible, quantified pipeline of guest pain that product squads can prioritize. | jira_createnotify | Approval unlocks B10 |
| B10 | Remediation Engineering | Fix plus regression test, deployed to a lower environment only. The agent receives the structured finding — never the raw guest text. Production deployment is not in this agent's capability set at any phase. | branch_createtest_rundeploy_lowerpr_open_draft | Lower-env build |
| B11 | Session Evidence + Quantification | Agentic validation in the lower environment. Synthetic traffic drives the repaired flow; the agent inspects the resulting sessions and confirms the defect no longer reproduces. This step requires Fullstory instrumentation in lower environments — a named prerequisite, not an assumption. | get_sessionssession_opensession_viewsession_closecompute_funnel | Validation verdict |
| B12•HITL 2 | Release management | Production release approval. Normal release train, normal change control. The validation verdict is an input to the human decision, not a substitute for it. | — | Production release |
| B13 | Orchestrator | Close on theme decay, not on merge. Monitoring continues; the case closes when the Qualtrics theme rate returns to baseline and the behavioral segment shrinks. A merged PR is not a resolved guest experience. | cluster_evalcompute_metric | Close, or reopen at B3 |
Loop C · Always-on funnel vigilance
Loops A and B are both reactive — they wait for infrastructure to notice or a guest to complain. The highest-value failures are the silent ones that do neither.
A scheduled job recomputes the checkout funnels hourly against statistical control limits rather than fixed thresholds, so normal daily and weekly seasonality does not generate noise. A step-level drop outside the control band opens a case at step A1 of Loop A, bypassing the Dynatrace trigger entirely. This is the pattern that catches the failure class that never throws a server error and never produces a survey response: the rate card that renders empty, the date that silently reverts, the points balance that displays as unavailable. Early cases from this loop should run evidence-only, because without a server-side signature there is nothing to localize — which means Loop C feeds human investigation first, and earns code-generation rights last.
Ten controls. The first three are the ones that matter most, and none of them is a prompt instruction — they are enforced by deterministic code between the agents.
One model gathers and explains
Root-cause hallucination has a specific origin. A single context window takes in partial evidence, feels pressure to produce an answer, and generates a plausible causal story that reads exactly like a real finding. It cites the session it looked at. It sounds confident. It is wrong in a way no reviewer can detect without redoing the work.
- Confident causes from a single source
- Dollar figures with no sample size
- Object IDs that do not exist in the org
- Behavioral explanations for tracking gaps
Observation and inference are different agents
The Session Evidence Agent is permitted to say what it saw and when. It is architecturally prevented from saying why. Inference happens only in the orchestrator's reconciliation step, which cannot run on one source. The model never gets to hold an uncited claim across a hand-off.
- Every claim carries a provenance token
- Two independent sources or the case stops
- Named-object allowlist, no invented IDs
- Abstention is a first-class outcome
1 · The citation contract
Every assertion written into a case file must carry a provenance token: a session ID plus timestamp, a Dynatrace problem ID, a saved-search hash, or a metric ID plus the computed value. A deterministic validator sits between every agent hand-off and strips any claim without one. The stripped claim is logged — a rising strip rate is an early warning that an agent is drifting, and it is one of the metrics the review board watches.
2 · How Fullstory MCP grounds the reasoning
"Use MCP as ground truth" is a slogan until it is four specific mechanisms. These are the four.
- Confirm the definition before computing it. The quantification agent must read a funnel, segment, or metric definition and state in the case file what it is about to measure, then compute. This single discipline eliminates the most common analytical error: reporting "checkout drop" from an object that measures something adjacent to checkout.
- Named-object allowlist. Agents may reference only objects that exist in the org — the checkout conversion funnels, the confirmation and rate-list pages, the sign-in failure segment, the watched elements already active on critical front-end error messages. If the object required to answer the question does not exist, the agent must say so and stop. Approximating a missing object is the mechanism by which an agent invents confidence.
- Visual and DOM verification. The accessibility tree and screenshot give falsifiable client-side truth. A claimed error message must be present in the captured DOM or visible in the frame. This turns "the agent believes guests saw an error" into a check that either passes or fails.
- Methodology persistence against MCP non-determinism. The analytics MCP routes through a reasoning intermediary, so the same question asked twice can resolve to different query strategies. Left alone, two agent runs on one incident produce two different dollar figures — and nothing destroys executive trust faster. Each case therefore persists the object IDs, the original query strings, the time window, and the exclusions applied. Re-runs use the saved IDs with a drift check; if a saved ID returns empty, the agent reports drift rather than silently rebuilding.
3 · Instrumentation fragility is a first-class outcome
The control is a pre-flight instrumentation health check on every case: are the elements and pages this case depends on being captured stably across the window under examination? If not, the agent returns a tracking cause, not a behavioral one, and routes to the capture-governance queue. This must be a rewarded outcome in the agent's evaluation, not a failed run. A system that can only ever find product defects will find product defects in instrumentation gaps.
4 · Confidence floors and abstention
- Cohort floor. No dollar figure below 100 sessions. Below the floor, the finding ships as directional with the sample size stated and no monetary estimate attached.
- Corroboration floor. Two independent sources minimum, as above. For Loop B, two of three.
- Reproduction floor. No PR without a regression test that fails before the change and passes after. This is checked by CI, not asserted by the agent.
- Abstention is success. "Insufficient evidence, escalating" is a valid terminal state and is measured as such. An agent whose abstention rate is near zero is not accurate — it is hallucinating, and the metric will show it.
5 · Destructive-action prevention
Capability-scoped at the platform layer. A prompt that says "do not merge" is a suggestion; a GitHub App without merge permission is a guarantee.
| Control | Enforcement |
|---|---|
| No merge, ever | Dedicated GitHub App with contents:write on unprotected branches and pull_requests:write only. No merge scope, no admin scope. Branch protection and CODEOWNERS enforced server-side. |
| Path denylist | CI configuration, infrastructure-as-code, secrets, database migrations, and authentication and payment modules are blocked at the pre-commit hook. Touching them requires an explicit human-set flag on the case. |
| Diff cap | Changes above roughly 200 lines or 5 files are refused. Above the cap the agent files an analysis issue with its findings instead of a branch — a large diff is a design decision, not a hotfix. |
| Sandboxed execution | Ephemeral worktree, no production network egress, no production credentials in the environment, destroyed at case close. |
| Provenance on every PR | agent-generated label plus the full evidence package in the body. No silent authorship anywhere in the repository. |
| Idempotency | Every external write carries an idempotency key derived from the case ID and step. A retry cannot double-file a Jira issue or double-open a PR. |
| Circuit breaker | Repeated failures on the same error signature trip the breaker and route to humans. Per-case step, token, and wall-clock budgets terminate runaway loops. |
| Session lease safety | Single-holder lease with TTL and a reaper process. Close precedes open, always. Only the Session Evidence Agent holds session tools at all. |
6 · Prompt injection through guest-supplied text
- All survey verbatims and all captured DOM content are wrapped and labelled as untrusted data, never as instruction, at every hand-off.
- The classifier emits a label from a closed taxonomy — it does not generate free text downstream.
- The Remediation Engineering Agent never receives raw verbatim or raw DOM text. It receives only the structured, human-approved finding. The agent with write capability is the agent furthest from untrusted input.
- Outbound tool calls are allowlisted by host. An injected instruction has nowhere to send anything.
7 · Human-in-the-loop checkpoints, stated plainly
| Gate | Owner | What cannot happen before it |
|---|---|---|
| Evidence review (A6 / B9) | Named engineering owner; CX and Product for Loop B | No repository access of any kind. No branch, no commit, no PR. The agent's investigation is complete and the human decides whether it is right. |
| Code review and merge (A8 / B12) | CODEOWNERS; Release management | No merge to any protected branch. No production deployment. Enforced by platform permissions, not policy. |
| Scope expansion | Agent Review Board | No new repository, path, or capability enters an agent's allowlist without a board decision backed by measured precision on the current scope. |
| Methodology change | Experimentation & Analytics owner | No change to the quantification formula or the cohort definitions behind a published figure. Dollar figures must stay comparable quarter over quarter. |
8 · Observability of the agents themselves
- Every prompt, tool call, response, and decision streams to Splunk as structured events under the existing retention policy, threaded by one correlation ID from Dynatrace problem through case, Jira key, PR number, and release.
- Case files land in Snowflake for longitudinal evaluation — this is the dataset that tells the board whether precision is improving.
- Dynatrace monitors the agent services as first-class services. The system that watches the booking path is itself watched.
9 · The evaluation harness the proof case already owns
At the proof case, twelve quantified frictions were identified in one year and eight have been found and fixed, each with a human-ratified root cause and a human-ratified dollar figure. That is a golden regression set. Before any agent is trusted with a live case, it replays those twelve incidents from their original signals and is scored on two things: did it reach the same root cause, and did it land inside a tolerance band of the ratified number. This converts "do we trust the agent" from a matter of opinion into a measurement, and it is the single highest-leverage thing to build in the first sixty days.
10 · Graceful degradation
- Fullstory unavailable: Loop A continues to ops-only triage with a flagged evidence gap. It does not proceed to quantification or code.
- Model endpoint degraded or rate-limited: cases queue durably. Nothing is dropped, nothing is retried into a duplicate.
- Reconciliation repeatedly failing across cases: the orchestrator trips a global breaker and reverts the whole system to human triage. The system is designed to fail into the current process, which already works.
Agents run inside the customer's cloud perimeter. Models are consumed as managed in-region services. Fullstory MCP is reached outbound-only. No inbound internet path to the agent runtime.
Reference architecture
| Layer | Recommendation | Rationale |
|---|---|---|
| Ingress & control | ||
| Trigger ingress | API gateway with mTLS and HMAC signature verification, fronting a durable queue | Dynatrace and Qualtrics webhooks are the only inbound path. The queue absorbs alert bursts, survives downstream outages, and makes replay possible for incident forensics. |
| Agent runtime | Containerized on the customer's existing Kubernetes platform, private VPC, no inbound internet, scale-to-zero workers | This workload is bursty, not steady. Event-driven workers with reserved capacity only for the always-on funnel job matches spend to incidents. |
| Egress control | Explicit allowlist: Fullstory API, Dynatrace, Qualtrics, GitHub Enterprise, the model endpoint. Everything else denied. | The primary technical mitigation for prompt injection. An injected instruction has no reachable destination. |
| Reasoning | ||
| Primary model layer | Managed in-region multimodal endpoint with customer-managed encryption keys, no-training and no-retention terms, and VPC service controls | Multimodal reasoning over session screenshots and DOM state is a hard requirement, and code generation quality is the gating factor on Loop A's value. This is where the capability actually lives. |
| Model abstraction | Provider-agnostic routing layer; per-agent model selection; versions pinned per agent and promoted deliberately | The right reasoning model will change inside twelve months. Pinning per agent means a model upgrade is a measurable change against the golden regression set, not a surprise in production. |
| Self-hosted tier | Small open-weight classifier for theme classification and signature dedupe only | High volume, low judgment, no data egress, trivially cheap. This is where self-hosting genuinely pays. We do not recommend self-hosting the reasoning or code-generation tier: open-weight quality at realistic internal scale is not there, and GPU capacity planning for bursty incident traffic is poor use of capital. |
| Data & evidence | ||
| Fullstory MCP access | Outbound-only, scoped service credential per agent role, read-only for every agent except the authoring scope the quantification agent needs | The Session Evidence Agent's credential cannot compute metrics. The quantification agent's credential cannot open sessions. Credential scope mirrors the topology. |
| Snowflake — grounding | Parameterized, version-controlled queries over the clickstream already landing from Phase 3. The agent calls a reviewed query; it does not write SQL. | This is the strongest grounding move available. The dollar figure comes from a reviewed, versioned data model that analytics owns — not from a language model's arithmetic. It also makes every published figure reproducible by a human running the same query. |
| Snowflake — evaluation | Case files, agent decisions, and outcomes persisted as the longitudinal evaluation dataset | Precision trends, dollar-estimate error bands, and PR acceptance rates are queries against this table. Autonomy decisions are made from it. |
| Splunk | Audit sink for all agent activity, and the read source for the Ops Correlation Agent via saved searches with bound windows | One correlation ID chain end to end is what makes the loop auditable to risk and audit functions. Saved searches prevent unbounded query cost. |
| Dynatrace | Trigger source via Davis AI, and APM for the agent services themselves | Already the detection layer. Extending it over the agent runtime means agent latency and error rates are visible in the same place as everything else. |
| Security & privacy | ||
| Identity & secrets | Per-agent workload identity, short-lived tokens from the existing enterprise vault, separate GitHub App per agent role. No long-lived personal access tokens anywhere. | Separation of duties is only real if the credentials are actually separate. One compromised agent identity reaches one system with one scope. |
| Guest data in context | Tokenized identifiers only. Loyalty tier carried as a cohort label, never as identity. Screenshots retained in customer-controlled storage on a short TTL. | Fullstory's private-by-default capture and element-level exclusion mean payment and personal fields are excluded at the source, so the agent context does not contain them to begin with. Tokenization closes the remainder. |
| Prompt & response logs | Classified sensitive, retained in Splunk under existing policy, access-controlled to the review board and security | These logs are the audit record. They are also, unavoidably, the place where guest context would surface if a control failed — so they are governed accordingly. |
| Lower environments | Fullstory instrumentation in staging on a separate org, driven by synthetic traffic | Named prerequisite. Loop B's agentic validation step cannot exist without it. Not currently in place at the proof case. |
Autonomy expands as a function of measured precision, never by calendar date. Each phase must clear its metrics on the golden regression set before the next unlocks.
0–60d
Human in every step. No code generation at all.
Complete the bridge attribute across booking-path services. Stand up the case file, the correlation ID chain, and the Splunk audit sink. Grant governed read-only MCP access. Build the golden regression set from the twelve quantified frictions and score a human-supervised agent against it. Loop A runs to the evidence package and stops there.
60–120d
Agentic investigation, human decision.
Both loops run end to end through evidence and quantification. Jira issues auto-file with full provenance. Loop C begins in observation mode. Still no pull requests. This phase is where a VoC pipeline of the kind the proof case has already scoped becomes continuous rather than episodic.
120–180d
Supervised remediation on an allowlisted scope.
Draft pull requests enabled on a small set of low-risk repositories and paths — front-end presentation defects first, the stay-date display class of bug, where client-side evidence is also the localization. Two human gates on every case. Every merge performed by a person.
180d+
Closed-loop validation, scope widened by evidence.
Lower-environment deployment and agentic validation go live for Loop B. The path allowlist widens one scope at a time, each expansion a board decision backed by measured precision on the current scope. Production deployment remains outside agent capability permanently.
The metrics that govern the program
| Metric | Why it governs |
|---|---|
| Time to containment | The headline. Current human baseline is 48 hours from identification to deployed temporary fix. This is the number the program is accountable to. |
| Root-cause precision | Agent-proposed cause versus human-ratified cause on the golden set and, later, on live cases. The primary trust metric. |
| Dollar-estimate error band | Agent figure versus the eventually ratified figure. Credibility with leadership depends more on this being tight and stable than on it being large. |
| PR acceptance rate | Agent-authored draft PRs merged without substantive rework. The gate on expanding the path allowlist. |
| Abstention rate | Cases the system declined to conclude. Watched for being too low, which indicates hallucination, as much as for being too high. |
| Provenance strip rate | Claims rejected by the citation validator. A rising rate is the earliest available signal of agent drift. |
| Tracking-cause share | Cases resolved as instrumentation gaps rather than product defects. Both a health signal for capture governance and proof the agents are not forcing behavioral explanations. |
| Escaped defect rate | Production issues traceable to an agent-authored change. Target is zero, and a single occurrence pauses the phase. |
Governance
We recommend a standing Agent Review Board with representation from technology, digital innovation, experimentation & analytics, capture and instrumentation governance, and enterprise security and architecture. It owns four decisions and only four: scope expansion, methodology change, phase advancement, and incident review on any agent-attributable defect. It meets on the existing cadence rather than creating a new one, and it is the natural first charter for a Center of Excellence.
What a Phase 4a engagement looks like
Nothing in the first sixty days generates code or touches a repository. The engagement stands up the evidence chain and proves precision against incidents the customer has already solved by hand.
A 60-day Phase 4a pilot on Loop A, evidence-only
Dynatrace trigger through to a human-reviewed evidence package. No code generation in scope. Scored against the twelve already-quantified frictions.
Two prerequisites, each with a named owner
Bridge attribute coverage across booking and authentication path services, and Fullstory instrumentation in lower environments. Both are engineering tasks with clear completion criteria, and both gate program value.
Hosting and privacy posture ratified with security and architecture
In-region managed model endpoint with customer-managed keys and no-retention terms, private VPC runtime, allowlisted egress, tokenized identifiers, Splunk audit retention. Confirm the pattern before the build rather than during review.
An Agent Review Board, with unlock thresholds agreed in advance
Specifically: the precision figure and the dollar tolerance band that unlock Phase 4b, agreed in advance. Setting the bar before the first result is what keeps the program honest.
Open risks, stated plainly
- Behavioral evidence does not localize backend defects. Anything server-side depends on the Ops Correlation Agent. Where the bridge attribute is missing, the system has no localization path and will decline the case. Loop A's code generation therefore starts on front-end defects by design, not by caution.
- MCP non-determinism. Mitigated by methodology persistence and drift checks, not eliminated. Published figures must be treated as reproducible only when accompanied by their persisted methodology.
- Instrumentation drift. Until data-attribute and selector standards are consistent, a share of cases will correctly resolve as tracking causes. This is the system working, and it needs to be framed that way to leadership before the first such case lands.
- At the proof case, lower-environment instrumentation does not exist today. Loop B's closed validation step is contingent on it. Until then, Loop B ends at a human-approved Jira issue.
- Guest-supplied text is an attack surface. Mitigated by data-labelling, closed-taxonomy classification, egress allowlisting, and keeping the write-capable agent away from raw input. Worth an explicit security review rather than a footnote.