Fullstory
REFERENCE ARCHITECTURE
Reference Architecture  ·  Agentic Issue Resolution  ·  GSI Partner Edition

From Signal to Shipped Fix,
Without Guessing

A resilient agentic architecture for autonomous issue detection, root cause analysis, and remediation — built so that every conclusion the system reaches is traceable to observed end-user behavior, and every line of code it proposes passes through a human.

✓ Orchestrator + Observation Lead + 5 scoped specialists 2 mandatory human gates per loop Baseline to beat: the client’s manual time-to-containment
To
Systems-integrator delivery, architecture & digital-experience practice leaders
From
Fullstory — Strategic Solutions
Re
Reference architecture, hosting, and governance for APM- and voice-of-customer-triggered remediation loops
Decision
Select one to three joint-customer accounts for a 60-day, evidence-only Stage 1 pilot

Build a routing orchestrator over narrow specialists. Keep humans on the merge button.

Most enterprises have already proven the manual version of this loop: a customer complaint becomes a confirmed behavioral finding, a quantified impact, and a deployed containment. The work is not to invent that loop. It is to compress it and run it continuously without lowering the evidentiary bar that made it credible.

Our recommendation is a central orchestrator and an Observation Lead, routing to five tool-scoped specialist agents, with Fullstory MCP as the behavioral ground truth that constrains what the reasoning model is permitted to assert. Observation fans out by modality: one agent reads the visual and DOM record of a session, another reads its event timeline, and any disagreement between them is recorded and ruled on explicitly rather than blended silently. Agents that gather evidence are architecturally forbidden from explaining it. Agents that explain are required to reconcile two independent sources. The agent that writes code never holds the credentials that produced the evidence.

Autonomy expands as a function of measured precision, not calendar. Stage 1 ships with a human in every step and no code generation at all, and it can still beat the client’s manual time-to-evidence. Because many clients have not yet chosen an agentic runtime, these loops also serve as the specification that should drive that decision; Section 4 treats it as a platform selection rather than a fit to existing infrastructure.

Detail panels
One constraint shapes everything below

Before the architecture: a limit worth stating plainly, because designing around it is the difference between a system leadership trusts and a system that gets switched off after its first confident wrong answer.

Behavioral evidence proves that a failure happened and to whom. It does not localize the defect to a line of code. A session replay can show a user clicking a primary button three times against an error message, the DOM state at that instant, the failing network call, and the console error. That is decisive proof of impact. It is not proof of cause. Any agent that generates a pull request from session evidence alone will confidently invent a root cause.

The architectural consequence: server-side correlation carries equal weight, not a supporting role. The APM trace and the log are what localize the defect; Fullstory is what proves the defect mattered and sizes the money. Neither source alone authorizes a code change. This is why the recommended topology separates observation from inference, why the orchestrator requires two independent sources to agree before a case advances, and why autonomous PR generation begins on front-end presentation defects — the on-screen display-state class of bug (a wrong date, price, or status shown to the user), where the client-side evidence is the localization — and expands outward only as measured precision earns it.

Agentic topology

Recommendation: a central Triage Orchestrator and an Observation Lead, routing to five tool-scoped specialist agents. Strict hub-and-spoke, two hops at most, no peer-to-peer calls. Not a monolith, and not a swarm.

Triage Orchestrator case file · state machine · rules on disagreement Observation Lead fans out · records concordance · never resolves Cohort Quantification quantify Ops Correlation localize Remediation Engineering draft PRs only Visual / DOM Agent Agentic Session Review Event Stream Agent Event Stream Analysis Neither observer sees the other’s output before reporting. Reports return up through the Lead; the Orchestrator rules. Session Inference Brokering
Control plane

Triage Orchestrator

Owns the case file and the state machine. Decides which agent runs next, enforces the two-source reconciliation rule, enforces step and cost budgets, rules on any disagreement the Observation Lead surfaces, and posts to the human gates. Deliberately has no analytical tools and no write access to any system of record except the case store and the notification channel — it cannot investigate, and it cannot ship. Case records land in the client’s existing warehouse as append-only rows, so no new warehouse is introduced; the case table is a purpose-built append-only ledger that becomes the longitudinal evaluation dataset for governing autonomy expansion.

case_store dispatch_agent notify (chat) ticket_create
Cannot: call Fullstory, the APM, the log platform, or the repository directly. Cannot assert a root cause. Cannot advance a case whose claims lack provenance tokens.
Sub-orchestrator  ·  observation

Observation Lead

Takes a single observation request and fans it out to both modality specialists in parallel. Collects the two reports unblended and writes an explicit concordance record: for each material claim, whether the visual and event-stream observations agree, disagree, or one is silent. It never resolves a disagreement and never merges two reports into one narrative. It forwards both reports plus the concordance record up to the Triage Orchestrator, which rules.

dispatch_agent concordance_write
Cannot: call Fullstory directly. Cannot use causal language. Cannot resolve or average away a disagreement. Cannot call any specialist outside the observation pair.
Specialist  ·  observe only  ·  visual / DOM

Visual / DOM Agent — Agentic Session Review

The sole holder of the Fullstory session lease. Opens the exact session and timestamp, reads DOM and accessibility state, captures the visible failure, diffs the interaction, closes the lease. Emits observations with timestamps — never causes.

session_open session_view session_diff session_screenshot session_get_a11y_tree session_close
Cannot: use causal language. Cannot read the event-stream report. Cannot see the code repository. Cannot compute cohort metrics. System prompt rejects any output containing “because”, “caused by”, or a proposed fix.
Specialist  ·  observe only  ·  event timeline

Event Stream Agent — Event Stream Analysis

Reasons over the session’s event timeline. Three rulesets are built in as expertise rather than split into separate agents: friction and error events, journey structure (navigation and click patterns), and page performance. Error and rage clicks are surfaced as high priority only when there is observed impact on the user’s journey, never merely because they are present. Emits timestamped observations, never causes.

get_session_events get_sessions
Cannot: use causal language. Cannot hold the session lease. Cannot read the visual / DOM report. Cannot see the code repository. Cannot compute cohort metrics.
Specialist  ·  quantify

Cohort Quantification Agent

Sizes the cohort and the money. Confirms what a funnel or segment actually measures before computing it, then applies the client’s already-ratified formula: conversion delta between sessions with and without the friction, annualized, times average order value (or the client’s equivalent value metric). Returns sample size with every figure.

build_segment get_segment build_funnel get_funnel compute_funnel compute_metric get_opportunities warehouse_query (parameterized)
Cannot: write free-form SQL. Cannot publish a dollar figure below the 100-session floor. Cannot reference an object ID that does not exist in the org.
Specialist  ·  localize

Ops Correlation Agent

Pulls the failing trace and the surrounding logs for the same correlation ID and time window, and produces the server-side failure signature that actually localizes the defect. This is the agent whose output authorizes engineering work.

apm_problem_detail apm_trace log_search (saved)
Cannot: read Fullstory. Cannot write to any system. Cannot run unbounded log queries — saved searches with bound time windows only.
Specialist  ·  propose code

Remediation Engineering Agent

Runs only after a human approves the evidence package. Works in an ephemeral sandboxed worktree with no production network egress. Must produce a regression test that fails before its change and passes after — a branch without one is rejected automatically. Opens a draft pull request and stops.

repo_read branch_create commit pr_open_draft test_run
Cannot: merge, force-push, or touch protected branches. Cannot modify CI config, infrastructure-as-code, secrets, database migrations, or auth and payment modules. Cannot see raw end-user survey text. Cannot exceed the diff cap.
Why observation fans out by modalitySession Inference Brokering

The pattern is Session Inference Brokering: keep each class of inference about a session separate, and force disagreement between them to surface explicitly instead of being blended inside one agent.

  • Two different inference jobs. Reading a structured event timeline and judging what is salient in a rendered page or DOM are different skills with different failure modes. Inside a single Session Evidence Agent they are reconciled silently, and that reconciliation is a judgment call nobody labeled as one.
  • Disagreement is evidence. When the two observers disagree, the case file records it in the concordance record and the Triage Orchestrator rules on it, with the ruling logged. The system never inherits a blended answer it cannot audit.
  • Independent by construction. Neither observer sees the other’s output before reporting, so agreement between them means something.
  • The event-stream agent is deliberately not split further. Friction and error events, journey structure, and page performance are separate rulesets, but they are non-overlapping and never conflict with one another, so they live as built-in expertise inside one agent. Split agents when tool scope genuinely differs, not merely because the interpretive judgment does.
  • Priority discipline is an instruction, not a hope. Error and rage clicks are easy to over-weight. The Event Stream Agent is instructed to rank them as high priority only when impact on the user journey is observed.
Why not one monolithic agentFive reasons, led by the session lease

Four of these reasons are standard enterprise architecture. The first is specific to this toolchain and is, on its own, decisive.

  • The session lease is a single-holder resource. Fullstory’s visual inspection tools are stateful: a session is opened, inspected, and must be closed before another is opened in the same run. A monolithic agent that fans out across several sessions in parallel will leak or deadlock leases. Confining the lease to one single-threaded specialist (the Visual / DOM Agent) with a TTL and a reaper makes close-safety an architectural property rather than a prompt instruction the model may forget under pressure.
  • Tool-selection accuracy degrades with surface area. Fullstory MCP alone exposes dozens of tools (35 when this architecture was first written; confirm the current count). Add the APM, log platform, VoC platform, repository host, ticketing, warehouse and chat tools and a single agent is choosing among well over a hundred — in an incident context, under time pressure, where the wrong tool choice is a wrong answer. Each specialist here sees between two and eight.
  • Separation of duties limits blast radius. The agent that can write to the repository holds no Fullstory, APM, or log credential. The agents that gather evidence cannot see code. No single compromised or drifting context can both fabricate evidence and act on it.
  • Narrow contracts are testable. “Did the evidence agent correctly report the console error present in this session?” is a regression test. “Did the agent handle the incident well?” is not. Each modality can be evaluated and improved independently, which is what makes staged autonomy expansion possible.
  • Model routing and cost. Multimodal session inspection is the expensive step; routing and classification are nearly free. Per-agent model selection lets the client spend reasoning budget where judgment actually lives, and swap models without rewriting the system.
Why not a peer-to-peer multi-agent swarmLinear audit trail, fixed specialist count
  • No agent calls a peer. Control flows down from the orchestrator and reports flow back up, through at most one lead. The two observation specialists never call each other. This keeps the audit trail linear — one ordered case file per incident — and makes emergent loops structurally impossible rather than merely discouraged.
  • Intake and routing are deterministic code, not inference. Webhook parsing, signature verification, dedupe, and threshold checks run as ordinary services. The model is invoked only where judgment is genuinely required, which shrinks both the hallucination surface and the bill.
  • Specialist count stays fixed. Capability grows by widening a specialist’s allowlist, never by spawning new agents at runtime.
Closed-loop workflows

Two trigger loops and one always-on monitor. Gold rows are mandatory human checkpoints. Red rows are stop conditions where the system is designed to decline rather than proceed.

Loop A  ·  Infrastructure / Ops (APM-triggered)A0–A9 · ops signal to merged fix

Target: ops signal to approved evidence package in minutes; signal to merged fix within the same business day. A session-link correlation attribute (the Fullstory session URL, carried on the failing server-side trace) is the bridge that makes the whole loop possible — and its absence is the loop's first stop condition.

StepOwnerAction & data inputTool callsHand-off
A0 Intake service
deterministic
APM problem event arrives by webhook. Verify HMAC signature, extract problem ID, affected service, error signature, impacted request count, time window, and the session-link attribute. Hash the error signature and dedupe against open cases. webhook_verifysignature_hashdedupe New case, or increment an existing one
A0•STOP Intake service No session link, no agentic RCA. If the failing trace carries no Fullstory URL attribute, the case is routed to human triage with the ops context only. No model is invoked. At high traffic volumes a single popular failure can emit thousands of events — dedupe by signature means one case per signature per window, not one case per event. — Human triage queue
A1 Orchestrator Open an append-only case file. Assign case ID. Thread the correlation ID: APM problem → case → ticket key → PR → release. Set step, token, and wall-clock budgets. case_store.create Dispatch A2 and A3 in parallel
A2 Ops Correlation Retrieve the failing trace and the surrounding logs for the same correlation ID and window. Produce the server-side failure signature: service, method, status, exception class, first-seen, rate. apm_problem_detailapm_tracelog_search Technical finding → case file
A3 Observation Lead Availability check first: parse device and session IDs from the bridge URL and confirm the session is within the retention window and has finished indexing. If unavailable — due to indexing lag, retention expiry, or a missing bridge attribute — surface the session IDs to the human triage queue with an unavailability note and close the case branch; do not proceed to A4. If available, fan out to the Visual / DOM Agent and the Event Stream Agent in parallel. Neither sees the other’s output. dispatch_agent Dispatch A3a and A3b
A3a Visual / DOM Agent Open the session at the event timestamp, capture the accessibility tree for DOM state, a screenshot of the visible failure, and a diff across the failing interaction. Close the lease. session_opensession_viewsession_get_a11y_treesession_screenshotsession_diffsession_close Visual observation record → Lead
A3b Event Stream Agent Read the session’s event timeline around the failure window: friction and error events, journey structure, and page performance. Rank error and rage clicks high only where journey impact is observed. Observations only, no causes. get_session_events Event-stream observation record → Lead
A3c Observation Lead Write the concordance record: per material claim, whether the two observation reports agree, disagree, or one is silent. Do not resolve or blend. Forward both reports and the record, unaltered, to the orchestrator. concordance_write Both reports + concordance → case file
A4•STOP Orchestrator Reconciliation gate. The orchestrator receives both observation reports and the Lead’s concordance record. Where the observers disagree on a material claim, it rules explicitly and logs the ruling. Then: does the client-side evidence corroborate the server-side signature inside the same time window? If not, the case is marked unreconciled and escalated to a human. It does not proceed to quantification and it certainly does not proceed to code. reconcile Escalate, or advance to A5
A5 Cohort Quantification Build the segment of sessions exhibiting the signature. Read the relevant funnel definition, confirm what it measures, then compute with and without the signature. Apply the ratified formula. Return affected users, conversion delta, annualized exposure, and sample size. build_segmentget_funnelcompute_funnelcompute_metricwarehouse_query Quantified impact → case file
A6•HITL 1 Named engineering owner Evidence review. The complete package — technical signature, behavioral observations with session links and timestamps, screenshot, dollar exposure with sample size, and the methodology used — posts to chat and files a ticket. Approve, reject, or request more evidence. Nothing reaches the repository before this. notifyticket_create Approval unlocks A7
A7 Remediation Engineering Clone to an ephemeral sandboxed worktree. Locate the presentation-layer defect — CSS state, DOM rendering, component logic — using the session observation and the reconciled server signature as the localization boundary. Write the minimal fix plus a regression test that reproduces the failing condition. Run the existing suite. Open a draft PR whose body carries the full evidence package and the agent-generated label. Initial scope: front-end and presentation-layer defects only. Backend service defects (server-side exception, API behavior) produce the evidence package but require human engineering investigation to localize; automated PR generation for backend defects is a later stage, gated on measured precision. repo_readbranch_createtest_runpr_open_draft Draft PR → CODEOWNERS
A8•HITL 2 CODEOWNERS reviewer Code review and merge. Normal review process, normal release train. The agent has no merge permission at the repository-app level — this gate is enforced by platform configuration, not by instruction. — Merge → release
A9 Cohort Quantification Post-release verification. Re-compute the identical funnel and segment on the post-deploy window and re-check the APM problem state. Confirm resolution, or reopen the case automatically. compute_funnelcompute_metricapm_problem_detail Close, or reopen at A1
Loop B  ·  Voice of Customer (survey-triggered)B0–B13 · survey signal, triangulated

This loop differs from Loop A in one structural way, and the difference matters more than anything else in it: a single negative survey is an anecdote. Loop B will not run root cause analysis on a sample of one.

StepOwnerAction & data inputTool callsHand-off
B0 Intake service
deterministic
Survey response webhook on a CSAT or NPS score below threshold, or a negative verbatim. Extract the embedded Fullstory session link, survey metadata, segment attributes (region, tier), and device. Treat the verbatim as untrusted data from this moment forward. webhook_verifyextract_session_link Theme classifier
B1 Classifier
small model
Map the verbatim to a fixed theme taxonomy — for example date handling, redemption, sign-in, price display, order change, payment. Closed vocabulary, not free-form generation. Output is a label and a confidence, nothing else. classify_theme Theme + confidence → cluster store
B2•STOP Cluster service Cluster threshold. A case opens only when at least K responses share a theme on the same page or flow inside a rolling window, or a single response correlates to a live APM problem. Everything else accumulates as aggregate VoC signal. Responses with no session link follow the aggregate-only path and never trigger RCA. Starting value: K = 5, 7-day rolling window. This is a calibration starting point derived from one engagement’s data, not a production commitment. The Agent Review Board ratifies the threshold before Loop B goes live and adjusts it after the first 90 days of production data; K is a governed parameter, not a hardcoded constant. cluster_eval Aggregate dashboard, or open a case
B3 Orchestrator Open the case. Select up to five representative sessions across the cluster — not one — spanning device, segment, and tier. Dispatch three independent investigations in parallel. case_store.createget_sessions Dispatch B4, B5, B6
B4 Classifier Subjective. Structured summary of what users say is wrong, across the cluster: theme, frequency, affected tiers, verbatim exemplars quoted and attributed, never paraphrased into a cause. voc_responses Subjective finding
B5 Ops Correlation Technical. Were there errors, latency excursions, or failing traces on the implicated flow during the cluster window? A negative result here is a finding, not a failure — it points the case toward a UX or content defect rather than an engineering one. apm_problem_detaillog_search Technical finding
B6 Observation Lead
Visual / DOM + Event Stream
Behavioral. For each representative session: availability check first — confirm the session is within retention and fully indexed. Sessions from the cluster window may be days old; sessions approaching retention expiry or not yet indexed are noted with their IDs and routed to human triage rather than inspected. For available sessions the Lead fans out to both modality specialists: the Visual / DOM Agent walks sessions in turn (lease held by one agent, one session at a time) and captures what the user actually saw, while the Event Stream Agent reads the timeline for dead and rage interactions, journey structure, and the point of abandonment. The Lead returns both reports plus a concordance record, unblended. A partially available cluster proceeds with the sessions that are available, provided at least two remain across distinct device and segment dimensions. session_opensession_viewsession_get_a11y_treesession_screenshotsession_closeget_session_events Observation records + concordance
B7•STOP Orchestrator Two-of-three triangulation. At least two of the three independent findings must agree on the same defect before the case is declared validated. One source alone — however vivid the verbatim — does not advance the case. Any observer disagreement inside the behavioral finding is ruled on and logged first. reconcile Escalate, or advance to B8
B8 Cohort Quantification Size the exposure beyond the survey respondents — the respondents are a sample, and a biased one. Build the behavioral segment matching the validated defect across all traffic, compute the conversion delta, annualize against AOV. build_segmentcompute_funnelcompute_metricwarehouse_query Quantified impact
B9•HITL 1 CX + Product owner Validated issue review. Ticket with the triangulated evidence, the cohort exposure, and the recommended disposition. This creates a visible, quantified pipeline of customer pain that product squads can prioritize. ticket_createnotify Approval unlocks B10
B10 Remediation Engineering Fix plus regression test, deployed to a lower environment only. The agent receives the structured finding — never the raw user text. Production deployment is not in this agent's capability set at any stage. branch_createtest_rundeploy_lowerpr_open_draft Lower-env build
B11 Observation Lead + Quantification Agentic validation in the lower environment. Synthetic traffic drives the repaired flow; the observation agents inspect the resulting sessions and confirms the defect no longer reproduces. Two named prerequisites, both required: (1) Fullstory instrumentation in staging on a separate org; (2) a confirmed synthetic traffic mechanism — APM synthetic monitors if the client already operates them, or a scripted browser automation suite (Playwright, Selenium) that can be aimed at the lower environment. The mechanism must be identified and owned before Loop B can be validated end-to-end. This is a scoping item for the Stage 1 kickoff, not an assumption to be resolved at build time. get_sessionssession_opensession_viewsession_closecompute_funnel Validation verdict
B12•HITL 2 Release management Production release approval. Normal release train, normal change control. The validation verdict is an input to the human decision, not a substitute for it. — Production release
B13 Orchestrator Close on theme decay, not on merge. Monitoring continues; the case closes when the survey theme rate returns to baseline and the behavioral segment shrinks. A merged PR is not a resolved customer experience. cluster_evalcompute_metric Close, or reopen at B3
Loop C  ·  Always-on funnel vigilanceHourly control-limit monitor

Loops A and B are both reactive — they wait for infrastructure to notice or a user to complain. The highest-value failures are the silent ones that do neither.

A scheduled job recomputes the critical funnels hourly against statistical control limits rather than fixed thresholds, so normal daily and weekly seasonality does not generate noise. A step-level drop outside the control band opens a case at step A1 of Loop A, bypassing the APM trigger entirely. This is the pattern that catches the failure class that never throws a server error and never produces a survey response: the price that renders empty, the date that silently reverts, the balance that displays as unavailable. Early cases from this loop should run evidence-only, because without a server-side signature there is nothing to localize — which means Loop C feeds human investigation first, and earns code-generation rights last.

Resiliency & hallucination defense

Ten controls. The first three are the ones that matter most, and none of them is a prompt instruction — they are enforced by deterministic code between the agents.

The failure mode

One model gathers and explains

Root-cause hallucination has a specific origin. A single context window takes in partial evidence, feels pressure to produce an answer, and generates a plausible causal story that reads exactly like a real finding. It cites the session it looked at. It sounds confident. It is wrong in a way no reviewer can detect without redoing the work.

  • Confident causes from a single source
  • Dollar figures with no sample size
  • Object IDs that do not exist in the org
  • Behavioral explanations for tracking gaps
The control

Observation and inference are different agents

The Session Evidence Agent is permitted to say what it saw and when. It is architecturally prevented from saying why. Inference happens only in the orchestrator's reconciliation step, which cannot run on one source. The model never gets to hold an uncited claim across a hand-off.

  • Every claim carries a provenance token
  • Two independent sources or the case stops
  • Named-object allowlist, no invented IDs
  • Abstention is a first-class outcome
1  ·  The citation contractEvery claim carries provenance

Every assertion written into a case file must carry a provenance token: a session ID plus timestamp, an APM problem ID, a saved-search hash, or a metric ID plus the computed value. A deterministic validator sits between every agent hand-off and strips any claim without one. The stripped claim is logged — a rising strip rate is an early warning that an agent is drifting, and it is one of the metrics the review board watches.

// case-file assertion — rejected without provenance { "claim": "Users saw an error message when submitting an order", "provenance": [ { "type": "session_observation", "session": "<id>", "ts": "20:52:14.330Z", "artifact": "a11y_tree + screenshot" }, { "type": "apm_problem", "id": "<problem>", "signature": "POST /order/commit 500" } ], "inference_allowed": true // two independent sources present }
2  ·  How Fullstory MCP grounds the reasoningFour mechanisms

"Use MCP as ground truth" is a slogan until it is four specific mechanisms. These are the four.

  • Confirm the definition before computing it. The quantification agent must read a funnel, segment, or metric definition and state in the case file what it is about to measure, then compute. This single discipline eliminates the most common analytical error: reporting "checkout drop" from an object that measures something adjacent to checkout.
  • Named-object allowlist. Agents may reference only objects that exist in the org — the critical-path conversion funnels, key confirmation pages, the sign-in failure segment, the watched elements already active on critical front-end error messages. If the object required to answer the question does not exist, the agent must say so and stop. Approximating a missing object is the mechanism by which an agent invents confidence.
  • Visual and DOM verification. The accessibility tree and screenshot give falsifiable client-side truth. A claimed error message must be present in the captured DOM or visible in the frame. This turns "the agent believes users saw an error" into a check that either passes or fails.
  • Methodology persistence against MCP non-determinism. The analytics MCP routes through a reasoning intermediary, so the same question asked twice can resolve to different query strategies. Left alone, two agent runs on one incident produce two different dollar figures — and nothing destroys executive trust faster. Each case therefore persists the object IDs, the original query strings, the time window, and the exclusions applied. Re-runs use the saved IDs with a drift check; if a saved ID returns empty, the agent reports drift rather than silently rebuilding.
3  ·  Instrumentation fragility is a first-class outcomeA tracking cause is a rewarded outcome
Most accounts' configuration audits surface two gaps that bound agent reliability: inconsistent data attributes and CSS selector standards across the site, and iOS and Android SDKs not yet at full parity with web. Element-level agent reasoning inherits that fragility directly. An agent that cannot tell "the user did not click it" from "we stopped capturing it" will report a behavior change that is actually a selector change.

The control is a pre-flight instrumentation health check on every case: are the elements and pages this case depends on being captured stably across the window under examination? If not, the agent returns a tracking cause, not a behavioral one, and routes to the capture-governance queue. This must be a rewarded outcome in the agent's evaluation, not a failed run. A system that can only ever find product defects will find product defects in instrumentation gaps.

4  ·  Confidence floors and abstentionAbstention is a success state
  • Cohort floor. No dollar figure below 100 sessions. Below the floor, the finding ships as directional with the sample size stated and no monetary estimate attached.
  • Corroboration floor. Two independent sources minimum, as above. For Loop B, two of three.
  • Reproduction floor. No PR without a regression test that fails before the change and passes after. This is checked by CI, not asserted by the agent.
  • Abstention is success. "Insufficient evidence, escalating" is a valid terminal state and is measured as such. An agent whose abstention rate is near zero is not accurate — it is hallucinating, and the metric will show it.
5  ·  Destructive-action preventionCapability-scoped at the platform layer

Capability-scoped at the platform layer. A prompt that says "do not merge" is a suggestion; a repository app without merge permission is a guarantee.

ControlEnforcement
No merge, everDedicated repository app with contents:write on unprotected branches and pull_requests:write only. No merge scope, no admin scope. Branch protection and CODEOWNERS enforced server-side.
Path denylistCI configuration, infrastructure-as-code, secrets, database migrations, and authentication and payment modules are blocked at the pre-commit hook. Touching them requires an explicit human-set flag on the case.
Diff capChanges above roughly 200 lines or 5 files are refused. Above the cap the agent files an analysis issue with its findings instead of a branch — a large diff is a design decision, not a hotfix.
Sandboxed executionEphemeral worktree, no production network egress, no production credentials in the environment, destroyed at case close.
Provenance on every PRagent-generated label plus the full evidence package in the body. No silent authorship anywhere in the repository.
IdempotencyEvery external write carries an idempotency key derived from the case ID and step. A retry cannot double-file a ticket or double-open a PR.
Circuit breakerRepeated failures on the same error signature trip the breaker and route to humans. Per-case step, token, and wall-clock budgets terminate runaway loops.
Session lease safetySingle-holder lease with TTL and a reaper process. Close precedes open, always. Only the Session Evidence Agent holds session tools at all.
6  ·  Prompt injection through end-user-supplied textNo agent holds a credential
Loop B ingests free text written by members of the public. A user can type instructions into a survey comment box. Session DOM content can contain attacker-controlled strings. Any of it can reach a model context that also holds repository write capability.
  • No agent holds a credential. An agent holding a ticketing or repository token is one successful injection away from an attacker holding that token. Agents hold capability references and a broker holds the secrets, so a compromised context can do exactly what it was granted and nothing else. This is the strongest available mitigation and it is architectural, not textual — see 4.3b.
  • All survey verbatims and all captured DOM content are wrapped and labelled as untrusted data, never as instruction, at every hand-off.
  • The classifier emits a label from a closed taxonomy — it does not generate free text downstream.
  • The Remediation Engineering Agent never receives raw verbatim or raw DOM text. It receives only the structured, human-approved finding. The agent with write capability is the agent furthest from untrusted input.
  • Outbound tool calls are allowlisted by host. An injected instruction has nowhere to send anything.
7  ·  Human-in-the-loop checkpoints, stated plainlyFour gates; nothing ships without a human
GateOwnerWhat cannot happen before it
Evidence review (A6 / B9)Named engineering owner; CX and Product for Loop BNo repository access of any kind. No branch, no commit, no PR. The agent's investigation is complete and the human decides whether it is right.
Code review and merge (A8 / B12)CODEOWNERS; Release managementNo merge to any protected branch. No production deployment. Enforced by platform permissions, not policy.
Scope expansionAgent Review BoardNo new repository, path, or capability enters an agent's allowlist without a board decision backed by measured precision on the current scope.
Methodology changeExperimentation & Analytics ownerNo change to the quantification formula or the cohort definitions behind a published figure. Dollar figures must stay comparable quarter over quarter.
Production deployment is never an agent capability. Not in Stage 1, and not at full maturity. The ceiling on autonomy here is a validated draft PR and a lower-environment deployment. Every production change goes through the client's existing release train, reviewed and merged by a person.
8  ·  Observability of the agents themselvesThe agents are monitored too
  • Every prompt, tool call, response, and decision streams to the client's log platform as structured events under existing retention policy, threaded by one correlation ID from APM problem through case, ticket key, PR number, and release.
  • Case files land in the client's warehouse for longitudinal evaluation — this is the dataset that tells the board whether precision is improving.
  • The APM monitors the agent services as first-class services. The system that watches the critical path is itself watched.
9  ·  The evaluation harness the client already ownsThe golden regression set

Most enterprises hold a history of resolved incidents, each with a human-ratified root cause and a human-ratified dollar figure. That is a golden regression set. Before any agent is trusted with a live case, it replays those incidents from their original signals and is scored on two things: did it reach the same root cause, and did it land inside a tolerance band of the ratified number. This converts “do we trust the agent” from a matter of opinion into a measurement, and it is the single highest-leverage thing to build in the first sixty days. It is also how an integrator proves a build is trustworthy rather than asserting it.

10  ·  Graceful degradationFails into the current process
  • Fullstory unavailable: Loop A continues to ops-only triage with a flagged evidence gap. It does not proceed to quantification or code.
  • Model endpoint degraded or rate-limited: cases queue durably. Nothing is dropped, nothing is retried into a duplicate.
  • Reconciliation repeatedly failing across cases: the orchestrator trips a global breaker and reverts the whole system to human triage. The system is designed to fail into the current process, which already works.
Platform selection

No agentic runtime has been chosen yet. That is an advantage, not a gap — these two loops are a sharp enough specification to drive the decision, and they are a far better selection test than a generic pilot would be.

What follows is a decision framework, not a procurement recommendation. Vendor capabilities in this category are moving monthly, and the specifics below should be probed directly in vendor conversations rather than taken as settled. What is stable, and what should anchor the evaluation, is the set of requirements this workload places on a platform. Those come from the loops and will not change.
4.1  ·  Six requirements that should drive the decisionWhat the loops demand of a platform

Most agent-platform evaluations start from vendor feature matrices and end in a tie. Starting from what these loops actually demand eliminates most of the field on the first pass.

RequirementWhere it comes fromWhat it rules out
Multi-week durable pauses Loop A pauses at the evidence gate for hours or days. Loop B pauses at issue review, lower-environment deploy, validation, and a release train — then closes on theme decay, which is weeks. Case state must outlive any process, any redeploy, and any model call. Message queues as the control plane. State held in a pod or a function. Most agent frameworks' built-in persistence, which is designed for conversation memory rather than multi-week business process state.
Stateful single-holder tool leases Fullstory's session inspection tools are opened, used, and must be closed before the next opens. Exactly one agent may hold the lease, with a TTL and a reaper. Naive parallel fan-out over a shared tool registry. Any runtime that cannot express a mutex or a singleton activity.
Per-agent credential scoping The whole separation-of-duties argument is real only if the credentials are actually separate: the evidence agent's token cannot compute metrics, the remediation agent's token reaches the repository and nothing else. A single shared service account across agents. Agents inheriting a human's permissions. This is the requirement the cloud platforms are weakest on — and the one we recommend solving with a dedicated control plane rather than with a cloud primitive. See 4.3b.
Multimodal reasoning plus strong code generation Session screenshots and accessibility trees are visual evidence. Loop A's value ceiling is set by code-generation quality. Open-weight-only model tiers at realistic internal scale. Text-only endpoints.
Minutes-long jobs with real disk The remediation agent clones a repository and runs a test suite. The evidence agent walks up to five sessions in sequence. A functions-only architecture. Short execution ceilings and ephemeral-only storage.
Immutable audit trail, signal to release One correlation ID from APM problem through case, ticket key, PR number, and release. Required for risk and audit, and it is the dataset that governs autonomy expansion. Any runtime that will not export structured per-step traces. Black-box managed services whose internal state you cannot query.
4.1b  ·  Three compute tiers — why each existsDurable engine, container jobs, functions

These loops require three distinct layers of compute. Conflating them is the most common platform selection error.

  • Durable execution engine (Temporal, Step Functions, Cloud Workflows, or Durable Functions): holds the case state machine across multi-week pauses, human-approval gates, redeployments, and model failures. This is the control plane — it does not run the agents, it orchestrates their invocation and records every state transition for audit. Without this layer, a case that pauses at HITL review loses its context the moment the container restarts.
  • Container jobs (scale-to-zero): the agent compute tier. Session inspection, cohort quantification, and remediation all run as minutes-long jobs with real disk — the remediation agent clones a repository and runs a test suite; the session evidence agent walks up to five sessions in sequence. A functions-only substrate cannot host these. Container jobs that scale to zero when idle keep cost proportional to case volume.
  • Lightweight functions: webhook receipt, HMAC signature verification, alert deduplication, and theme classification. These are deterministic, milliseconds-long, high-volume operations that require no model call. Running them in the durable engine or in a container job wastes capacity and adds latency at the ingest boundary.

The table below maps each layer to the cloud-specific implementation options. The client's cloud decision will be driven by commercial terms and existing operational muscle — the capability requirements above are cloud-neutral and should anchor the evaluation regardless of which path is chosen.

4.2  ·  Candidate stacksAWS, Google, Azure, cloud-neutral

The cloud decision will most likely be made on grounds that have nothing to do with agents — existing gravity, commercial terms, where enterprise architecture already has operational muscle. A client's current estate often gives weak or conflicting signals. So rather than pick a cloud, here is what each path looks like and where each one is weak for this workload.

LayerAWS-anchoredGoogle-anchoredAzure-anchoredCloud-neutral option
Durable execution Step Functions, or Temporal on AWS Cloud Workflows, or Temporal on GCP Durable Functions / Durable Task, or Temporal on Azure Temporal (Cloud or self-hosted). Also Restate, Inngest.
Agent compute ECS Fargate or Bedrock AgentCore runtime Cloud Run jobs (recommended); GKE is a viable option if the client already operates Kubernetes at scale — the additional cluster overhead is not justified for this workload alone Azure Container Apps jobs Containers on any managed scale-to-zero job service
Model access Bedrock — broad multi-model including strong code-generation options Vertex AI — Gemini-native Azure AI Foundry — widest catalog, fastest model turnover Provider-agnostic gateway in front of one or two endpoints
Agent identity IAM roles — not bound to workforce identity today Service accounts — integration work required Entra ID — furthest along on agent-native identity and conditional access Sekizui — agents hold certificates (mTLS, SPIFFE), never secrets; grants are deny-by-default and revocable
Tool governance No cloud ships this as a complete answer yet — see 4.3b Sekizui — policy enforcement point and credential broker; MCP tool lists pinned under version control
Grounding store The client's existing warehouse in all three cases Existing warehouse, with reviewed parameterized queries
Audit & APM The client's existing log and APM platforms in all three cases — no reason to change either Existing log and APM platforms
One cloud-native caveat worth surfacing early. Vertex's native grounding advantage is against BigQuery. If the client's behavioral warehouse is a different vendor, that particular benefit does not transfer — the quantification path runs through reviewed warehouse queries regardless of cloud. Weigh the Google path on Gemini and operational fit, not on native grounding.
4.3  ·  The two decisions that actually require deliberationDurable engine; framework or thin orchestration

Most of the table above resolves itself once a cloud is picked. Two choices do not, they are genuinely portable, and they carry the most consequence. These are where we would spend the architecture group's time.

Decision 1 — the durable execution engine

✓

Recommendation: Temporal

Multi-month workflow state and human-approval signals are its core competency, not an extension of it. Workflows are defined in code, which matters because these loops branch on evidence quality rather than following a linear path.

⋮

It survives the cloud decision

Since no cloud is chosen, a portable control plane means this choice is not re-litigated later. Namespace isolation also maps cleanly onto per-specialist credential boundaries.

⚠

The honest tradeoff

Operational burden if self-hosted — Temporal Cloud removes most of it at a cost. The cloud-native engines are cheaper and simpler if the client is confident in a single cloud and comfortable expressing this branching in a state-machine DSL. Step Functions is genuinely excellent; it is also AWS-bound.

Decision 2 — framework or thin orchestration

✓

Recommendation: thin, with a framework only inside single steps

These loops are deterministic workflows with model calls inside them — not autonomous agent loops. The orchestrator is an explicit state machine. That is workflow code, and an agent framework is the wrong tool for it.

✗

Two sources of truth is the failure mode

Most frameworks bring their own state and persistence. Layered over a durable execution engine, you get two competing records of where a case stands — and case state is the thing auditors will ask about.

⋮

Where a framework does earn its place

Inside an individual reasoning step — the session evidence extraction, the classifier — running as one activity with scoped tools. Framework choice then becomes a reversible, step-local decision rather than a platform commitment.

4.3b  ·  The agent-to-system control planeCredential broker and policy enforcement

Requirement three in 4.1 — per-agent credential scoping — is the one no cloud platform answers completely, and it is the requirement the entire separation-of-duties argument rests on. We recommend solving it with a dedicated control plane rather than waiting for a cloud primitive to mature.

Recommendation: Sekizui (github.com/fullstorydev/sekizui) — a policy enforcement point and credential broker sitting between the agent fleet and every system it acts on. Stated plainly: this is Fullstory open source, MIT licensed. We are recommending a component we wrote, and the license is why that should be acceptable rather than suspect — the client or integrator can read it, run it, fork it, and keep running it with no dependence on any commercial relationship with us. Evaluate it on the architecture, not on who authored it.

The obvious benefit is arithmetic: five agents needing six systems is thirty integrations, each with its own credential, audit gap, rate limit and error handling. A broker makes that N + M instead of N × M. That collapse is worth having but it is not the reason to adopt it. The reason is that it turns three controls this proposal describes as design intent into properties the infrastructure actually enforces.

What this memo specifiedWhat the control plane enforces instead of trusting
Per-agent credential scoping
requirement 3
Agents hold certificates — mutual TLS with SPIFFE identities — and never secrets. Grants are deny-by-default, written and reviewed as action × target × constraints, and revocable. A compromised agent can do exactly what it was granted and nothing more. Break-glass refuses new work, cancels calls already in flight, and records how many it killed.
Scoped tool allowlists per agent MCP servers sit behind a vetted contract: the vendor's tool list is pinned under version control and compared against what the server actually serves. Unvetted tools are unreachable, unvetted arguments are refused, results that break their declared shape are refused, and drift is recorded. Critically — a capability denied on a native connector cannot be obtained by going through that vendor's MCP server instead.
Immutable audit trail
requirement 6
One hash-chained log that records decisions, not just actions — including refusals, which rule matched, who asked on whose behalf, and what it cost. "What did this agent do to production this week" becomes a query rather than an investigation. This is the record the Agent Review Board governs autonomy expansion from.
Prompt-injection containment
section 3.6
Loop B ingests text written by the public. An agent holding a repository token is one successful injection away from an attacker holding that token. With credentials brokered, the blast radius of a successful injection is the grant set, not the credential set. Egress allowlisting and data-labelling remain; this sits beneath them.
End-user data minimisation Two mechanisms. Connectors declare field by field what they may return, and anything undeclared is stripped before a consumer, a rule, or the audit log sees it — a connector whose declaration is unsound is quarantined at startup rather than run. Then per-consumer lenses that can only ever remove: "analytics never receives email addresses" is one line, enforced on the way out. A lens never narrows the audit log.
Cost and loop control
section 3.7
Every call and every scheduled poll is priced at its worst case against the target's budget and the deployment ceiling before it is sent. Ceilings exist that no grant can exceed — policy answers "may you," ceilings answer "should anyone, ever."
Data residency A deployment ceiling plus a grant layer, with a cross-region resolution refused before any credential moves. Relevant for a global estate where EU end-user data cannot be processed outside its region.

What it changes in the loops

Sections 2 and 3 specify that intake, deduplication and theme classification must be deterministic code rather than model inference, precisely so those steps present no injection surface. That primitive has a name here: a reflex — a deterministic event-to-action rule that fires with no model in the loop, with its condition checked against the schema at boot, its own firing budget, and shadow mode as the default so the dangerous option is never the easy one. A reflex is a principal and takes the identical enforcement path an agent does, so it is not a bypass.

  • A0, B0 and B1 — webhook verification, signature-hash deduplication, and closed-taxonomy theme classification — are reflexes rather than bespoke services.
  • Loop C, the hourly funnel check against control limits, becomes a scheduled poll feeding a reflex. Durable cursors matter here: a poll cursor survives kill -9, re-delivers at most one event, and names a window it could not re-read rather than silently skipping it. A monitor that loses an hour without saying so is worse than no monitor.
  • Shadow mode as default maps directly onto the phased rollout. Stage 1 can run the full reflex set in shadow — recording what each rule would have fired — which is a materially better Stage 1 than one that simply has the capability switched off.
The honest cost, in the project's own words: centralising credentials makes this control plane the highest-value target in the system. The trade is that one small, hardened, audited surface is more defensible than secrets distributed across a fleet of prompt-injectable agents. That is the right trade for this workload, but it is a real concentration of risk and it should be reviewed as such by enterprise security — not waved through because the component is well-designed.

What the client should probe before committing

  • Connector coverage is the practical gating question. These loops need the APM, log platform, VoC platform, repository host and warehouse alongside Fullstory. Confirm which connectors exist today versus which are scoped work, because a missing connector is either an integration to build or a system that sits outside the governed path — and a system outside the path defeats the purpose.
  • Maturity. It reached v1 recently. Ask for production references, the results of any external hardening review, and what the upgrade cadence looks like.
  • Who operates it. MIT licensing means the client or integrator can run it; it also means someone must run it. Decide whether that sits with platform engineering or security, and what the support path is when it refuses to start — which it will do by design rather than hold a credential it cannot rotate or audit.
  • Secret-manager integration is a prerequisite, not a follow-on. Credentials arrive from a mounted file under a declared root, from platform workload identity, or from OAuth. That integration has to be settled before the Stage 1 pilot, not during it.
4.4  ·  What we would deliberately not buy yetFour things to defer
  • A managed agent runtime as the foundation. The hosted agent services from all three clouds are credible and worth a look — several are framework-agnostic, so they are closer to runtimes than to frameworks. But they are optimized for conversational and single-agent shapes, their built-in session and memory models overlap awkwardly with durable execution plus the citation contract, and agent-native identity — the requirement that matters most here — is the least mature thing any of them ships. Pilot one. Do not found the program on one.
  • An eval platform on the critical path. The golden-set scoring is plain code against the warehouse, where case files already land. Tracing UX is worth buying once the team feels the pain of not having it — not in the first sixty days, and not as a dependency of the pilot.
  • A dedicated vector store. Nothing in either loop is a retrieval problem. Evidence arrives by ID from the APM and Fullstory; quantification runs on governed SQL. If a vector database appears in a proposed architecture for this workload, ask what question it answers.
  • GPU capacity. Managed in-region inference with customer-managed keys and no-retention terms covers it. Capacity planning for bursty incident traffic is a poor use of capital, and the self-hosted tier we do recommend — theme classification and signature dedupe — is small enough to run on CPU.
4.5  ·  Security and privacy posture — cloud-independentCloud-independent controls

None of this changes with the platform decision, so it can be ratified with security and architecture in parallel with the selection rather than after it.

ControlPosture
IngressAPI gateway with mTLS and HMAC signature verification is the only inbound path. APM and survey webhooks land on a durable queue that feeds the workflow engine, so alert bursts are absorbed and replay is possible for forensics.
EgressHost allowlist: Fullstory, the APM, the VoC platform, the repository host, the model endpoint. Everything else denied. This is the primary technical mitigation for prompt injection — an injected instruction has no reachable destination.
Agent identity & secretsAgents hold capability references, not credentials. Identity by mutual TLS and SPIFFE; grants deny-by-default, reviewable and revocable; break-glass cancels calls already in flight. Separate repository app per agent role remains, but no agent context ever contains a token. This is what makes separation of duties real rather than aspirational — see 4.3b.
Fullstory accessOutbound-only, brokered, scoped per agent role. The evidence agent cannot compute metrics; the quantification agent cannot open sessions. Enforced as grants on a single path rather than by issuing each agent its own key — and a capability denied on the native connector cannot be obtained through the MCP server instead.
End-user data in contextTokenized identifiers only. Segment or tier as a cohort label, never as identity. Fullstory's private-by-default capture and element-level exclusion mean payment and personal fields are excluded at source, so the agent context does not contain them to begin with.
Model termsIn-region inference, customer-managed encryption keys, no training on client data, no retention. Versions pinned per agent; a model change is a measured event against the golden regression set, not a silent upgrade.
Prompt & response logsClassified sensitive, retained in the client's log platform under existing policy, access-controlled to the review board and security. These are the audit record.
Lower environmentsFullstory instrumentation in staging on a separate org, driven by synthetic traffic. Named prerequisite — Loop B's validation step cannot exist without it, and it is not in place today.
4.6  ·  Why this should be the first agent a client buildsWhy start here, and the fair counterargument

If the platform decision is open, the choice of first workload is a decision about what the platform has to prove. We would argue for this one, on four grounds — and then give the honest counterargument.

  • It has a hard, pre-existing success criterion. Most first agents fail quietly because nobody can say whether they worked. This one is measured against the client's manual time-to-containment and its already-ratified incident history. The verdict is arithmetic, not opinion.
  • Stage 1 is read-only. No write capability for the first sixty days, so the blast radius during platform shakedown is close to zero. You are load-testing a runtime with an agent that cannot break anything.
  • It exercises every hard capability the platform will ever need. Multi-week durable pauses, human gates, stateful tool leases, per-agent credential scoping, multimodal reasoning, write-capable tooling under blast-radius control, and an end-to-end audit trail. A conversational pilot exercises roughly one of those. Anything the client builds after this one inherits a runtime that has already been proven against the difficult case.
  • It pays for the platform. Recovered revenue is the funding argument for everything downstream, and this is the only candidate first workload that produces a dollar figure as its native output.
The counterargument, stated fairly: this is not the easiest first agent. A read-only question-answering agent over the warehouse would ship faster, with less governance overhead and fewer prerequisites. If the goal is a quick visible win, build that one. If the goal is to make a platform decision that holds — and to find out in sixty days rather than eighteen months whether the chosen runtime can carry real agentic work — this is the better first build. We think the second goal is the one worth optimizing for, and the read-only Stage 1 scope is what makes the risk of it acceptable.
Two prerequisites carry the whole program regardless of platform, and both are concrete and assignable. First, the session-link correlation attribute must be present on every service in the critical paths (for example checkout and authentication) — where it is missing, Loop A is blind, and the system will correctly refuse to run. Second, lower-environment Fullstory instrumentation, without which Loop B's closed validation step is aspiration rather than architecture. Neither depends on the runtime decision, so both can start now.
Phasing, governance & decision asks

Autonomy expands as a function of measured precision, never by calendar date. Each stage must clear its metrics on the golden regression set before the next unlocks.

Staged autonomyFour stages, unlocked by measured precision
Stage 1
0–60d

Human in every step. No code generation at all.

Complete the bridge attribute across critical-path services. Stand up the case file, the correlation ID chain, and the audit sink. Grant governed read-only MCP access. Build the golden regression set from the client’s resolved incidents and score a human-supervised agent against it. Loop A runs to the evidence package and stops there.

Unlock criteria: root-cause precision on the regression set, dollar estimates inside the agreed tolerance band, and a measured time-to-evidence materially under the client's manual containment baseline.
Stage 2
60–120d

Agentic investigation, human decision.

Both loops run end to end through evidence and quantification. Tickets auto-file with full provenance. Loop C begins in observation mode. Still no pull requests. This is the stage where the voice-of-customer pipeline becomes continuous rather than episodic.

Unlock criteria: stable precision across both loops, abstention rate inside the healthy band, and zero unreconciled cases advanced in error.
Stage 3
120–180d

Supervised remediation on an allowlisted scope.

Draft pull requests enabled on a small set of low-risk repositories and paths — front-end presentation defects first, the on-screen display-state class of bug, where client-side evidence is also the localization. Two human gates on every case. Every merge performed by a person.

Unlock criteria: PR acceptance rate above the agreed threshold, zero escaped defects attributable to an agent-authored change, and no denylist violations.
Stage 4
180d+

Closed-loop validation, scope widened by evidence.

Lower-environment deployment and agentic validation go live for Loop B. The path allowlist widens one scope at a time, each expansion a board decision backed by measured precision on the current scope. Production deployment remains outside agent capability permanently.

Steady state: recovered revenue attributed to agent-originated cases becomes a reported line in the quarterly value readout.
The metrics that govern the programEight metrics
MetricWhy it governs
Time to containmentThe headline. Measured against the client's manual baseline from identification to deployed temporary fix. This is the number the program is accountable to.
Root-cause precisionAgent-proposed cause versus human-ratified cause on the golden set and, later, on live cases. The primary trust metric.
Dollar-estimate error bandAgent figure versus the eventually ratified figure. Credibility with leadership depends more on this being tight and stable than on it being large.
PR acceptance rateAgent-authored draft PRs merged without substantive rework. The gate on expanding the path allowlist.
Abstention rateCases the system declined to conclude. Watched for being too low, which indicates hallucination, as much as for being too high.
Provenance strip rateClaims rejected by the citation validator. A rising rate is the earliest available signal of agent drift.
Tracking-cause shareCases resolved as instrumentation gaps rather than product defects. Both a health signal for capture governance and proof the agents are not forcing behavioral explanations.
Escaped defect rateProduction issues traceable to an agent-authored change. Target is zero, and a single occurrence pauses the stage.
GovernanceThe Agent Review Board

We recommend a standing Agent Review Board with representation from digital technology, digital innovation, experimentation & analytics, capture and instrumentation governance, and enterprise security and architecture. It owns four decisions and only four: scope expansion, methodology change, stage advancement, and incident review on any agent-attributable defect. It meets on the client’s existing cadence rather than creating a new one, and it is a natural first charter for an AI or data Center of Excellence.

Four asks to begin a joint Stage 1 pilot

Nothing in the first sixty days generates code or touches a repository. The ask is to stand up the evidence chain and prove precision against incidents the client has already solved by hand.

Ask 1

Name two or three overlap accounts as friendly pilot candidates

Joint customers where the integrator already has delivery presence and the client already runs Fullstory. Stage 1 is Loop A, evidence-only, scored against the client’s own resolved incidents.

Ask 2

Name a champion on each side

One delivery or architecture lead who carries the pilot internally, and one counterpart who owns the instrumentation prerequisites.

Ask 3

Assign the two prerequisites with named owners per account

Session-link attribute coverage across critical-path services, and Fullstory instrumentation in lower environments. Both are engineering tasks with clear completion criteria, and both gate program value.

Ask 4

Charter the Agent Review Board and agree the unlock thresholds

Specifically: the precision figure and the dollar tolerance band that unlock Stage 2, agreed in advance. Setting the bar before the first result is what keeps the program honest.

Open risks, stated plainlyKnown risks and how each is mitigated
  • Behavioral evidence does not localize backend defects. The Ops Correlation Agent carries the server-side weight. Where the bridge attribute is missing, the system has no localization path and will decline the case. Loop A's code generation therefore starts on front-end defects by design, not by caution.
  • Fan-out adds a hop. The Observation Lead adds a coordination point and some latency. It is justified by making observer disagreement explicit and auditable, and should be measured against a single-agent baseline during Stage 1.
  • MCP non-determinism. Mitigated by methodology persistence and drift checks, not eliminated. Published figures must be treated as reproducible only when accompanied by their persisted methodology.
  • Instrumentation drift. Until data-attribute and selector standards are consistent, a share of cases will correctly resolve as tracking causes. This is the system working, and it needs to be framed that way to leadership before the first such case lands.
  • Lower-environment instrumentation may not exist. Loop B's closed validation step is contingent on it. Until then, Loop B ends at a human-approved ticket.
  • End-user-supplied text is an attack surface. Mitigated by data-labelling, closed-taxonomy classification, egress allowlisting, and keeping the write-capable agent away from raw input. Worth an explicit security review rather than a footnote.