Before the architecture: a limit worth stating plainly, because designing around it is the difference between a system leadership trusts and a system that gets switched off after its first confident wrong answer.
The architectural consequence: server-side correlation is load-bearing, not supporting. The Dynatrace trace and the Splunk log are what localize the defect; Fullstory is what proves the defect mattered and sizes the money. Neither source alone authorizes a code change. This is why the recommended topology separates observation from inference, why the orchestrator requires two independent sources to agree before a case advances, and why autonomous PR generation begins on front-end presentation defects — the stay-date display class of bug, where the client-side evidence is the localization — and expands outward only as measured precision earns it.
Recommendation: a central Triage Orchestrator routing to four tool-scoped specialist agents, strict hub-and-spoke, no peer-to-peer calls. Not a monolith, and not a swarm.
Triage Orchestrator
Owns the case file and the state machine. Decides which specialist runs next, enforces the two-source reconciliation rule, enforces step and cost budgets, and posts to the human gates. Deliberately has no analytical tools and no write access to any system of record except the case store and the notification channel — it cannot investigate, and it cannot ship. Case records land in Snowflake as append-only rows — Snowflake already holds Marriott's behavioral clickstream from Phase 3, so no new warehouse is introduced; the case table is a purpose-built append-only ledger that becomes the longitudinal evaluation dataset for governing autonomy expansion.
Session Evidence Agent
The sole holder of the Fullstory session lease. Opens the exact session and timestamp, reads DOM and accessibility state, captures the visible failure, diffs the interaction, closes the lease. Emits observations with timestamps — never causes.
Cohort Quantification Agent
Sizes the cohort and the money. Confirms what a funnel or segment actually measures before computing it, then applies Marriott's already-ratified formula: conversion delta between sessions with and without the friction, annualized, times AOV. Returns sample size with every figure.
Ops Correlation Agent
Pulls the failing trace and the surrounding logs for the same correlation ID and time window, and produces the server-side failure signature that actually localizes the defect. This is the agent whose output authorizes engineering work.
Remediation Engineering Agent
Runs only after a human approves the evidence package. Works in an ephemeral sandboxed worktree with no production network egress. Must produce a regression test that fails before its change and passes after — a branch without one is rejected automatically. Opens a draft pull request and stops.
Why not one monolithic agent
Four of these reasons are standard enterprise architecture. The first is specific to this toolchain and is, on its own, decisive.
- The session lease is a single-holder resource. Fullstory's visual inspection tools are stateful: a session is opened, inspected, and must be closed before another is opened in the same run. A monolithic agent that fans out across several sessions in parallel will leak or deadlock leases. Confining the lease to one single-threaded specialist with a TTL and a reaper makes close-safety an architectural property rather than a prompt instruction the model may forget under pressure.
- Tool-selection accuracy degrades with surface area. Fullstory MCP alone exposes 35 tools. Add Dynatrace, Splunk, Qualtrics, GitHub, Jira, Snowflake and Teams and a single agent is choosing among well over a hundred — in an incident context, under time pressure, where the wrong tool choice is a wrong answer. Each specialist here sees between three and eight.
- Separation of duties limits blast radius. The agent that can write to the repository holds no Fullstory, Dynatrace, or Splunk credential. The agent that gathers evidence cannot see code. No single compromised or drifting context can both fabricate evidence and act on it.
- Narrow contracts are testable. "Did the evidence agent correctly report the console error present in this session?" is a regression test. "Did the agent handle the incident well?" is not. Specialists can be evaluated and improved independently, which is what makes staged autonomy expansion possible.
- Model routing and cost. Multimodal session inspection is the expensive step; routing and classification are nearly free. Per-agent model selection lets Marriott spend reasoning budget where judgment actually lives, and swap models without rewriting the system.
Why not a peer-to-peer multi-agent swarm
- No agent calls another agent. All control flows through the orchestrator. This keeps the audit trail linear — one ordered case file per incident — and makes emergent loops structurally impossible rather than merely discouraged.
- Intake and routing are deterministic code, not inference. Webhook parsing, signature verification, dedupe, and threshold checks run as ordinary services. The model is invoked only where judgment is genuinely required, which shrinks both the hallucination surface and the bill.
- Specialist count stays fixed. Capability grows by widening a specialist's allowlist, never by spawning new agents at runtime.
Two trigger loops and one always-on monitor. Gold rows are mandatory human checkpoints. Red rows are stop conditions where the system is designed to decline rather than proceed.
Loop A · Infrastructure / Ops (Dynatrace-triggered)
Target: ops signal to approved evidence package in minutes; signal to merged fix within the same business day. The X-FullStory-URL request attribute on the failing PurePath is the bridge that makes the whole loop possible — and its absence is the loop's first stop condition.
| Step | Owner | Action & data input | Tool calls | Hand-off |
|---|---|---|---|---|
| A0 | Intake service deterministic |
Davis AI problem event arrives by webhook. Verify HMAC signature, extract problem ID, affected service, error signature, impacted request count, time window, and the X-FullStory-URL attribute. Hash the error signature and dedupe against open cases. | webhook_verifysignature_hashdedupe | New case, or increment an existing one |
| A0•STOP | Intake service | No session link, no agentic RCA. If the failing trace carries no Fullstory URL attribute, the case is routed to human triage with the ops context only. No model is invoked. At 4.7M visits a day a single popular failure can emit thousands of events — dedupe by signature means one case per signature per window, not one case per event. | — | Human triage queue |
| A1 | Orchestrator | Open an append-only case file. Assign case ID. Thread the correlation ID: Dynatrace problem → case → Jira key → PR → release. Set step, token, and wall-clock budgets. | case_store.create | Dispatch A2 and A3 in parallel |
| A2 | Ops Correlation | Retrieve the failing PurePath and the surrounding logs for the same correlation ID and window. Produce the server-side failure signature: service, method, status, exception class, first-seen, rate. | dynatrace_problem_detaildynatrace_purepathsplunk_search | Technical finding → case file |
| A3 | Session Evidence | Parse device and session IDs from the bridge URL. Availability check first: confirm the session is within the retention window and has finished indexing before opening the lease. If unavailable — due to indexing lag, retention expiry, or a missing bridge attribute — surface the session IDs to the human triage queue with an unavailability note and close the case branch; do not proceed to A4. If available: open the session at the event timestamp, capture the accessibility tree for DOM state, a screenshot of the visible failure, and a diff across the failing interaction. Close the lease. | session_opensession_viewsession_get_a11y_treesession_screenshotsession_diffsession_close | Observation record → case file |
| A4•STOP | Orchestrator | Reconciliation gate. Does the client-side observation corroborate the server-side signature inside the same time window? If the two do not agree, the case is marked unreconciled and escalated to a human. It does not proceed to quantification and it certainly does not proceed to code. | reconcile | Escalate, or advance to A5 |
| A5 | Cohort Quantification | Build the segment of sessions exhibiting the signature. Read the checkout funnel definition, confirm what it measures, then compute with and without the signature. Apply the ratified formula. Return affected users, conversion delta, annualized exposure, and sample size. | build_segmentget_funnelcompute_funnelcompute_metricsnowflake_query | Quantified impact → case file |
| A6•HITL 1 | Named engineering owner | Evidence review. The complete package — technical signature, behavioral observations with session links and timestamps, screenshot, dollar exposure with sample size, and the methodology used — posts to Teams or Slack and files a Jira issue. Approve, reject, or request more evidence. Nothing reaches the repository before this. | notifyjira_create | Approval unlocks A7 |
| A7 | Remediation Engineering | Clone to an ephemeral sandboxed worktree. Locate the presentation-layer defect — CSS state, DOM rendering, component logic — using the session observation and the reconciled server signature as the localization boundary. Write the minimal fix plus a regression test that reproduces the failing condition. Run the existing suite. Open a draft PR whose body carries the full evidence package and the agent-generated label. Phase 4a scope: front-end and presentation-layer defects only. Backend service defects (server-side exception, API behavior) produce the evidence package but require human engineering investigation to localize; automated PR generation for backend defects is Phase 4b, gated on measured precision from Phase 4a. | repo_readbranch_createtest_runpr_open_draft | Draft PR → CODEOWNERS |
| A8•HITL 2 | CODEOWNERS reviewer | Code review and merge. Normal review process, normal release train. The agent has no merge permission at the GitHub App level — this gate is enforced by platform configuration, not by instruction. | — | Merge → release |
| A9 | Cohort Quantification | Post-release verification. Re-compute the identical funnel and segment on the post-deploy window and re-check the Dynatrace problem state. Confirm resolution, or reopen the case automatically. | compute_funnelcompute_metricdynatrace_problem_detail | Close, or reopen at A1 |
Loop B · Voice of Customer (Qualtrics-triggered)
This loop differs from Loop A in one structural way, and the difference matters more than anything else in it: a single negative survey is an anecdote. Loop B will not run root cause analysis on a sample of one.
| Step | Owner | Action & data input | Tool calls | Hand-off |
|---|---|---|---|---|
| B0 | Intake service deterministic |
Qualtrics response webhook on a CSAT or NPS score below threshold, or a negative verbatim. Extract the embedded Fullstory session link, survey metadata, brand, property, loyalty tier, and device. Treat the verbatim as untrusted data from this moment forward. | webhook_verifyextract_session_link | Theme classifier |
| B1 | Classifier small model |
Map the verbatim to a fixed theme taxonomy — date handling, points redemption, sign-in, rate display, reservation change, payment. Closed vocabulary, not free-form generation. Output is a label and a confidence, nothing else. | classify_theme | Theme + confidence → cluster store |
| B2•STOP | Cluster service | Cluster threshold. A case opens only when at least K responses share a theme on the same page or flow inside a rolling window, or a single response correlates to a live Dynatrace problem. Everything else accumulates as aggregate VoC signal. Responses with no session link follow the aggregate-only path and never trigger RCA. Starting value: K = 5, 7-day rolling window. This is a calibration starting point derived from the 12-friction dataset, not a production commitment. The Agent Review Board ratifies the threshold before Loop B goes live and adjusts it after the first 90 days of production data; K is a governed parameter, not a hardcoded constant. | cluster_eval | Aggregate dashboard, or open a case |
| B3 | Orchestrator | Open the case. Select up to five representative sessions across the cluster — not one — spanning device, brand, and loyalty tier. Dispatch three independent investigations in parallel. | case_store.createget_sessions | Dispatch B4, B5, B6 |
| B4 | Classifier | Subjective. Structured summary of what guests say is wrong, across the cluster: theme, frequency, affected tiers, verbatim exemplars quoted and attributed, never paraphrased into a cause. | qualtrics_responses | Subjective finding |
| B5 | Ops Correlation | Technical. Were there errors, latency excursions, or failing traces on the implicated flow during the cluster window? A negative result here is a finding, not a failure — it points the case toward a UX or content defect rather than an engineering one. | dynatrace_problem_detailsplunk_search | Technical finding |
| B6 | Session Evidence | Behavioral. For each representative session in the cluster: availability check first — confirm the session is within retention and fully indexed. Sessions from the cluster window may be days old; sessions approaching retention expiry or not yet indexed are noted with their IDs and routed to human triage rather than inspected. Walk available sessions in turn — lease held by one agent, one session at a time. Capture what the guest actually saw and did: DOM state, visible messaging, dead and rage interactions, the point of abandonment. A partially available cluster (some sessions unreachable) proceeds with the sessions that are available, provided at least two remain across distinct device and brand dimensions. | session_opensession_viewsession_get_a11y_treesession_screenshotsession_close | Observation records |
| B7•STOP | Orchestrator | Two-of-three triangulation. At least two of the three independent findings must agree on the same defect before the case is declared validated. One source alone — however vivid the verbatim — does not advance the case. | reconcile | Escalate, or advance to B8 |
| B8 | Cohort Quantification | Size the exposure beyond the survey respondents — the respondents are a sample, and a biased one. Build the behavioral segment matching the validated defect across all traffic, compute the conversion delta, annualize against AOV. | build_segmentcompute_funnelcompute_metricsnowflake_query | Quantified impact |
| B9•HITL 1 | CX + Product owner | Validated issue review. Jira issue with the triangulated evidence, the cohort exposure, and the recommended disposition. This is the queue the account's VoC initiative already calls for: a visible, quantified pipeline of guest pain that product squads can prioritize. | jira_createnotify | Approval unlocks B10 |
| B10 | Remediation Engineering | Fix plus regression test, deployed to a lower environment only. The agent receives the structured finding — never the raw guest text. Production deployment is not in this agent's capability set at any phase. | branch_createtest_rundeploy_lowerpr_open_draft | Lower-env build |
| B11 | Session Evidence + Quantification | Agentic validation in the lower environment. Synthetic traffic drives the repaired flow; the agent inspects the resulting sessions and confirms the defect no longer reproduces. Two named prerequisites, both required: (1) Fullstory instrumentation in staging on a separate org — not in place today; (2) a confirmed synthetic traffic mechanism — Dynatrace synthetic monitors if Marriott already operates them, or a scripted browser automation suite (Playwright, Selenium) that can be aimed at the lower environment. The mechanism must be identified and owned before Loop B can be validated end-to-end. This is a scoping item for the Phase 4a kickoff, not an assumption to be resolved at build time. | get_sessionssession_opensession_viewsession_closecompute_funnel | Validation verdict |
| B12•HITL 2 | Release management | Production release approval. Normal release train, normal change control. The validation verdict is an input to the human decision, not a substitute for it. | — | Production release |
| B13 | Orchestrator | Close on theme decay, not on merge. Monitoring continues; the case closes when the Qualtrics theme rate returns to baseline and the behavioral segment shrinks. A merged PR is not a resolved guest experience. | cluster_evalcompute_metric | Close, or reopen at B3 |
Loop C · Always-on funnel vigilance
Loops A and B are both reactive — they wait for infrastructure to notice or a guest to complain. The highest-value failures are the silent ones that do neither.
A scheduled job recomputes the checkout funnels hourly against statistical control limits rather than fixed thresholds, so normal daily and weekly seasonality does not generate noise. A step-level drop outside the control band opens a case at step A1 of Loop A, bypassing the Dynatrace trigger entirely. This is the pattern that catches the failure class that never throws a server error and never produces a survey response: the rate card that renders empty, the date that silently reverts, the points balance that displays as unavailable. Early cases from this loop should run evidence-only, because without a server-side signature there is nothing to localize — which means Loop C feeds human investigation first, and earns code-generation rights last.
Ten controls. The first three are the ones that matter most, and none of them is a prompt instruction — they are enforced by deterministic code between the agents.
One model gathers and explains
Root-cause hallucination has a specific origin. A single context window takes in partial evidence, feels pressure to produce an answer, and generates a plausible causal story that reads exactly like a real finding. It cites the session it looked at. It sounds confident. It is wrong in a way no reviewer can detect without redoing the work.
- Confident causes from a single source
- Dollar figures with no sample size
- Object IDs that do not exist in the org
- Behavioral explanations for tracking gaps
Observation and inference are different agents
The Session Evidence Agent is permitted to say what it saw and when. It is architecturally prevented from saying why. Inference happens only in the orchestrator's reconciliation step, which cannot run on one source. The model never gets to hold an uncited claim across a hand-off.
- Every claim carries a provenance token
- Two independent sources or the case stops
- Named-object allowlist, no invented IDs
- Abstention is a first-class outcome
1 · The citation contract
Every assertion written into a case file must carry a provenance token: a session ID plus timestamp, a Dynatrace problem ID, a saved-search hash, or a metric ID plus the computed value. A deterministic validator sits between every agent hand-off and strips any claim without one. The stripped claim is logged — a rising strip rate is an early warning that an agent is drifting, and it is one of the metrics the review board watches.
2 · How Fullstory MCP grounds the reasoning
"Use MCP as ground truth" is a slogan until it is four specific mechanisms. These are the four.
- Confirm the definition before computing it. The quantification agent must read a funnel, segment, or metric definition and state in the case file what it is about to measure, then compute. This single discipline eliminates the most common analytical error: reporting "checkout drop" from an object that measures something adjacent to checkout.
- Named-object allowlist. Agents may reference only objects that exist in the org — the checkout conversion funnels, the confirmation and rate-list pages, the sign-in failure segment, the watched elements already active on critical front-end error messages. If the object required to answer the question does not exist, the agent must say so and stop. Approximating a missing object is the mechanism by which an agent invents confidence.
- Visual and DOM verification. The accessibility tree and screenshot give falsifiable client-side truth. A claimed error message must be present in the captured DOM or visible in the frame. This turns "the agent believes guests saw an error" into a check that either passes or fails.
- Methodology persistence against MCP non-determinism. The analytics MCP routes through a reasoning intermediary, so the same question asked twice can resolve to different query strategies. Left alone, two agent runs on one incident produce two different dollar figures — and nothing destroys executive trust faster. Each case therefore persists the object IDs, the original query strings, the time window, and the exclusions applied. Re-runs use the saved IDs with a drift check; if a saved ID returns empty, the agent reports drift rather than silently rebuilding.
3 · Instrumentation fragility is a first-class outcome
The control is a pre-flight instrumentation health check on every case: are the elements and pages this case depends on being captured stably across the window under examination? If not, the agent returns a tracking cause, not a behavioral one, and routes to the capture-governance queue. This must be a rewarded outcome in the agent's evaluation, not a failed run. A system that can only ever find product defects will find product defects in instrumentation gaps.
4 · Confidence floors and abstention
- Cohort floor. No dollar figure below 100 sessions. Below the floor, the finding ships as directional with the sample size stated and no monetary estimate attached.
- Corroboration floor. Two independent sources minimum, as above. For Loop B, two of three.
- Reproduction floor. No PR without a regression test that fails before the change and passes after. This is checked by CI, not asserted by the agent.
- Abstention is success. "Insufficient evidence, escalating" is a valid terminal state and is measured as such. An agent whose abstention rate is near zero is not accurate — it is hallucinating, and the metric will show it.
5 · Destructive-action prevention
Capability-scoped at the platform layer. A prompt that says "do not merge" is a suggestion; a GitHub App without merge permission is a guarantee.
| Control | Enforcement |
|---|---|
| No merge, ever | Dedicated GitHub App with contents:write on unprotected branches and pull_requests:write only. No merge scope, no admin scope. Branch protection and CODEOWNERS enforced server-side. |
| Path denylist | CI configuration, infrastructure-as-code, secrets, database migrations, and authentication and payment modules are blocked at the pre-commit hook. Touching them requires an explicit human-set flag on the case. |
| Diff cap | Changes above roughly 200 lines or 5 files are refused. Above the cap the agent files an analysis issue with its findings instead of a branch — a large diff is a design decision, not a hotfix. |
| Sandboxed execution | Ephemeral worktree, no production network egress, no production credentials in the environment, destroyed at case close. |
| Provenance on every PR | agent-generated label plus the full evidence package in the body. No silent authorship anywhere in the repository. |
| Idempotency | Every external write carries an idempotency key derived from the case ID and step. A retry cannot double-file a Jira issue or double-open a PR. |
| Circuit breaker | Repeated failures on the same error signature trip the breaker and route to humans. Per-case step, token, and wall-clock budgets terminate runaway loops. |
| Session lease safety | Single-holder lease with TTL and a reaper process. Close precedes open, always. Only the Session Evidence Agent holds session tools at all. |
6 · Prompt injection through guest-supplied text
- No agent holds a credential. An agent holding a Jira or GitHub token is one successful injection away from an attacker holding that token. Agents hold capability references and a broker holds the secrets, so a compromised context can do exactly what it was granted and nothing else. This is the strongest available mitigation and it is architectural, not textual — see 4.3b.
- All survey verbatims and all captured DOM content are wrapped and labelled as untrusted data, never as instruction, at every hand-off.
- The classifier emits a label from a closed taxonomy — it does not generate free text downstream.
- The Remediation Engineering Agent never receives raw verbatim or raw DOM text. It receives only the structured, human-approved finding. The agent with write capability is the agent furthest from untrusted input.
- Outbound tool calls are allowlisted by host. An injected instruction has nowhere to send anything.
7 · Human-in-the-loop checkpoints, stated plainly
| Gate | Owner | What cannot happen before it |
|---|---|---|
| Evidence review (A6 / B9) | Named engineering owner; CX and Product for Loop B | No repository access of any kind. No branch, no commit, no PR. The agent's investigation is complete and the human decides whether it is right. |
| Code review and merge (A8 / B12) | CODEOWNERS; Release management | No merge to any protected branch. No production deployment. Enforced by platform permissions, not policy. |
| Scope expansion | Agent Review Board | No new repository, path, or capability enters an agent's allowlist without a board decision backed by measured precision on the current scope. |
| Methodology change | Experimentation & Analytics owner | No change to the quantification formula or the cohort definitions behind a published figure. Dollar figures must stay comparable quarter over quarter. |
8 · Observability of the agents themselves
- Every prompt, tool call, response, and decision streams to Splunk as structured events under the existing retention policy, threaded by one correlation ID from Dynatrace problem through case, Jira key, PR number, and release.
- Case files land in Snowflake for longitudinal evaluation — this is the dataset that tells the board whether precision is improving.
- Dynatrace monitors the agent services as first-class services. The system that watches the booking path is itself watched.
9 · The evaluation harness Marriott already owns
Twelve quantified frictions have been identified this year and eight have been found and fixed, each with a human-ratified root cause and a human-ratified dollar figure. That is a golden regression set. Before any agent is trusted with a live case, it replays those twelve incidents from their original signals and is scored on two things: did it reach the same root cause, and did it land inside a tolerance band of the ratified number. This converts "do we trust the agent" from a matter of opinion into a measurement, and it is the single highest-leverage thing to build in the first sixty days.
10 · Graceful degradation
- Fullstory unavailable: Loop A continues to ops-only triage with a flagged evidence gap. It does not proceed to quantification or code.
- Model endpoint degraded or rate-limited: cases queue durably. Nothing is dropped, nothing is retried into a duplicate.
- Reconciliation repeatedly failing across cases: the orchestrator trips a global breaker and reverts the whole system to human triage. The system is designed to fail into the current process, which already works.
No agentic runtime has been chosen yet. That is an advantage, not a gap — these two loops are a sharp enough specification to drive the decision, and they are a far better selection test than a generic pilot would be.
4.1 · Six requirements that should drive the decision
Most agent-platform evaluations start from vendor feature matrices and end in a tie. Starting from what these loops actually demand eliminates most of the field on the first pass.
| Requirement | Where it comes from | What it rules out |
|---|---|---|
| Multi-week durable pauses | Loop A pauses at the evidence gate for hours or days. Loop B pauses at issue review, lower-environment deploy, validation, and a release train — then closes on theme decay, which is weeks. Case state must outlive any process, any redeploy, and any model call. | Message queues as the control plane. State held in a pod or a function. Most agent frameworks' built-in persistence, which is designed for conversation memory rather than multi-week business process state. |
| Stateful single-holder tool leases | Fullstory's session inspection tools are opened, used, and must be closed before the next opens. Exactly one agent may hold the lease, with a TTL and a reaper. | Naive parallel fan-out over a shared tool registry. Any runtime that cannot express a mutex or a singleton activity. |
| Per-agent credential scoping | The whole separation-of-duties argument is real only if the credentials are actually separate: the evidence agent's token cannot compute metrics, the remediation agent's token reaches GitHub and nothing else. | A single shared service account across agents. Agents inheriting a human's permissions. This is the requirement the cloud platforms are weakest on — and the one we recommend solving with a dedicated control plane rather than with a cloud primitive. See 4.3b. |
| Multimodal reasoning plus strong code generation | Session screenshots and accessibility trees are visual evidence. Loop A's value ceiling is set by code-generation quality. | Open-weight-only model tiers at realistic internal scale. Text-only endpoints. |
| Minutes-long jobs with real disk | The remediation agent clones a repository and runs a test suite. The evidence agent walks up to five sessions in sequence. | A functions-only architecture. Short execution ceilings and ephemeral-only storage. |
| Immutable audit trail, signal to release | One correlation ID from Dynatrace problem through case, Jira key, PR number, and release. Required for risk and audit, and it is the dataset that governs autonomy expansion. | Any runtime that will not export structured per-step traces. Black-box managed services whose internal state you cannot query. |
4.1b · Three compute tiers — why each exists
These loops require three distinct layers of compute. Conflating them is the most common platform selection error.
- Durable execution engine (Temporal, Step Functions, Cloud Workflows, or Durable Functions): holds the case state machine across multi-week pauses, human-approval gates, redeployments, and model failures. This is the control plane — it does not run the agents, it orchestrates their invocation and records every state transition for audit. Without this layer, a case that pauses at HITL review loses its context the moment the container restarts.
- Container jobs (scale-to-zero): the agent compute tier. Session inspection, cohort quantification, and remediation all run as minutes-long jobs with real disk — the remediation agent clones a repository and runs a test suite; the session evidence agent walks up to five sessions in sequence. A functions-only substrate cannot host these. Container jobs that scale to zero when idle keep cost proportional to case volume.
- Lightweight functions: webhook receipt, HMAC signature verification, alert deduplication, and theme classification. These are deterministic, milliseconds-long, high-volume operations that require no model call. Running them in the durable engine or in a container job wastes capacity and adds latency at the ingest boundary.
The table below maps each layer to the cloud-specific implementation options. Marriott's cloud decision will be driven by commercial terms and existing operational muscle — the capability requirements above are cloud-neutral and should anchor the evaluation regardless of which path is chosen.
4.2 · Candidate stacks
The cloud decision will most likely be made on grounds that have nothing to do with agents — existing gravity, commercial terms, where enterprise architecture already has operational muscle. Marriott's current estate gives weak and conflicting signals: Snowflake is cloud-neutral, and the Gemini reference on the agentic platform slide hints Google without establishing it. So rather than pick a cloud, here is what each path looks like and where each one is weak for this workload.
| Layer | AWS-anchored | Google-anchored | Azure-anchored | Cloud-neutral option |
|---|---|---|---|---|
| Durable execution | Step Functions, or Temporal on AWS | Cloud Workflows, or Temporal on GCP | Durable Functions / Durable Task, or Temporal on Azure | Temporal (Cloud or self-hosted). Also Restate, Inngest. |
| Agent compute | ECS Fargate or Bedrock AgentCore runtime | Cloud Run jobs (recommended); GKE is a viable option if Marriott already operates Kubernetes at scale — the additional cluster overhead is not justified for this workload alone | Azure Container Apps jobs | Containers on any managed scale-to-zero job service |
| Model access | Bedrock — broad multi-model including strong code-generation options | Vertex AI — Gemini-native, matches the stated reasoning layer | Azure AI Foundry — widest catalog, fastest model turnover | Provider-agnostic gateway in front of one or two endpoints |
| Agent identity | IAM roles — not bound to workforce identity today | Service accounts — integration work required | Entra ID — furthest along on agent-native identity and conditional access | Sekizui — agents hold certificates (mTLS, SPIFFE), never secrets; grants are deny-by-default and revocable |
| Tool governance | No cloud ships this as a complete answer yet — see 4.3b | Sekizui — policy enforcement point and credential broker; MCP tool lists pinned under version control | ||
| Grounding store | Snowflake in all three cases — already landing clickstream from Phase 3 | Snowflake, with reviewed parameterized queries | ||
| Audit & APM | Splunk and Dynatrace in all three cases — no reason to change either | Splunk, Dynatrace | ||
4.3 · The two decisions that actually require deliberation
Most of the table above resolves itself once a cloud is picked. Two choices do not, they are genuinely portable, and they carry the most consequence. These are where we would spend the architecture group's time.
Decision 1 — the durable execution engine
Recommendation: Temporal
Multi-month workflow state and human-approval signals are its core competency, not an extension of it. Workflows are defined in code, which matters because these loops branch on evidence quality rather than following a linear path.
It survives the cloud decision
Since no cloud is chosen, a portable control plane means this choice is not re-litigated later. Namespace isolation also maps cleanly onto per-specialist credential boundaries.
The honest tradeoff
Operational burden if self-hosted — Temporal Cloud removes most of it at a cost. The cloud-native engines are cheaper and simpler if Marriott is confident in a single cloud and comfortable expressing this branching in a state-machine DSL. Step Functions is genuinely excellent; it is also AWS-bound.
Decision 2 — framework or thin orchestration
Recommendation: thin, with a framework only inside single steps
These loops are deterministic workflows with model calls inside them — not autonomous agent loops. The orchestrator is an explicit state machine. That is workflow code, and an agent framework is the wrong tool for it.
Two sources of truth is the failure mode
Most frameworks bring their own state and persistence. Layered over a durable execution engine, you get two competing records of where a case stands — and case state is the thing auditors will ask about.
Where a framework does earn its place
Inside an individual reasoning step — the session evidence extraction, the classifier — running as one activity with scoped tools. Framework choice then becomes a reversible, step-local decision rather than a platform commitment.
4.3b · The agent-to-system control plane
Requirement three in 4.1 — per-agent credential scoping — is the one no cloud platform answers completely, and it is the requirement the entire separation-of-duties argument rests on. We recommend solving it with a dedicated control plane rather than waiting for a cloud primitive to mature.
The obvious benefit is arithmetic: five agents needing six systems is thirty integrations, each with its own credential, audit gap, rate limit and error handling. A broker makes that N + M instead of N × M. That collapse is worth having but it is not the reason to adopt it. The reason is that it turns three controls this proposal describes as design intent into properties the infrastructure actually enforces.
| What this memo specified | What the control plane enforces instead of trusting |
|---|---|
| Per-agent credential scoping requirement 3 |
Agents hold certificates — mutual TLS with SPIFFE identities — and never secrets. Grants are deny-by-default, written and reviewed as action × target × constraints, and revocable. A compromised agent can do exactly what it was granted and nothing more. Break-glass refuses new work, cancels calls already in flight, and records how many it killed. |
| Scoped tool allowlists per agent | MCP servers sit behind a vetted contract: the vendor's tool list is pinned under version control and compared against what the server actually serves. Unvetted tools are unreachable, unvetted arguments are refused, results that break their declared shape are refused, and drift is recorded. Critically — a capability denied on a native connector cannot be obtained by going through that vendor's MCP server instead. |
| Immutable audit trail requirement 6 |
One hash-chained log that records decisions, not just actions — including refusals, which rule matched, who asked on whose behalf, and what it cost. "What did this agent do to production this week" becomes a query rather than an investigation. This is the record the Agent Review Board governs autonomy expansion from. |
| Prompt-injection containment section 3.6 |
Loop B ingests text written by the public. An agent holding a GitHub token is one successful injection away from an attacker holding that token. With credentials brokered, the blast radius of a successful injection is the grant set, not the credential set. Egress allowlisting and data-labelling remain; this sits beneath them. |
| Guest-data minimisation | Two mechanisms. Connectors declare field by field what they may return, and anything undeclared is stripped before a consumer, a rule, or the audit log sees it — a connector whose declaration is unsound is quarantined at startup rather than run. Then per-consumer lenses that can only ever remove: "analytics never receives email addresses" is one line, enforced on the way out. A lens never narrows the audit log. |
| Cost and loop control section 3.7 |
Every call and every scheduled poll is priced at its worst case against the target's budget and the deployment ceiling before it is sent. Ceilings exist that no grant can exceed — policy answers "may you," ceilings answer "should anyone, ever." |
| Data residency | A deployment ceiling plus a grant layer, with a cross-region resolution refused before any credential moves. Relevant for a global estate where EU guest data cannot be processed outside its region. |
What it changes in the loops
Sections 2 and 3 specify that intake, deduplication and theme classification must be deterministic code rather than model inference, precisely so those steps present no injection surface. That primitive has a name here: a reflex — a deterministic event-to-action rule that fires with no model in the loop, with its condition checked against the schema at boot, its own firing budget, and shadow mode as the default so the dangerous option is never the easy one. A reflex is a principal and takes the identical enforcement path an agent does, so it is not a bypass.
- A0, B0 and B1 — webhook verification, signature-hash deduplication, and closed-taxonomy theme classification — are reflexes rather than bespoke services.
- Loop C, the hourly funnel check against control limits, becomes a scheduled poll feeding a reflex. Durable cursors matter here: a poll cursor survives kill -9, re-delivers at most one event, and names a window it could not re-read rather than silently skipping it. A monitor that loses an hour without saying so is worse than no monitor.
- Shadow mode as default maps directly onto the phased rollout. Phase 4a can run the full reflex set in shadow — recording what each rule would have fired — which is a materially better Phase 4a than one that simply has the capability switched off.
What Marriott should probe before committing
- Connector coverage is the practical gating question. These loops need Dynatrace, Qualtrics, Splunk, GitHub and Snowflake alongside Fullstory. Confirm which connectors exist today versus which are scoped work, because a missing connector is either an integration to build or a system that sits outside the governed path — and a system outside the path defeats the purpose.
- Maturity. It reached v1 recently. Ask for production references, the results of any external hardening review, and what the upgrade cadence looks like.
- Who operates it. MIT licensing means Marriott can run it; it also means Marriott must run it. Decide whether that sits with platform engineering or security, and what the support path is when it refuses to start — which it will do by design rather than hold a credential it cannot rotate or audit.
- Secret-manager integration is a prerequisite, not a follow-on. Credentials arrive from a mounted file under a declared root, from platform workload identity, or from OAuth. That integration has to be settled before the Phase 4a pilot, not during it.
4.4 · What we would deliberately not buy yet
- A managed agent runtime as the foundation. The hosted agent services from all three clouds are credible and worth a look — several are framework-agnostic, so they are closer to runtimes than to frameworks. But they are optimized for conversational and single-agent shapes, their built-in session and memory models overlap awkwardly with durable execution plus the citation contract, and agent-native identity — the requirement that matters most here — is the least mature thing any of them ships. Pilot one. Do not found the program on one.
- An eval platform on the critical path. The golden-set scoring is plain code against Snowflake, where case files already land. Tracing UX is worth buying once the team feels the pain of not having it — not in the first sixty days, and not as a dependency of the pilot.
- A dedicated vector store. Nothing in either loop is a retrieval problem. Evidence arrives by ID from Dynatrace and Fullstory; quantification runs on governed SQL. If a vector database appears in a proposed architecture for this workload, ask what question it answers.
- GPU capacity. Managed in-region inference with customer-managed keys and no-retention terms covers it. Capacity planning for bursty incident traffic is a poor use of capital, and the self-hosted tier we do recommend — theme classification and signature dedupe — is small enough to run on CPU.
4.5 · Security and privacy posture — cloud-independent
None of this changes with the platform decision, so it can be ratified with security and architecture in parallel with the selection rather than after it.
| Control | Posture |
|---|---|
| Ingress | API gateway with mTLS and HMAC signature verification is the only inbound path. Dynatrace and Qualtrics webhooks land on a durable queue that feeds the workflow engine, so alert bursts are absorbed and replay is possible for forensics. |
| Egress | Host allowlist: Fullstory, Dynatrace, Qualtrics, GitHub Enterprise, the model endpoint. Everything else denied. This is the primary technical mitigation for prompt injection — an injected instruction has no reachable destination. |
| Agent identity & secrets | Agents hold capability references, not credentials. Identity by mutual TLS and SPIFFE; grants deny-by-default, reviewable and revocable; break-glass cancels calls already in flight. Separate GitHub App per agent role remains, but no agent context ever contains a token. This is what makes separation of duties real rather than aspirational — see 4.3b. |
| Fullstory access | Outbound-only, brokered, scoped per agent role. The evidence agent cannot compute metrics; the quantification agent cannot open sessions. Enforced as grants on a single path rather than by issuing each agent its own key — and a capability denied on the native connector cannot be obtained through the MCP server instead. |
| Guest data in context | Tokenized identifiers only. Loyalty tier as a cohort label, never as identity. Fullstory's private-by-default capture and element-level exclusion mean payment and personal fields are excluded at source, so the agent context does not contain them to begin with. |
| Model terms | In-region inference, customer-managed encryption keys, no training on Marriott data, no retention. Versions pinned per agent; a model change is a measured event against the golden regression set, not a silent upgrade. |
| Prompt & response logs | Classified sensitive, retained in Splunk under existing policy, access-controlled to the review board and security. These are the audit record. |
| Lower environments | Fullstory instrumentation in staging on a separate org, driven by synthetic traffic. Named prerequisite — Loop B's validation step cannot exist without it, and it is not in place today. |
4.6 · Why this should be the first agent Marriott builds
If the platform decision is open, the choice of first workload is a decision about what the platform has to prove. We would argue for this one, on four grounds — and then give the honest counterargument.
- It has a hard, pre-existing success criterion. Most first agents fail quietly because nobody can say whether they worked. This one is measured against 48 hours to containment and twelve already-ratified frictions. The verdict is arithmetic, not opinion.
- Phase 4a is read-only. No write capability for the first sixty days, so the blast radius during platform shakedown is close to zero. You are load-testing a runtime with an agent that cannot break anything.
- It exercises every hard capability the platform will ever need. Multi-week durable pauses, human gates, stateful tool leases, per-agent credential scoping, multimodal reasoning, write-capable tooling under blast-radius control, and an end-to-end audit trail. A conversational pilot exercises roughly one of those. Anything Marriott builds after this one inherits a runtime that has already been proven against the difficult case.
- It pays for the platform. Recovered revenue is the funding argument for everything downstream, and this is the only candidate first workload that produces a dollar figure as its native output.
Autonomy expands as a function of measured precision, never by calendar date. Each phase must clear its metrics on the golden regression set before the next unlocks.
0–60d
Human in every step. No code generation at all.
Complete the bridge attribute across booking-path services. Stand up the case file, the correlation ID chain, and the Splunk audit sink. Grant governed read-only MCP access. Build the golden regression set from the twelve quantified frictions and score a human-supervised agent against it. Loop A runs to the evidence package and stops there.
60–120d
Agentic investigation, human decision.
Both loops run end to end through evidence and quantification. Jira issues auto-file with full provenance. Loop C begins in observation mode. Still no pull requests. This phase is where the VoC pipeline the account has already scoped becomes continuous rather than episodic.
120–180d
Supervised remediation on an allowlisted scope.
Draft pull requests enabled on a small set of low-risk repositories and paths — front-end presentation defects first, the stay-date display class of bug, where client-side evidence is also the localization. Two human gates on every case. Every merge performed by a person.
180d+
Closed-loop validation, scope widened by evidence.
Lower-environment deployment and agentic validation go live for Loop B. The path allowlist widens one scope at a time, each expansion a board decision backed by measured precision on the current scope. Production deployment remains outside agent capability permanently.
The metrics that govern the program
| Metric | Why it governs |
|---|---|
| Time to containment | The headline. Current human baseline is 48 hours from identification to deployed temporary fix. This is the number the program is accountable to. |
| Root-cause precision | Agent-proposed cause versus human-ratified cause on the golden set and, later, on live cases. The primary trust metric. |
| Dollar-estimate error band | Agent figure versus the eventually ratified figure. Credibility with leadership depends more on this being tight and stable than on it being large. |
| PR acceptance rate | Agent-authored draft PRs merged without substantive rework. The gate on expanding the path allowlist. |
| Abstention rate | Cases the system declined to conclude. Watched for being too low, which indicates hallucination, as much as for being too high. |
| Provenance strip rate | Claims rejected by the citation validator. A rising rate is the earliest available signal of agent drift. |
| Tracking-cause share | Cases resolved as instrumentation gaps rather than product defects. Both a health signal for capture governance and proof the agents are not forcing behavioral explanations. |
| Escaped defect rate | Production issues traceable to an agent-authored change. Target is zero, and a single occurrence pauses the phase. |
Governance
We recommend a standing Agent Review Board with representation from Global Digital Technology, Digital Innovation, Experimentation & Analytics, capture and instrumentation governance, and enterprise security and architecture. It owns four decisions and only four: scope expansion, methodology change, phase advancement, and incident review on any agent-attributable defect. It meets on the existing cadence rather than creating a new one, and it is the natural first charter for the Center of Excellence the account audit already recommends.
Four approvals to begin Phase 4a
Nothing in the first sixty days generates code or touches a repository. The ask is to stand up the evidence chain and prove precision against incidents Marriott has already solved by hand.
Approve the 60-day Phase 4a pilot on Loop A, evidence-only
Dynatrace trigger through to a human-reviewed evidence package. No code generation in scope. Scored against the twelve already-quantified frictions.
Assign the two prerequisites with named owners
Bridge attribute coverage across booking and authentication path services, and Fullstory instrumentation in lower environments. Both are engineering tasks with clear completion criteria, and both gate program value.
Run a platform selection against the six requirements, not a vendor matrix
Two decisions need the architecture group’s time: the durable execution engine and whether to adopt an agent framework at all. The cloud-independent security posture can be ratified in parallel rather than after.
Charter the Agent Review Board and agree the unlock thresholds
Specifically: the precision figure and the dollar tolerance band that unlock Phase 4b, agreed in advance. Setting the bar before the first result is what keeps the program honest.
Open risks, stated plainly
- Behavioral evidence does not localize backend defects. The Ops Correlation Agent is load-bearing for anything server-side. Where the bridge attribute is missing, the system has no localization path and will decline the case. Loop A's code generation therefore starts on front-end defects by design, not by caution.
- MCP non-determinism. Mitigated by methodology persistence and drift checks, not eliminated. Published figures must be treated as reproducible only when accompanied by their persisted methodology.
- Instrumentation drift. Until data-attribute and selector standards are consistent, a share of cases will correctly resolve as tracking causes. This is the system working, and it needs to be framed that way to leadership before the first such case lands.
- Lower-environment instrumentation does not exist today. Loop B's closed validation step is contingent on it. Until then, Loop B ends at a human-approved Jira issue.
- Guest-supplied text is an attack surface. Mitigated by data-labelling, closed-taxonomy classification, egress allowlisting, and keeping the write-capable agent away from raw input. Worth an explicit security review rather than a footnote.