Marriott Agentic Architecture
Fullstory
MARRIOTT INTERNATIONAL
Business Strategy Memo  ·  Phase 4: Agentic Intelligence

From Signal to Shipped Fix,
Without Guessing

A resilient agentic architecture for autonomous issue detection, root cause analysis, and remediation — built so that every conclusion the system reaches is traceable to observed guest behavior, and every line of code it proposes passes through a human.

✓ Orchestrator + 4 scoped specialists 2 mandatory human gates per loop Baseline to beat: 48h to containment
To
Marriott Digital Leadership & Engineering VPs
From
Fullstory — Strategic Solutions
Re
Agentic architecture, hosting, and governance for the Dynatrace and Qualtrics remediation loops
Decision
Approve a 60-day Phase 4a pilot, two named prerequisites, and an Agent Review Board

Build a routing orchestrator over narrow specialists. Keep humans on the merge button.

Marriott has already proven the manual version of this loop. A Qualtrics verbatim about shifting stay dates became a confirmed behavioral finding, a 2.73-point conversion delta, a $200M annualized exposure, and a deployed containment in 48 hours. The work of Phase 4 is not to invent that loop. It is to compress it and run it continuously without lowering the evidentiary bar that made it credible.

Our recommendation is a central orchestrator routing to four tool-scoped specialist agents, with Fullstory MCP as the behavioral ground truth that constrains what the reasoning model is permitted to assert. Agents that gather evidence are architecturally forbidden from explaining it. Agents that explain are required to reconcile two independent sources. The agent that writes code never holds the credentials that produced the evidence.

Autonomy expands as a function of measured precision, not calendar. Phase 4a ships with a human in every step and no code generation at all — and it still beats the 48-hour baseline. Because no agentic runtime has been chosen yet, these loops also serve as the specification that should drive that decision; Section 4 treats it as a platform selection rather than a fit to existing infrastructure.

One constraint shapes everything below

Before the architecture: a limit worth stating plainly, because designing around it is the difference between a system leadership trusts and a system that gets switched off after its first confident wrong answer.

Behavioral evidence proves that a failure happened and to whom. It does not localize the defect to a line of code. A session replay can show a guest clicking Book Now three times against a "Correction needed" message, the DOM state at that instant, the failing network call, and the console error. That is decisive proof of impact. It is not proof of cause. Any agent that generates a pull request from session evidence alone will confidently invent a root cause.

The architectural consequence: server-side correlation is load-bearing, not supporting. The Dynatrace trace and the Splunk log are what localize the defect; Fullstory is what proves the defect mattered and sizes the money. Neither source alone authorizes a code change. This is why the recommended topology separates observation from inference, why the orchestrator requires two independent sources to agree before a case advances, and why autonomous PR generation begins on front-end presentation defects — the stay-date display class of bug, where the client-side evidence is the localization — and expands outward only as measured precision earns it.

Agentic topology

Recommendation: a central Triage Orchestrator routing to four tool-scoped specialist agents, strict hub-and-spoke, no peer-to-peer calls. Not a monolith, and not a swarm.

Control plane

Triage Orchestrator

Owns the case file and the state machine. Decides which specialist runs next, enforces the two-source reconciliation rule, enforces step and cost budgets, and posts to the human gates. Deliberately has no analytical tools and no write access to any system of record except the case store and the notification channel — it cannot investigate, and it cannot ship. Case records land in Snowflake as append-only rows — Snowflake already holds Marriott's behavioral clickstream from Phase 3, so no new warehouse is introduced; the case table is a purpose-built append-only ledger that becomes the longitudinal evaluation dataset for governing autonomy expansion.

case_store dispatch_agent notify (Teams / Slack) jira_create
Cannot: call Fullstory, Dynatrace, Splunk, or GitHub directly. Cannot assert a root cause. Cannot advance a case whose claims lack provenance tokens.
Specialist  ·  observe only

Session Evidence Agent

The sole holder of the Fullstory session lease. Opens the exact session and timestamp, reads DOM and accessibility state, captures the visible failure, diffs the interaction, closes the lease. Emits observations with timestamps — never causes.

session_open session_view session_diff session_screenshot session_get_a11y_tree session_close get_session_events
Cannot: use causal language. Cannot see the code repository. Cannot compute cohort metrics. System prompt rejects any output containing "because", "caused by", or a proposed fix.
Specialist  ·  quantify

Cohort Quantification Agent

Sizes the cohort and the money. Confirms what a funnel or segment actually measures before computing it, then applies Marriott's already-ratified formula: conversion delta between sessions with and without the friction, annualized, times AOV. Returns sample size with every figure.

build_segment get_segment build_funnel get_funnel compute_funnel compute_metric get_opportunities snowflake_query (parameterized)
Cannot: write free-form SQL. Cannot publish a dollar figure below the 100-session floor. Cannot reference an object ID that does not exist in the org.
Specialist  ·  localize

Ops Correlation Agent

Pulls the failing trace and the surrounding logs for the same correlation ID and time window, and produces the server-side failure signature that actually localizes the defect. This is the agent whose output authorizes engineering work.

dynatrace_problem_detail dynatrace_purepath splunk_search (saved)
Cannot: read Fullstory. Cannot write to any system. Cannot run unbounded Splunk queries — saved searches with bound time windows only.
Specialist  ·  propose code

Remediation Engineering Agent

Runs only after a human approves the evidence package. Works in an ephemeral sandboxed worktree with no production network egress. Must produce a regression test that fails before its change and passes after — a branch without one is rejected automatically. Opens a draft pull request and stops.

repo_read branch_create commit pr_open_draft test_run
Cannot: merge, force-push, or touch protected branches. Cannot modify CI config, infrastructure-as-code, secrets, database migrations, or auth and payment modules. Cannot see raw guest survey text. Cannot exceed the diff cap.

Why not one monolithic agent

Four of these reasons are standard enterprise architecture. The first is specific to this toolchain and is, on its own, decisive.

Why not a peer-to-peer multi-agent swarm

Closed-loop workflows

Two trigger loops and one always-on monitor. Gold rows are mandatory human checkpoints. Red rows are stop conditions where the system is designed to decline rather than proceed.

Loop A  ·  Infrastructure / Ops (Dynatrace-triggered)

Target: ops signal to approved evidence package in minutes; signal to merged fix within the same business day. The X-FullStory-URL request attribute on the failing PurePath is the bridge that makes the whole loop possible — and its absence is the loop's first stop condition.

StepOwnerAction & data inputTool callsHand-off
A0 Intake service
deterministic
Davis AI problem event arrives by webhook. Verify HMAC signature, extract problem ID, affected service, error signature, impacted request count, time window, and the X-FullStory-URL attribute. Hash the error signature and dedupe against open cases. webhook_verifysignature_hashdedupe New case, or increment an existing one
A0•STOP Intake service No session link, no agentic RCA. If the failing trace carries no Fullstory URL attribute, the case is routed to human triage with the ops context only. No model is invoked. At 4.7M visits a day a single popular failure can emit thousands of events — dedupe by signature means one case per signature per window, not one case per event. — Human triage queue
A1 Orchestrator Open an append-only case file. Assign case ID. Thread the correlation ID: Dynatrace problem → case → Jira key → PR → release. Set step, token, and wall-clock budgets. case_store.create Dispatch A2 and A3 in parallel
A2 Ops Correlation Retrieve the failing PurePath and the surrounding logs for the same correlation ID and window. Produce the server-side failure signature: service, method, status, exception class, first-seen, rate. dynatrace_problem_detaildynatrace_purepathsplunk_search Technical finding → case file
A3 Session Evidence Parse device and session IDs from the bridge URL. Availability check first: confirm the session is within the retention window and has finished indexing before opening the lease. If unavailable — due to indexing lag, retention expiry, or a missing bridge attribute — surface the session IDs to the human triage queue with an unavailability note and close the case branch; do not proceed to A4. If available: open the session at the event timestamp, capture the accessibility tree for DOM state, a screenshot of the visible failure, and a diff across the failing interaction. Close the lease. session_opensession_viewsession_get_a11y_treesession_screenshotsession_diffsession_close Observation record → case file
A4•STOP Orchestrator Reconciliation gate. Does the client-side observation corroborate the server-side signature inside the same time window? If the two do not agree, the case is marked unreconciled and escalated to a human. It does not proceed to quantification and it certainly does not proceed to code. reconcile Escalate, or advance to A5
A5 Cohort Quantification Build the segment of sessions exhibiting the signature. Read the checkout funnel definition, confirm what it measures, then compute with and without the signature. Apply the ratified formula. Return affected users, conversion delta, annualized exposure, and sample size. build_segmentget_funnelcompute_funnelcompute_metricsnowflake_query Quantified impact → case file
A6•HITL 1 Named engineering owner Evidence review. The complete package — technical signature, behavioral observations with session links and timestamps, screenshot, dollar exposure with sample size, and the methodology used — posts to Teams or Slack and files a Jira issue. Approve, reject, or request more evidence. Nothing reaches the repository before this. notifyjira_create Approval unlocks A7
A7 Remediation Engineering Clone to an ephemeral sandboxed worktree. Locate the presentation-layer defect — CSS state, DOM rendering, component logic — using the session observation and the reconciled server signature as the localization boundary. Write the minimal fix plus a regression test that reproduces the failing condition. Run the existing suite. Open a draft PR whose body carries the full evidence package and the agent-generated label. Phase 4a scope: front-end and presentation-layer defects only. Backend service defects (server-side exception, API behavior) produce the evidence package but require human engineering investigation to localize; automated PR generation for backend defects is Phase 4b, gated on measured precision from Phase 4a. repo_readbranch_createtest_runpr_open_draft Draft PR → CODEOWNERS
A8•HITL 2 CODEOWNERS reviewer Code review and merge. Normal review process, normal release train. The agent has no merge permission at the GitHub App level — this gate is enforced by platform configuration, not by instruction. — Merge → release
A9 Cohort Quantification Post-release verification. Re-compute the identical funnel and segment on the post-deploy window and re-check the Dynatrace problem state. Confirm resolution, or reopen the case automatically. compute_funnelcompute_metricdynatrace_problem_detail Close, or reopen at A1

Loop B  ·  Voice of Customer (Qualtrics-triggered)

This loop differs from Loop A in one structural way, and the difference matters more than anything else in it: a single negative survey is an anecdote. Loop B will not run root cause analysis on a sample of one.

StepOwnerAction & data inputTool callsHand-off
B0 Intake service
deterministic
Qualtrics response webhook on a CSAT or NPS score below threshold, or a negative verbatim. Extract the embedded Fullstory session link, survey metadata, brand, property, loyalty tier, and device. Treat the verbatim as untrusted data from this moment forward. webhook_verifyextract_session_link Theme classifier
B1 Classifier
small model
Map the verbatim to a fixed theme taxonomy — date handling, points redemption, sign-in, rate display, reservation change, payment. Closed vocabulary, not free-form generation. Output is a label and a confidence, nothing else. classify_theme Theme + confidence → cluster store
B2•STOP Cluster service Cluster threshold. A case opens only when at least K responses share a theme on the same page or flow inside a rolling window, or a single response correlates to a live Dynatrace problem. Everything else accumulates as aggregate VoC signal. Responses with no session link follow the aggregate-only path and never trigger RCA. Starting value: K = 5, 7-day rolling window. This is a calibration starting point derived from the 12-friction dataset, not a production commitment. The Agent Review Board ratifies the threshold before Loop B goes live and adjusts it after the first 90 days of production data; K is a governed parameter, not a hardcoded constant. cluster_eval Aggregate dashboard, or open a case
B3 Orchestrator Open the case. Select up to five representative sessions across the cluster — not one — spanning device, brand, and loyalty tier. Dispatch three independent investigations in parallel. case_store.createget_sessions Dispatch B4, B5, B6
B4 Classifier Subjective. Structured summary of what guests say is wrong, across the cluster: theme, frequency, affected tiers, verbatim exemplars quoted and attributed, never paraphrased into a cause. qualtrics_responses Subjective finding
B5 Ops Correlation Technical. Were there errors, latency excursions, or failing traces on the implicated flow during the cluster window? A negative result here is a finding, not a failure — it points the case toward a UX or content defect rather than an engineering one. dynatrace_problem_detailsplunk_search Technical finding
B6 Session Evidence Behavioral. For each representative session in the cluster: availability check first — confirm the session is within retention and fully indexed. Sessions from the cluster window may be days old; sessions approaching retention expiry or not yet indexed are noted with their IDs and routed to human triage rather than inspected. Walk available sessions in turn — lease held by one agent, one session at a time. Capture what the guest actually saw and did: DOM state, visible messaging, dead and rage interactions, the point of abandonment. A partially available cluster (some sessions unreachable) proceeds with the sessions that are available, provided at least two remain across distinct device and brand dimensions. session_opensession_viewsession_get_a11y_treesession_screenshotsession_close Observation records
B7•STOP Orchestrator Two-of-three triangulation. At least two of the three independent findings must agree on the same defect before the case is declared validated. One source alone — however vivid the verbatim — does not advance the case. reconcile Escalate, or advance to B8
B8 Cohort Quantification Size the exposure beyond the survey respondents — the respondents are a sample, and a biased one. Build the behavioral segment matching the validated defect across all traffic, compute the conversion delta, annualize against AOV. build_segmentcompute_funnelcompute_metricsnowflake_query Quantified impact
B9•HITL 1 CX + Product owner Validated issue review. Jira issue with the triangulated evidence, the cohort exposure, and the recommended disposition. This is the queue the account's VoC initiative already calls for: a visible, quantified pipeline of guest pain that product squads can prioritize. jira_createnotify Approval unlocks B10
B10 Remediation Engineering Fix plus regression test, deployed to a lower environment only. The agent receives the structured finding — never the raw guest text. Production deployment is not in this agent's capability set at any phase. branch_createtest_rundeploy_lowerpr_open_draft Lower-env build
B11 Session Evidence + Quantification Agentic validation in the lower environment. Synthetic traffic drives the repaired flow; the agent inspects the resulting sessions and confirms the defect no longer reproduces. Two named prerequisites, both required: (1) Fullstory instrumentation in staging on a separate org — not in place today; (2) a confirmed synthetic traffic mechanism — Dynatrace synthetic monitors if Marriott already operates them, or a scripted browser automation suite (Playwright, Selenium) that can be aimed at the lower environment. The mechanism must be identified and owned before Loop B can be validated end-to-end. This is a scoping item for the Phase 4a kickoff, not an assumption to be resolved at build time. get_sessionssession_opensession_viewsession_closecompute_funnel Validation verdict
B12•HITL 2 Release management Production release approval. Normal release train, normal change control. The validation verdict is an input to the human decision, not a substitute for it. — Production release
B13 Orchestrator Close on theme decay, not on merge. Monitoring continues; the case closes when the Qualtrics theme rate returns to baseline and the behavioral segment shrinks. A merged PR is not a resolved guest experience. cluster_evalcompute_metric Close, or reopen at B3

Loop C  ·  Always-on funnel vigilance

Loops A and B are both reactive — they wait for infrastructure to notice or a guest to complain. The highest-value failures are the silent ones that do neither.

A scheduled job recomputes the checkout funnels hourly against statistical control limits rather than fixed thresholds, so normal daily and weekly seasonality does not generate noise. A step-level drop outside the control band opens a case at step A1 of Loop A, bypassing the Dynatrace trigger entirely. This is the pattern that catches the failure class that never throws a server error and never produces a survey response: the rate card that renders empty, the date that silently reverts, the points balance that displays as unavailable. Early cases from this loop should run evidence-only, because without a server-side signature there is nothing to localize — which means Loop C feeds human investigation first, and earns code-generation rights last.

Resiliency & hallucination defense

Ten controls. The first three are the ones that matter most, and none of them is a prompt instruction — they are enforced by deterministic code between the agents.

The failure mode

One model gathers and explains

Root-cause hallucination has a specific origin. A single context window takes in partial evidence, feels pressure to produce an answer, and generates a plausible causal story that reads exactly like a real finding. It cites the session it looked at. It sounds confident. It is wrong in a way no reviewer can detect without redoing the work.

  • Confident causes from a single source
  • Dollar figures with no sample size
  • Object IDs that do not exist in the org
  • Behavioral explanations for tracking gaps
The control

Observation and inference are different agents

The Session Evidence Agent is permitted to say what it saw and when. It is architecturally prevented from saying why. Inference happens only in the orchestrator's reconciliation step, which cannot run on one source. The model never gets to hold an uncited claim across a hand-off.

  • Every claim carries a provenance token
  • Two independent sources or the case stops
  • Named-object allowlist, no invented IDs
  • Abstention is a first-class outcome

1  ·  The citation contract

Every assertion written into a case file must carry a provenance token: a session ID plus timestamp, a Dynatrace problem ID, a saved-search hash, or a metric ID plus the computed value. A deterministic validator sits between every agent hand-off and strips any claim without one. The stripped claim is logged — a rising strip rate is an early warning that an agent is drifting, and it is one of the metrics the review board watches.

// case-file assertion — rejected without provenance { "claim": "Guests saw 'Correction needed' when committing a 14-night reservation", "provenance": [ { "type": "session_observation", "session": "<id>", "ts": "20:52:14.330Z", "artifact": "a11y_tree + screenshot" }, { "type": "dynatrace_problem", "id": "<problem>", "signature": "POST /reservation/commit 500" } ], "inference_allowed": true // two independent sources present }

2  ·  How Fullstory MCP grounds the reasoning

"Use MCP as ground truth" is a slogan until it is four specific mechanisms. These are the four.

3  ·  Instrumentation fragility is a first-class outcome

The account's own configuration audit flags two gaps that bound agent reliability: inconsistent data attributes and CSS selector standards across the site, and iOS and Android SDKs not yet at full parity with web. Element-level agent reasoning inherits that fragility directly. An agent that cannot tell "the guest did not click it" from "we stopped capturing it" will report a behavior change that is actually a selector change.

The control is a pre-flight instrumentation health check on every case: are the elements and pages this case depends on being captured stably across the window under examination? If not, the agent returns a tracking cause, not a behavioral one, and routes to the capture-governance queue. This must be a rewarded outcome in the agent's evaluation, not a failed run. A system that can only ever find product defects will find product defects in instrumentation gaps.

4  ·  Confidence floors and abstention

5  ·  Destructive-action prevention

Capability-scoped at the platform layer. A prompt that says "do not merge" is a suggestion; a GitHub App without merge permission is a guarantee.

ControlEnforcement
No merge, everDedicated GitHub App with contents:write on unprotected branches and pull_requests:write only. No merge scope, no admin scope. Branch protection and CODEOWNERS enforced server-side.
Path denylistCI configuration, infrastructure-as-code, secrets, database migrations, and authentication and payment modules are blocked at the pre-commit hook. Touching them requires an explicit human-set flag on the case.
Diff capChanges above roughly 200 lines or 5 files are refused. Above the cap the agent files an analysis issue with its findings instead of a branch — a large diff is a design decision, not a hotfix.
Sandboxed executionEphemeral worktree, no production network egress, no production credentials in the environment, destroyed at case close.
Provenance on every PRagent-generated label plus the full evidence package in the body. No silent authorship anywhere in the repository.
IdempotencyEvery external write carries an idempotency key derived from the case ID and step. A retry cannot double-file a Jira issue or double-open a PR.
Circuit breakerRepeated failures on the same error signature trip the breaker and route to humans. Per-case step, token, and wall-clock budgets terminate runaway loops.
Session lease safetySingle-holder lease with TTL and a reaper process. Close precedes open, always. Only the Session Evidence Agent holds session tools at all.

6  ·  Prompt injection through guest-supplied text

Loop B ingests free text written by members of the public. A guest can type instructions into a survey comment box. Session DOM content can contain attacker-controlled strings. Any of it can reach a model context that also holds repository write capability.

7  ·  Human-in-the-loop checkpoints, stated plainly

GateOwnerWhat cannot happen before it
Evidence review (A6 / B9)Named engineering owner; CX and Product for Loop BNo repository access of any kind. No branch, no commit, no PR. The agent's investigation is complete and the human decides whether it is right.
Code review and merge (A8 / B12)CODEOWNERS; Release managementNo merge to any protected branch. No production deployment. Enforced by platform permissions, not policy.
Scope expansionAgent Review BoardNo new repository, path, or capability enters an agent's allowlist without a board decision backed by measured precision on the current scope.
Methodology changeExperimentation & Analytics ownerNo change to the quantification formula or the cohort definitions behind a published figure. Dollar figures must stay comparable quarter over quarter.
Production deployment is never an agent capability. Not in Phase 4a, and not at full maturity. The ceiling on autonomy here is a validated draft PR and a lower-environment deployment. Every production change goes through Marriott's existing release train, reviewed and merged by a person.

8  ·  Observability of the agents themselves

9  ·  The evaluation harness Marriott already owns

Twelve quantified frictions have been identified this year and eight have been found and fixed, each with a human-ratified root cause and a human-ratified dollar figure. That is a golden regression set. Before any agent is trusted with a live case, it replays those twelve incidents from their original signals and is scored on two things: did it reach the same root cause, and did it land inside a tolerance band of the ratified number. This converts "do we trust the agent" from a matter of opinion into a measurement, and it is the single highest-leverage thing to build in the first sixty days.

10  ·  Graceful degradation

Platform selection

No agentic runtime has been chosen yet. That is an advantage, not a gap — these two loops are a sharp enough specification to drive the decision, and they are a far better selection test than a generic pilot would be.

What follows is a decision framework, not a procurement recommendation. Vendor capabilities in this category are moving monthly, and the specifics below should be probed directly in vendor conversations rather than taken as settled. What is stable, and what should anchor the evaluation, is the set of requirements this workload places on a platform. Those come from the loops and will not change.

4.1  ·  Six requirements that should drive the decision

Most agent-platform evaluations start from vendor feature matrices and end in a tie. Starting from what these loops actually demand eliminates most of the field on the first pass.

RequirementWhere it comes fromWhat it rules out
Multi-week durable pauses Loop A pauses at the evidence gate for hours or days. Loop B pauses at issue review, lower-environment deploy, validation, and a release train — then closes on theme decay, which is weeks. Case state must outlive any process, any redeploy, and any model call. Message queues as the control plane. State held in a pod or a function. Most agent frameworks' built-in persistence, which is designed for conversation memory rather than multi-week business process state.
Stateful single-holder tool leases Fullstory's session inspection tools are opened, used, and must be closed before the next opens. Exactly one agent may hold the lease, with a TTL and a reaper. Naive parallel fan-out over a shared tool registry. Any runtime that cannot express a mutex or a singleton activity.
Per-agent credential scoping The whole separation-of-duties argument is real only if the credentials are actually separate: the evidence agent's token cannot compute metrics, the remediation agent's token reaches GitHub and nothing else. A single shared service account across agents. Agents inheriting a human's permissions. This is the requirement the cloud platforms are weakest on — and the one we recommend solving with a dedicated control plane rather than with a cloud primitive. See 4.3b.
Multimodal reasoning plus strong code generation Session screenshots and accessibility trees are visual evidence. Loop A's value ceiling is set by code-generation quality. Open-weight-only model tiers at realistic internal scale. Text-only endpoints.
Minutes-long jobs with real disk The remediation agent clones a repository and runs a test suite. The evidence agent walks up to five sessions in sequence. A functions-only architecture. Short execution ceilings and ephemeral-only storage.
Immutable audit trail, signal to release One correlation ID from Dynatrace problem through case, Jira key, PR number, and release. Required for risk and audit, and it is the dataset that governs autonomy expansion. Any runtime that will not export structured per-step traces. Black-box managed services whose internal state you cannot query.

4.1b  ·  Three compute tiers — why each exists

These loops require three distinct layers of compute. Conflating them is the most common platform selection error.

The table below maps each layer to the cloud-specific implementation options. Marriott's cloud decision will be driven by commercial terms and existing operational muscle — the capability requirements above are cloud-neutral and should anchor the evaluation regardless of which path is chosen.

4.2  ·  Candidate stacks

The cloud decision will most likely be made on grounds that have nothing to do with agents — existing gravity, commercial terms, where enterprise architecture already has operational muscle. Marriott's current estate gives weak and conflicting signals: Snowflake is cloud-neutral, and the Gemini reference on the agentic platform slide hints Google without establishing it. So rather than pick a cloud, here is what each path looks like and where each one is weak for this workload.

LayerAWS-anchoredGoogle-anchoredAzure-anchoredCloud-neutral option
Durable execution Step Functions, or Temporal on AWS Cloud Workflows, or Temporal on GCP Durable Functions / Durable Task, or Temporal on Azure Temporal (Cloud or self-hosted). Also Restate, Inngest.
Agent compute ECS Fargate or Bedrock AgentCore runtime Cloud Run jobs (recommended); GKE is a viable option if Marriott already operates Kubernetes at scale — the additional cluster overhead is not justified for this workload alone Azure Container Apps jobs Containers on any managed scale-to-zero job service
Model access Bedrock — broad multi-model including strong code-generation options Vertex AI — Gemini-native, matches the stated reasoning layer Azure AI Foundry — widest catalog, fastest model turnover Provider-agnostic gateway in front of one or two endpoints
Agent identity IAM roles — not bound to workforce identity today Service accounts — integration work required Entra ID — furthest along on agent-native identity and conditional access Sekizui — agents hold certificates (mTLS, SPIFFE), never secrets; grants are deny-by-default and revocable
Tool governance No cloud ships this as a complete answer yet — see 4.3b Sekizui — policy enforcement point and credential broker; MCP tool lists pinned under version control
Grounding store Snowflake in all three cases — already landing clickstream from Phase 3 Snowflake, with reviewed parameterized queries
Audit & APM Splunk and Dynatrace in all three cases — no reason to change either Splunk, Dynatrace
One cloud-native caveat worth surfacing early. Vertex's native grounding advantage is against BigQuery. Marriott's behavioral warehouse is Snowflake, so that particular benefit does not transfer — the quantification path runs through reviewed Snowflake queries regardless of cloud. Weigh the Google path on Gemini and operational fit, not on native grounding.

4.3  ·  The two decisions that actually require deliberation

Most of the table above resolves itself once a cloud is picked. Two choices do not, they are genuinely portable, and they carry the most consequence. These are where we would spend the architecture group's time.

Decision 1 — the durable execution engine

✓

Recommendation: Temporal

Multi-month workflow state and human-approval signals are its core competency, not an extension of it. Workflows are defined in code, which matters because these loops branch on evidence quality rather than following a linear path.

⋮

It survives the cloud decision

Since no cloud is chosen, a portable control plane means this choice is not re-litigated later. Namespace isolation also maps cleanly onto per-specialist credential boundaries.

⚠

The honest tradeoff

Operational burden if self-hosted — Temporal Cloud removes most of it at a cost. The cloud-native engines are cheaper and simpler if Marriott is confident in a single cloud and comfortable expressing this branching in a state-machine DSL. Step Functions is genuinely excellent; it is also AWS-bound.

Decision 2 — framework or thin orchestration

✓

Recommendation: thin, with a framework only inside single steps

These loops are deterministic workflows with model calls inside them — not autonomous agent loops. The orchestrator is an explicit state machine. That is workflow code, and an agent framework is the wrong tool for it.

✗

Two sources of truth is the failure mode

Most frameworks bring their own state and persistence. Layered over a durable execution engine, you get two competing records of where a case stands — and case state is the thing auditors will ask about.

⋮

Where a framework does earn its place

Inside an individual reasoning step — the session evidence extraction, the classifier — running as one activity with scoped tools. Framework choice then becomes a reversible, step-local decision rather than a platform commitment.

4.3b  ·  The agent-to-system control plane

Requirement three in 4.1 — per-agent credential scoping — is the one no cloud platform answers completely, and it is the requirement the entire separation-of-duties argument rests on. We recommend solving it with a dedicated control plane rather than waiting for a cloud primitive to mature.

Recommendation: Sekizui (github.com/fullstorydev/sekizui) — a policy enforcement point and credential broker sitting between the agent fleet and every system it acts on. Stated plainly: this is Fullstory open source, MIT licensed. We are recommending a component we wrote, and the license is why that should be acceptable rather than suspect — Marriott can read it, run it, fork it, and keep running it with no dependence on any commercial relationship with us. Evaluate it on the architecture, not on who authored it.

The obvious benefit is arithmetic: five agents needing six systems is thirty integrations, each with its own credential, audit gap, rate limit and error handling. A broker makes that N + M instead of N × M. That collapse is worth having but it is not the reason to adopt it. The reason is that it turns three controls this proposal describes as design intent into properties the infrastructure actually enforces.

What this memo specifiedWhat the control plane enforces instead of trusting
Per-agent credential scoping
requirement 3
Agents hold certificates — mutual TLS with SPIFFE identities — and never secrets. Grants are deny-by-default, written and reviewed as action × target × constraints, and revocable. A compromised agent can do exactly what it was granted and nothing more. Break-glass refuses new work, cancels calls already in flight, and records how many it killed.
Scoped tool allowlists per agent MCP servers sit behind a vetted contract: the vendor's tool list is pinned under version control and compared against what the server actually serves. Unvetted tools are unreachable, unvetted arguments are refused, results that break their declared shape are refused, and drift is recorded. Critically — a capability denied on a native connector cannot be obtained by going through that vendor's MCP server instead.
Immutable audit trail
requirement 6
One hash-chained log that records decisions, not just actions — including refusals, which rule matched, who asked on whose behalf, and what it cost. "What did this agent do to production this week" becomes a query rather than an investigation. This is the record the Agent Review Board governs autonomy expansion from.
Prompt-injection containment
section 3.6
Loop B ingests text written by the public. An agent holding a GitHub token is one successful injection away from an attacker holding that token. With credentials brokered, the blast radius of a successful injection is the grant set, not the credential set. Egress allowlisting and data-labelling remain; this sits beneath them.
Guest-data minimisation Two mechanisms. Connectors declare field by field what they may return, and anything undeclared is stripped before a consumer, a rule, or the audit log sees it — a connector whose declaration is unsound is quarantined at startup rather than run. Then per-consumer lenses that can only ever remove: "analytics never receives email addresses" is one line, enforced on the way out. A lens never narrows the audit log.
Cost and loop control
section 3.7
Every call and every scheduled poll is priced at its worst case against the target's budget and the deployment ceiling before it is sent. Ceilings exist that no grant can exceed — policy answers "may you," ceilings answer "should anyone, ever."
Data residency A deployment ceiling plus a grant layer, with a cross-region resolution refused before any credential moves. Relevant for a global estate where EU guest data cannot be processed outside its region.

What it changes in the loops

Sections 2 and 3 specify that intake, deduplication and theme classification must be deterministic code rather than model inference, precisely so those steps present no injection surface. That primitive has a name here: a reflex — a deterministic event-to-action rule that fires with no model in the loop, with its condition checked against the schema at boot, its own firing budget, and shadow mode as the default so the dangerous option is never the easy one. A reflex is a principal and takes the identical enforcement path an agent does, so it is not a bypass.

The honest cost, in the project's own words: centralising credentials makes this control plane the highest-value target in the system. The trade is that one small, hardened, audited surface is more defensible than secrets distributed across a fleet of prompt-injectable agents. That is the right trade for this workload, but it is a real concentration of risk and it should be reviewed as such by enterprise security — not waved through because the component is well-designed.

What Marriott should probe before committing

4.4  ·  What we would deliberately not buy yet

4.5  ·  Security and privacy posture — cloud-independent

None of this changes with the platform decision, so it can be ratified with security and architecture in parallel with the selection rather than after it.

ControlPosture
IngressAPI gateway with mTLS and HMAC signature verification is the only inbound path. Dynatrace and Qualtrics webhooks land on a durable queue that feeds the workflow engine, so alert bursts are absorbed and replay is possible for forensics.
EgressHost allowlist: Fullstory, Dynatrace, Qualtrics, GitHub Enterprise, the model endpoint. Everything else denied. This is the primary technical mitigation for prompt injection — an injected instruction has no reachable destination.
Agent identity & secretsAgents hold capability references, not credentials. Identity by mutual TLS and SPIFFE; grants deny-by-default, reviewable and revocable; break-glass cancels calls already in flight. Separate GitHub App per agent role remains, but no agent context ever contains a token. This is what makes separation of duties real rather than aspirational — see 4.3b.
Fullstory accessOutbound-only, brokered, scoped per agent role. The evidence agent cannot compute metrics; the quantification agent cannot open sessions. Enforced as grants on a single path rather than by issuing each agent its own key — and a capability denied on the native connector cannot be obtained through the MCP server instead.
Guest data in contextTokenized identifiers only. Loyalty tier as a cohort label, never as identity. Fullstory's private-by-default capture and element-level exclusion mean payment and personal fields are excluded at source, so the agent context does not contain them to begin with.
Model termsIn-region inference, customer-managed encryption keys, no training on Marriott data, no retention. Versions pinned per agent; a model change is a measured event against the golden regression set, not a silent upgrade.
Prompt & response logsClassified sensitive, retained in Splunk under existing policy, access-controlled to the review board and security. These are the audit record.
Lower environmentsFullstory instrumentation in staging on a separate org, driven by synthetic traffic. Named prerequisite — Loop B's validation step cannot exist without it, and it is not in place today.

4.6  ·  Why this should be the first agent Marriott builds

If the platform decision is open, the choice of first workload is a decision about what the platform has to prove. We would argue for this one, on four grounds — and then give the honest counterargument.

The counterargument, stated fairly: this is not the easiest first agent. A read-only question-answering agent over Snowflake would ship faster, with less governance overhead and fewer prerequisites. If the goal is a quick visible win, build that one. If the goal is to make a platform decision that holds — and to find out in sixty days rather than eighteen months whether the chosen runtime can carry real agentic work — this is the better first build. We think the second goal is the one worth optimizing for, and the read-only Phase 4a scope is what makes the risk of it acceptable.
Two prerequisites carry the whole program regardless of platform, and both are concrete and assignable. First, the X-FullStory-URL request attribute must be present on every service in the booking and authentication paths — where it is missing, Loop A is blind, and the system will correctly refuse to run. Second, lower-environment Fullstory instrumentation, without which Loop B's closed validation step is aspiration rather than architecture. Neither depends on the runtime decision, so both can start now.
Phasing, governance & decision asks

Autonomy expands as a function of measured precision, never by calendar date. Each phase must clear its metrics on the golden regression set before the next unlocks.

4a
0–60d

Human in every step. No code generation at all.

Complete the bridge attribute across booking-path services. Stand up the case file, the correlation ID chain, and the Splunk audit sink. Grant governed read-only MCP access. Build the golden regression set from the twelve quantified frictions and score a human-supervised agent against it. Loop A runs to the evidence package and stops there.

Unlock criteria: root-cause precision on the regression set, dollar estimates inside the agreed tolerance band, and a measured time-to-evidence materially under the 48-hour human containment baseline.
4b
60–120d

Agentic investigation, human decision.

Both loops run end to end through evidence and quantification. Jira issues auto-file with full provenance. Loop C begins in observation mode. Still no pull requests. This phase is where the VoC pipeline the account has already scoped becomes continuous rather than episodic.

Unlock criteria: stable precision across both loops, abstention rate inside the healthy band, and zero unreconciled cases advanced in error.
4c
120–180d

Supervised remediation on an allowlisted scope.

Draft pull requests enabled on a small set of low-risk repositories and paths — front-end presentation defects first, the stay-date display class of bug, where client-side evidence is also the localization. Two human gates on every case. Every merge performed by a person.

Unlock criteria: PR acceptance rate above the agreed threshold, zero escaped defects attributable to an agent-authored change, and no denylist violations.
4d
180d+

Closed-loop validation, scope widened by evidence.

Lower-environment deployment and agentic validation go live for Loop B. The path allowlist widens one scope at a time, each expansion a board decision backed by measured precision on the current scope. Production deployment remains outside agent capability permanently.

Steady state: recovered revenue attributed to agent-originated cases becomes a reported line in the quarterly value readout.

The metrics that govern the program

MetricWhy it governs
Time to containmentThe headline. Current human baseline is 48 hours from identification to deployed temporary fix. This is the number the program is accountable to.
Root-cause precisionAgent-proposed cause versus human-ratified cause on the golden set and, later, on live cases. The primary trust metric.
Dollar-estimate error bandAgent figure versus the eventually ratified figure. Credibility with leadership depends more on this being tight and stable than on it being large.
PR acceptance rateAgent-authored draft PRs merged without substantive rework. The gate on expanding the path allowlist.
Abstention rateCases the system declined to conclude. Watched for being too low, which indicates hallucination, as much as for being too high.
Provenance strip rateClaims rejected by the citation validator. A rising rate is the earliest available signal of agent drift.
Tracking-cause shareCases resolved as instrumentation gaps rather than product defects. Both a health signal for capture governance and proof the agents are not forcing behavioral explanations.
Escaped defect rateProduction issues traceable to an agent-authored change. Target is zero, and a single occurrence pauses the phase.

Governance

We recommend a standing Agent Review Board with representation from Global Digital Technology, Digital Innovation, Experimentation & Analytics, capture and instrumentation governance, and enterprise security and architecture. It owns four decisions and only four: scope expansion, methodology change, phase advancement, and incident review on any agent-attributable defect. It meets on the existing cadence rather than creating a new one, and it is the natural first charter for the Center of Excellence the account audit already recommends.

Four approvals to begin Phase 4a

Nothing in the first sixty days generates code or touches a repository. The ask is to stand up the evidence chain and prove precision against incidents Marriott has already solved by hand.

Ask 1

Approve the 60-day Phase 4a pilot on Loop A, evidence-only

Dynatrace trigger through to a human-reviewed evidence package. No code generation in scope. Scored against the twelve already-quantified frictions.

Ask 2

Assign the two prerequisites with named owners

Bridge attribute coverage across booking and authentication path services, and Fullstory instrumentation in lower environments. Both are engineering tasks with clear completion criteria, and both gate program value.

Ask 3

Run a platform selection against the six requirements, not a vendor matrix

Two decisions need the architecture group’s time: the durable execution engine and whether to adopt an agent framework at all. The cloud-independent security posture can be ratified in parallel rather than after.

Ask 4

Charter the Agent Review Board and agree the unlock thresholds

Specifically: the precision figure and the dollar tolerance band that unlock Phase 4b, agreed in advance. Setting the bar before the first result is what keeps the program honest.

Open risks, stated plainly