Using vector embeddings across millions of user sessions to derive multi-dimensional behavioral clarity, cluster users into archetypes, and power descriptive and prescriptive enterprise data strategies.
Dashboards aggregate. Funnels measure steps. Session tools replay moments. None of them tell you which type of user you're talking to โ or how to treat them differently across every downstream system.
A "12% abandonment" number is an average across deal-seekers, loyalists, first-timers, and frustrated power users โ each requiring a different response.
CRM segments on what users bought. Behavioral archetypes reveal how they navigate, where they hesitate, and what drives them toward or away from conversion.
One session is a snapshot. Behavioral shape emerges across dozens of sessions โ the pattern of return, explore, hesitate, convert, and lapse is what defines an archetype.
CDP outreach, CRM personalization, AI chat agents, and fraud models all make decisions without knowing the behavioral archetype of the user they're acting on.
Fullcapture records the complete behavioral stream โ not a 5% sample, not predefined events only. That completeness is what makes cross-session archetype clustering possible. You can only cluster on what you captured.
fs-skills is a JavaScript SDK toolkit that decorates Fullstory events with business-meaningful context at capture time โ before the event hits the data pipeline. Named elements are the foundation: they translate opaque CSS selectors into human-readable action labels that cluster analysis can reason about.
Raw selector [data-id="atc-btn"] becomes click:Add_to_Order_Button โ a token with conversion semantics the embedding model can cluster on.
URL patterns mapped to page names create navigation tokens like page:Deals โ page:Order_Menu โ page:Cart โ readable funnel trajectories.
Business events like custom:order_completed with order value, coupon code, and store ID become the terminal signal every cluster analysis centers on.
Session-level context โ is_logged_in, scroll_depth_pct, has_store_id โ feeds the behavioral feature matrix alongside the journey embedding.
Stale named element selectors (post-app-rewrite) show zero click counts on critical buttons. The pipeline surfaces these gaps, routing fixes back to the Fullstory UI โ documented in SELECTOR_FIXES.md.
Client-side injection requires engineers to add decoration to the DOM or view tree โ Fullstory captures it as semantic labels. Server-side extraction works the other direction: Fullstory reads network traffic, URL patterns, and element attributes and extracts structured semantic variables automatically, via rules configured in the Fullstory UI or API. No frontend changes required.
JSON fields from XHR/fetch responses โ order totals, cart contents, user tier, product IDs โ surfaced as semantic event properties without any frontend code changes.
Structured data parsed from URL paths โ /store/[storeId]/menu โ storeId attached as a session variable on every event in that page.
Search queries, filter selections, and configuration parameters from POST bodies โ capturing intent before the user takes a visible action.
data-* attributes and ARIA labels already present in the DOM โ product SKUs, price tiers, promo slot IDs โ mapped onto interaction events without re-instrumentation.
Extracted variables arrive in the NDJSON export pre-attached to events as first-class fields. The pipeline reads them directly โ no regex parsing, no post-hoc joining.
When a Fullstory customer names an element "Add to Order Button | Web," builds a funnel that tracks it, and reviews that funnel dashboard every morning โ that is a declared priority embedded in their Fullstory configuration. The pipeline reads that signal back and uses it to calibrate the semantic weight tier automatically, replacing hand-coded constants with customer-derived evidence.
Any named element or custom event referenced as a funnel step is, by definition, on the customer's critical path. These get the highest weight: 1.5 โ ร3 repetitions.
Metrics that track clicks or events on specific named elements โ and are viewed frequently โ indicate the team watches those elements as business KPIs. Weight: 1.3 โ ร2 repetitions.
Custom segments defined by behavior on specific pages or elements signal that the customer segments their user base by those interactions โ high behavioral relevance. Weight: 1.2 โ ร2 repetitions.
Descriptive, business-intentful names ("Apply Promo Code," "Guest Checkout Button") indicate intentional instrumentation. Generic or selector-derived names indicate lower signal confidence. Baseline: 1.0 โ ร1 repetition.
Without named elements, conversion funnels, and high-signal metrics in place, the weight map cannot be derived from customer intent โ it degrades to guessing. Clusters will still form, but they will reflect raw DOM structure and URL shape, not business behavior. The analysis becomes a technical exercise rather than an actionable one. Org tuning is the prerequisite work, not a nice-to-have enhancement.
Client-side injection (slide 03) creates named elements via DOM decoration. Server-side extraction (slide 04) enriches events with structured business variables from network traffic. Weight inference calibrates how loudly each element speaks in the journey string โ using the customer's own instrumentation choices as the authority on importance.
discover_org_context tool for spot-checking. Write all three files to semantic/ before running Phase 4.element_name. Named page URL matching โ page_name. Token prefix assignment.Each user's 30-day session history is flattened into a weighted token sequence โ a "behavioral sentence" the embedding model reads the same way it reads text. Word frequency signals importance; we exploit this by repeating high-signal tokens.
Cap: 120 tokens per user (most recent events). Noise events excluded: identify, log, consent, performance.
all-MiniLM-L6-v2 is a general-purpose text model. It learned from text that word frequency signals importance. Repeating click:add_to_order ร3 in the journey string makes the model treat it as the dominant behavioral signal โ without retraining the model.
Two algorithms in sequence โ one to make the geometry tractable, one to find the clusters.
Users are positioned by behavioral similarity โ the shape of their navigation across dozens of sessions. Clusters share behavioral DNA. Distance is meaningful.
NDJSON export (API v2) for internal POC or partner demos. Warehouse Ready-to-Analyze Views or Raw Storage Bucket for customer production pipelines.
Identify the user population: buyers, signed-in, high-session. The segment defines the behavioral space โ narrow it around a conversion goal.
Capture from the Fullstory UI. Ensure conversion-critical actions โ add to cart, place order, apply promo โ have clean, current selectors. Fix stale ones first.
Sample 500K events to confirm timestamp format, field sparsity, and event type distribution. Applies regardless of how the data arrives.
Phases 1โ6: resolve โ feature engineer โ embed โ cluster. DuckDB + sentence-transformers + HDBSCAN. ~2 hours on 57M events, 32GB RAM.
5 centroid users + aggregate cluster metrics โ 2โ4 word archetype name + one-sentence description. ~$0.80 for 103 clusters.
Export archetype labels to CDP, CRM, or AI agent system prompt. Behavioral archetype is now a first-class field in every enterprise data system.
A named, described, and quantified library of behavioral archetypes for your customer base โ derived from real session data, not survey responses or demographic proxies. Each archetype comes with behavioral flags, aggregate statistics, and centroid session replays you can watch in Fullstory.
behavioral-clustering/CLAUDE.mdpapajohns/ (103 clusters, 100K users)