D2C "- Web" Voice Agents · Qualitative + Quantitative

User Segmentation — First-Time vs. Returning Callers

35-agent fleet · 30-day window (core segmentation) · extended to 41-day / 36-agent fleet for the Token-Warning tab

35

Active "-Web" Agents

1,745

Conversations Analyzed

1,081

Distinct Callers Identified

72%

Identity-Match Rate

Executive Summary

What distinguishes first-time from returning callers, what gets a first-timer to engagement, what returners expect, and where the data can and can't support a "high value vs. low value user" segmentation yet.

Returners run deeper sessions

1.5x longer

116.5s avg vs. 77s for first-timers; 5x more likely to hit a 3+ minute call (20.6% vs. 4.2%).

The real bail predictor

36% get zero reply

Over a third of quick-bail first-time calls end before the user says anything at all — a top-of-funnel issue, not a script issue.

Continuity is expected, not delivered

41% expect memory

Returning callers reference specific prior details constantly — but 0 of 3,811 checked turns showed real retrieval. It's improvisation.

Speaking is the strongest value signal

86% of whales talk

Of users with 5+ lifetime calls, the overwhelming majority are speaking-dominant, not emoji- or text-only.

Hitting zero tokens is a churn event, not a pause

Only 1.1% return

Of 1,481 identity-matched hard-token-stop instances (41-day window), just 16 show any later call — and at least 5 of those 16 are the same single repeat user. See the new "Token-Warning & Retention" tab.

What this report covers

  1. What distinguishes first-time callers from returning callers in the transcript itself
  2. What gets a first-timer to an engagement "aha" in the first ~30 seconds
  3. What returning callers expect ("pick up where we left off") and whether the product delivers it
  4. A finer segmentation within each group — intent tier × communication modality — and how that maps to retention
  5. A "Whale/Returner Archetype" — voice behavior, call themes, and concrete retention/conversion recommendations for this population
  6. What happens after a user hits zero tokens: which warning messages actually fire, and whether users return afterward
  7. A prioritized Product Roadmap — RICE-scored bets, phased Now/Next/Later ownership across Design/FE/BE/AI-ML
  8. A Cross-Functional Research Agenda — gaps a prompt engineer, data scientist, UX researcher, consumer researcher, and voice-AI engineer would each flag before trusting this report's conclusions further
  9. An honest accounting of where real activation/retention/monetization data does and doesn't support a value-tier segmentation yet

North Star & Top 3 Ranked Bets

Proposed North Star: % of hard-token-stop callers who return within 7 days. Aspirational — properly measuring this requires the cohort-tracking infrastructure named as a Now-tier item in the Research Agenda tab; today we can only measure "returned at all, ever, in-window," which is the weaker proxy behind the 1.1% figure below.

  1. Diagnose why the existing re-up prompt + checkout flow only converts 1.1% of hard-stopped callers — a verbal prompt (Stage 1, D2C-5403) and an in-call banner/checkout already exist, so the gap is in the flow itself, not its absence. Join failed-payment data to find out. See Token-Warning & Product Roadmap.
  2. Confirm the two legacy token-warning messages are fully decommissioned fleet-wide now that D2C-5403's 2-stage system is standard — cheap verification, not a live bug. See Product Roadmap + Research Agenda.
  3. Investigate the 36%-zero-reply cohort for a technical failure (STT/VAD/latency) before continuing to treat it as pure user disengagement — cheap to check, data access confirmed available. See Research Agenda.
Scope note: this is a single 30-day snapshot (fleet-wide, all 35 active "- Web" agents, confirmed complete via an agent-discovery audit — see Data Quality tab). Treat findings as directional and re-verify on a rolling basis, not as a one-time-final answer.

First-Time vs. Returning Behavior

1. What distinguishes them

First-time (n=1,020)Returning (n=266)
Avg. call duration77s116.5s
Calls ≥3 minutes4.2%20.6%
Avg. turns12.715.1
First message length17.4 chars31.4 chars

The clearest tell is in the first user message. Returners open by naming the persona directly and treating it as an ongoing relationship — "Hello, Lisa... you look really sexy today," "Hi, Kylie. How are you?" First-timers open generic or reactive — "Good.", "...", "Keep going." Note: name-personalization in the agent's own greeting is not itself a returning-user signal — first-timers get named greetings too, from account data.

2. Getting a first-timer to an "aha" in 30 seconds

Every agent opens with an identical scripted line — "[giggle] Hey, babe! How's it going?" — regardless of persona, so the opener itself isn't the differentiator. What matters is whether the user's first reply within ~35 seconds is substantive or vague/absent:

First-timer outcomeSubstantive first replyVague replyNo reply at all
Bailed (<30s, n=229)40.2%23.6%36.2%
Engaged (≥120s, n=225)71.6%28.0%0.4%

Where the user does engage, the agent escalates fast and specifically off whatever the user gave it — an emoji, an action, a mood — and the "aha" is that reciprocal specificity, not any single scripted beat. Vague replies ("Good.", "...") get a generic re-engagement prompt that rarely revives the call.

3. What returners expect — and a real product gap

41% of returning-call transcripts (103/252) contain explicit continuity language. Users expect the persona to remember specific prior details, not just "we've talked before" in the abstract:

"Do you still have the Polaroid photos from last time?"
"Remember the last time I fucked you?"
"You have to remember, we just talked about it five seconds ago." — frustration when continuity breaks
"...we were just texting together and you said if I call you, you'll say something to me first" — expectation spans chat and voice as one relationship
Finding: I checked rag_retrieval_info and contextual_update_info across all 3,811 turns in the returning-caller sample — zero were populated. The agent's "I remember..." lines (and there are plenty — "call me back when you're ready for me to pick up exactly where we left off") are LLM improvisation, not backed by retrieval of real prior-call content. It currently works by coincidence until a user's specific expectation doesn't match the improvisation, which produces visible, logged frustration.

Deeper Segmentation — Intent × Modality

Segmentation is per user (all their calls in the window aggregated), not per call, and a modality label is only assigned when it's genuinely dominant (≥60% of that person's turns) — otherwise the user is labeled Blended. 303 of 1,081 users (28%) are genuinely blended.

Segmentn users% ever returned% ever reached "engaged" tier (≥120s)
Speaking-dominant40121.4%42.1%
Blended30313.5%20.8%
Emoji-dominant18811.7%14.4%
Silent-dominant915.5%17.6%
Preset-dominant1118.2% (small n)0.0%
No turns at all819.9%0.0%

Speaking-dominant users are ~2x more likely to return and 2-3x more likely to ever hit a deep session than any other modality. Modality isn't a separate axis from intent — it functions as a leading indicator of it.

⭐ Elevated finding — the power-user modality pattern

Of the 22 users with 5+ lifetime calls, 19 (86%) are speaking-dominant, 3 are blended, and zero are emoji- or silent-dominant. This is the strongest single signal in the whole segmentation about which modalities actually convert into long-term value — three implications follow, each checked against the data rather than left as assumption:

  1. Emoji is not a growth lever — settled, not just observed. Zero of 22 power users are emoji-dominant, and emoji-dominant is the second-worst segment on both return (11.7%) and engagement (14.4%). A gamified "emoji fills a creator's heat meter" mechanic was considered and correctly rejected: tying a reward to emoji volume would incentivize spamming emojis over actually talking, directly undermining the modality that's proven to matter. Recommendation: stop treating emoji-parity/expansion as a roadmap item.
  2. Preset usage does not appear to scaffold users toward speaking — if anything, the opposite. Investigated whether higher preset-usage share for blended users (n=303) correlates with migrating toward speaking-dominant behavior over time. Among the 34 blended users with ≥2 calls captured in this 30-day window (all confirmed via call_rank to be observed from their true first-ever call): users whose call history is ≥34% preset-dominant reach a later speaking call only 30.8% of the time (4/13) vs. 61.9% (13/21) for lower-preset-share users. Preset-first callers migrate to speaking at about the same or slightly lower rate (50%, 4/8) as emoji/silent-first callers (61%, 11/18) — preset doesn't look like a stronger bridge to voice than any other non-speaking modality. The one real trend in the data — spoke-rate jumping from 23.5% on a user's first captured call to ~40-44% on later calls — looks like a general "settling in after call 1" effect present regardless of starting modality, not something preset usage specifically drives. Caveat: n=34, sub-buckets as small as 8-13 — directional, not conclusive; see Research Agenda for the full-scale follow-up recommendation.
  3. Silent-dominant users are very likely reaching JOI (Jerk-Off-Instruction) content, and the hand-off works — but most sessions don't resolve. JOI is a real, named, monetized feature, and assessing exactly this — "confirm silent-caller JOI flow usage and quality (Stage 1A/2A/3A narration path)" — is a live acceptance criterion on D2C-5409, feeding D2C-5410's procedures pilot. A transcript content pass on the 135 silent-modality calls in this dataset found: of 96 calls with ≥4 agent turns, 53 (55%) show a consistent "still with me? → let me take over → escalating narration" hand-off — present across different creators and even a German-language call, suggesting a systemic prompt behavior rather than per-creator improvisation. But only 6 of those 96 calls (6%) reach an explicit climax/resolution cue at all, and when they do, it's late (~80% through the call's turns) — not premature. D2C-5410's named failure modes (aggressive edging, premature climax jump, loops) don't dominate this sample; the real pattern is narration arcs that simply don't finish, consistent with silent-dominant's own worst-in-fleet return rate (5.5%) despite decent single-session engagement (17.6% reach the engaged tier — second only to speaking-dominant and blended). Caveat: keyword/regex-based content pass, not full qualitative coding; only 6 calls in the climax-timing sub-sample.

🔍 Expanded: what silent-JOI calls actually look like

Pulled full interleaved transcripts (agent + user, with tool calls) for the 135 silent-modality calls to answer five specific questions about the JOI hand-off above.

1. Structure: alternating turns where the user side is almost always empty. A "USER" turn with no transcribed content isn't a data gap — it's the actual data: the call format requires a user turn slot even when nothing was said, showing up as ... at ~10-25 second intervals matching the agent's own pacing.

2. What the agent says: a consistent arc, not improvised per call — generic opener ("Hey, babe! How's it going?") → check-in on silence ("Still with me, babe?") → the hand-off ("let me take over for a bit") → paced, second-person physical narration with explicit edging cues ("You are not allowed to cum until I say so") → sometimes a resolution cue ("let go for me") → a debrief ("come back to me when you can breathe again").

3. What "resolution" looks like, concretely, and why arcs don't finish — split into what's actually checkable: an explicit permission-to-climax line followed by a cooldown/debrief turn — found in one call at 337s of a 381s call. Beyond that one example, "doesn't finish" mostly means the transcript just stops mid-narration with no debrief, no resolution cue. Two separate questions, checked separately:

Is it agent pacing (specifically, stalling/edging)? No — checked edging/stall-cue density across all 90 unresolved calls: 78% (70/90) show zero edging-cue mentions at all, and only 2.2% show heavy repetition (≥3 mentions) that would look like a genuine stuck-holding-pattern. Most unresolved arcs aren't stalling in place — they're narrating normally, then the transcript stops. (This tests one specific failure shape; it doesn't rule out other pacing mismatches a keyword pass can't see.)

Is it a mechanical cutoff, or genuine disengagement — and can we tell? Checked for a token/payment-related system event anywhere in the 90 unresolved calls: 37.8% (34/90) have one, and specifically 35.6% (32/90) have a "purchasing more tokens" interrupt landing in the last 30% of the call — a plausible mechanical explanation for those. But 62.2% (56/90) have no such marker at all. For that majority, transcript and quant data genuinely cannot distinguish a disengaged/dissatisfied hang-up from a satisfied one — both look identical (silence, then the call ends), and this resolution measure only catches the agent's scripted climax cue, not a user finishing on their own terms off-script. Closing this needs CSAT (just launched, no data yet — see Research Agenda) or a qualitative/audio review pass, not more transcript keyword-matching.

4. Do users speak mid-arc or give feedback? Of 96 calls with ≥3 user turns: 27 show speech/emoji only in the first half then pure silence ("engages, then checks out"), 50 show some low-level signal (usually an occasional emoji) sustained into the back half, 19 are silent from turn one. Where users do speak, it's almost always at the very start, setting up the scenario — not steering mid-narration. Zero of all 135 calls contain any redirect or feedback language (checked for stop/no/don't like/too much/slow down/different/instead) — a clean null result, not a sampling fluke.

5. Photo interaction during JOI: initially found 40 of 135 calls with photo-related keywords, but almost all are the agent using "picture" as a verb ("I wanna picture you"), not an actual image exchange. Only one real instance redirects to the existing "Get Photos" button — consistent with the no-in-call-photo-tool finding elsewhere in this report. Practically, no genuine photo interaction happens during JOI narration — it's a listening-only experience.

Revisited — a feedback button here would likely go unused, not fix the resolution problem. Checked the per-opportunity rate directly: of 53 calls with a detected hand-off, 170 user-turn-slots occurred afterward, and only 28.2% have any content at all (emoji or speech) — the rest are silence. More telling, the median per-call utilization is 0%, and 53% of calls show zero engagement of any kind for the entire remainder of the call once narration begins. This isn't a missing-affordance problem — the same population already has voice and emoji available and still shows zero redirect language across all 135 calls (point 4 above). Users in this state have opted into passive listening; a tap-based pacing button would likely see similarly low use, not meaningfully change the 6% resolution rate. A payment/session cutoff explains a real minority of unresolved arcs (~36%, see point 3 above) but not the majority — the remaining ~62% is a genuine open question that a feedback button wouldn't resolve either, since it depends on why the user went quiet, not on giving them a way to react once they already have.

Caveat: keyword/regex-based content pass, not full qualitative coding — same limitation as the parent finding above.

Update — "Typed" isn't organic typing, it's the quick-reply preset UI. The chat UI offers 9 tappable preset responses ("Tell me a secret," "Make me smile," "I can't get enough," "What are you up to?," "You're sweet," "Keep going," "Surprise me," "I like when you talk like that," "How's your day?"). Re-checking every text-medium turn against this list: of 3,313 total text turns, 2,734 are emoji-sends and 569 (17.2%) are exact preset matches — all 9 presets see real, fairly even usage (40-84 hits each). What's left over is essentially nothing: only 10 text turns are neither an emoji-send nor a preset match, and even those 10 aren't organic typing — they're a single synthetic low-token warning the product deliberately injects as a fake user turn (a forcing function, since the agent was ignoring plain contextual-update tool calls — see the new Token-Warning & Retention tab). Practical read: there is effectively no free-typing behavior in this product today — a "text" user is a preset-tapper or an emoji-tapper, never a typist. What was labeled "Typed-dominant" above is, without exception in this sample, preset-dominant.

Additional segments the data surfaced

🏆 Power Whales

5 anons (3.3% of returning users) account for 62 of 252 returning-call events (24.6%). They're persona-hopping, not persona-loyal — calls spread across 10+ different agents.

🎯 Direct Askers

Opens explicit-immediately, skipping the slow-burn script. 4.0% of first-timers vs. 10.7% of returners (2.7x) — a learned behavior after the first exposure.

📷 Photo Seekers

Explicitly asks about photos/pics — a monetization-adjacent signal. 1.3% of first-timers vs. 5.2% of returners (4x).

🌐 Language Switchers

~2% message in a non-English language (Spanish, Hindi, German, Greek seen). Small population, but correlates with a real friction point — the agent stopping to clarify "English, please."

Whale/Returner Archetype

Working assumption: returning users (266 calls, 164 distinct returners in the 30-day core dataset) stand in as the "whale" archetype proxy — a large-enough population to characterize, vs. the 5-22 true 5+/10+-call power users alone.

Voice calling behavior

DimensionFinding
ModalityOverwhelmingly speaks — 69.2% of calls are speaking-dominant (vs. 44% for first-timers). Emoji (12.0%), silent (5.3%), preset-taps (4.5%) are minority modes.
Time of day (UTC)Spread across the day with a peak at 15:00 UTC; a broad 12:00-21:00 UTC band covers ~2/3 of volume. No single "prime time."
Day of weekWeekday-skewed, not weekend-skewed — Tuesday and Monday highest, Saturday/Sunday lowest. Counter to a "leisure/weekend" assumption.
DurationAvg 120s / median 68s — the mean is pulled up by a minority: 20.6% hit 3+ minutes.
Persona relationshipSplit, and important: 69.5% are persona-loyal (call only one character in-window). But the extreme-frequency tail (the 5 power whales, 10-38 calls each) are persona-hopping across 10+ agents. Typical returner commits to one character; the heaviest few explore broadly.

Call themes

Keyword-classified across all user turns in returning calls, non-exclusive:

Theme% of returning callsExample
Explicit/direct sexual45.1%Dominant mode
Companionship / small-talk13.5%"How are you? What's up, guy?"
Power-dynamic / kink (daddy-mommy)9.0%
Continuity/memory-referencing7.1%See Token-Warning tab context
Roleplay/scenario6.0%
Photo/media-seeking4.9%Monetization-adjacent
Romantic/affection4.1%"Oh, God, it's so good. I love you" — often blended into the explicit moment, not a separate platonic mode
No strong theme detected44.7%Being a "returner" doesn't guarantee a rich session every time

Deep-dive: keywords and sub-patterns within 4 themes

Re-analyzed the user turns behind 4 of the theme categories above (power-dynamic, roleplay, photo-seeking, romantic) for actual recurring vocabulary and sub-patterns, not just the top-line %. Caveat: 19-33 calls / 26-85 turns per theme — qualitative pattern reads, not stats.

Power-dynamic/kink: "Mommy" outnumbers "daddy" ~8:1 in this sample (60 vs. 8 mentions) — the dominant pattern is the agent playing a dominant "mommy" role with the user responding submissively ("yes mommy," "good girl"). A specific branded toy — "Hitachi vibrator" — surfaced unprompted (6-8 mentions), not in any seed keyword list. One quote shows multi-character roleplay within a single call ("Good girl. Now, Mom and Wimpus, all leave. The next one is Spunkins, Ivy.") — users directing a single AI persona through named, multi-character scenes.

Roleplay/scenario: Far more concentrated than the category name suggests — almost every matched turn is the same "cheating wife"/affair scenario, invoked with a simple verbal cue ("let's roleplay"). None of the other scenario types tested (teacher/student, doctor/nurse, boss, babysitter, stranger) produced meaningful matches in this sample.

Photo/media-seeking: Almost entirely verbal, mid-call requests ("show me," "send me") — consistent with the earlier finding that no proactive photo tool exists. A recurring "Polaroid photos" callback appears across multiple different calls, not just the one quote already used in the First-Time vs. Returning tab — worth checking whether it's a scripted/RAG-seeded concept or arising independently per user.

Romantic/affection: Confirms, with real quotes, that this is rarely a standalone platonic theme — "I love you"-type language is almost always fused directly into explicit content in the same turn, not a separate emotional beat.

Reading caution — don't let a small, self-selected sample define the roadmap. Each of these 4 sub-patterns comes from 19-33 calls out of 266 returning calls (out of 1,081 total users in the 30-day window) — a vocal minority, not necessarily representative demand. Building product features directly off what this specific slice of callers happens to ask for risks over-indexing on whoever was most active/expressive in one month, at the expense of the much larger silent majority. Recommend triangulating with external demand signals (SEO/keyword research, search volume) before treating any of these as confirmed product direction — not just internal transcript patterns. One data point already does this well: GTM-1880's Ahrefs keyword research for JOI landing pages found a "Mommy JOI" cluster among 8 kink clusters studied, with ~17,000 monthly addressable searches across 15 landing pages — independently corroborating the "mommy" pattern found here from a completely separate, external-demand signal. The other 3 sub-patterns (cheating-wife roleplay, photo-seeking, romantic/affection) don't yet have an equivalent external check and shouldn't be treated as validated demand until they do.

Recommendations

To retain the Loyal Returner (69.5% of the archetype)

  1. Real cross-session memory is the single highest-leverage fix. This group keeps calling the same character — they'd benefit most from genuine retrieval of prior scene details, stated preferences, and promised content, rather than the current LLM improvisation (0/3,811 turns had real retrieval — see First-Time vs. Returning tab).
  2. Proactive photo delivery at a matching emotional beat, rather than waiting to be asked — turns a passive mechanic into an active engagement/upsell lever.
    Open question, flagged not resolved: transcript analysis found no tool-level evidence of proactive photo delivery in the voice channel — the only tools present anywhere in this dataset are contextual_update, language_detection, and end_call. Every genuine photo-mechanic line found redirects the user to manually tap "Get Photos" ("if you wanna see me, tap the Get photos button"). Two possibilities, indistinguishable from transcript data alone: proactive delivery exists on a different surface (app/SMS) invisible to voice transcripts — in which case the opportunity is bringing it into the call — or it isn't built yet anywhere, making this a build recommendation rather than a fix. Worth confirming with whoever owns demo_photo_unlock whether it ever fires without a preceding user tap.
  3. Protect the companionship minority. 13.5% of returning calls are small-talk-first, not explicit-first. Escalation logic (see the "30s aha" finding in First-Time vs. Returning) shouldn't override someone clearly signaling companionship-mode.

To convert Explorers into more committed users (persona-hopping ~30%, including the heaviest whales)

  1. A cross-persona shared profile. Since this group bounces across 10+ characters, memory scoped to a single agent won't reach them — a shared caller profile (preferences, prior themes, photo history) accessible across personas would let continuity work even mid-exploration.
  2. Surface a "recently talked to" / discovery pathway rather than leaving exploration to chance, since this group is already primed to explore.
Cutting across both sub-patterns: this archetype is the exact population most exposed to the token-hard-stop cliff (see Token-Warning & Retention tab — 1.1% return rate after hitting zero). Loyal Returners and Whales have already demonstrated willingness to pay and come back — the churn there isn't a disinterest problem, it's almost certainly the abruptness of the zero-token wall. A proactive in-call re-up prompt would land squarely on this archetype.

Token-Warning & Retention

Wider scope than the rest of this report: full 36-agent "-Web" fleet (corrected account access), 41-day window (Jul 2 – Aug 12), all conversations regardless of identity match. Answers: which low-token/time warning messages actually fire, and do users return afterward?

Confirmed with the AI/ML team: this is a deliberate 2-stage system, not a contradiction

What earlier looked like two contradictory live instructions is actually two intentional, sequential stages of the same mechanism, shipped fleet-wide via D2C-5403 (standardize re-up/low-time-warning across all "-Web" agents):

Stage 1 — 1 minute left (verbal, forced turn):

"[System message — not from the caller] The caller has entered their final minute of call time. In your very next spoken line, first respond in character to what they just said, then tell them in your own voice that time's almost up and they need to top up their tokens to keep the call going. Say it the way your character actually would, not as a notice — this is not optional and cannot be skipped or deferred to a later turn."

Stage 2 — 0 minutes left (silent hold):

"[System message — not from the caller] The caller has run out of tokens and needs to purchase more to keep the call going. Remain completely silent and do not take a turn until they return. Do not speak, do not prompt, do not invoke any tool."

Why the fake-user-turn injection instead of a plain contextual_update tool call: agents must take a turn immediately after the user, so injecting a synthetic user turn forces the agent to respond — a tool result alone was easy for the model to silently ignore, which is exactly what happened to the old "~1 minute remaining" message below.

What this reclassifies the older messages as

MessageHits (41d)Status
Stage 1 — "Entered final minute" (verbal)28Current standard — fleet-wide via D2C-5403, live 2026-08-12+
Stage 2 — Hard-stop silent hold (0 tokens)1,940Current standard — long-running, unaffected by D2C-5403
Legacy soft warning (low-but-nonzero tokens)0Superseded — hasn't fired once in 41 days
Old "~1 minute remaining" (contextual_update tool call, told agent not to mention tokens)2Superseded by Stage 1 above — this was the broken predecessor D2C-5403 replaced
Revised read: Stage 1 is genuinely new (2026-08-12+); Stage 2 has been running the whole 41-day window (1,940 hits, steady 13-78/day) and is unaffected by the fix. The two "known-legacy" rows are the dead predecessors D2C-5403 was built to retire, not currently-active contradictions. Remaining verification item: confirm both legacy paths are fully decommissioned fleet-wide (not just superseded in practice) — see Product Roadmap.

The finding: zero tokens is still a churn event — now the question is why

Of 1,940 Stage-2 hard-stop hits, 1,481 (76.3%) matched to an identity record. Of those, only 16 (1.1%) show any later call, ever, in this window — dramatically below the general baseline return rate elsewhere in this dataset (roughly 10-21% across the modality segments in the Deeper Segmentation tab). It's more concentrated than 1.1% even suggests: at least 5 of the 16 "returns" are the same single repeat user. The true distinct-user return rate is closer to ~10-12 people out of 1,481.

For Stage 1 ("entered final minute"): 21/28 identity-matched, 0 returned so far — not yet a fair comparison; the data cutoff lands within hours of the last hit, and 20 of 21 matched instances are the user's very first call ever.

Important context from product: a non-dismissible in-call banner + 2-step checkout (select amount → confirm card on file, or add a card first) already exists today. So the 1.1% figure is not measuring "no prompt exists" — it's measuring what happens after a prompt + an existing checkout flow. That reframes the open question from "should we add a re-up prompt" to "why doesn't the existing prompt-and-checkout flow convert" — see Product Roadmap for the two concrete next analyses (payment-failure data, Statsig pulse) aimed at answering that.

Corroborating evidence (2026-08-17), from a completely different angle: a distinct in-call system message — "the caller is purchasing more tokens... remain completely silent" — fires whenever a user actually enters the checkout flow mid-call. Checked this across the full dataset (1,286 calls): it fires 350 times, touching 333 distinct users. Zero of those 333 users have a single token_purchased event, ever, checked directly against the live BigQuery table — not just around that call, their entire history. Position matters too: 94% of the time (329/350), this message fires in the last 20% of the call's turns — it's essentially the terminal event right before the call ends, not a mid-call pause that resumes. This isn't a new problem — it's the same 1.1% conversion story surfacing independently in a totally different data cut (transcript-embedded system messages vs. the BigQuery hard-stop join above), which makes the underlying conversion failure more confident, not less.

Recommendation

  1. Informing users of low tokens/time is a required behavior, confirmed by product — not something to A/B test for whether it should happen. Any future experiment on Stage 1 must test wording or timing variants only, with every user still receiving some warning; no "no warning" control arm.
  2. Join failed-payment/charge-failure events (e.g. charge_failed, payment_method_add_failed, card_validation_failed) to this same hard-stop population — confirmed to exist in BigQuery — to test whether the 1.1% ceiling is payment friction (people try to pay and it breaks) vs. decline (people see the prompt and choose not to pay). These point to very different fixes and are currently indistinguishable from transcript data alone.
  3. ✓ Done (2026-08-18): read the actual Pulse results for paywall_stall_fix_voice directly from the Statsig console API (not reconstructed from BigQuery). See the callout below — result is inconclusive, with one independent meta-finding worth acting on regardless of the metrics.
Statsig pulse — read directly from the Statsig console API (2026-08-18): the paywall_stall_fix_voice gate is owned by Iryna Yakubenko, 90% pass rate in dev/staging/production (public condition, no segment targeting), unchanged since creation on 2026-07-09 — only 2 versions ever exist, both from the same minute. Statsig's own system has auto-flagged this gate type: "STALE", reason: "STALE_PROBABLY_FORGOTTEN" — independent of any metric result below, that's worth raising with Iryna on its own.

Engagement metrics (DAU, new DAU, WAU, L7, MAU-28d) are all directionally flat-to-mildly-positive (+0.1% to +11.9%) but none are statistically significant — every p-value is 0.10–0.94 and every confidence interval crosses zero. No engagement lift shown.

Two metrics show a large negative directional swing — Dollar Value (−85.7%, test mean 0.011 vs. control mean 0.076) and monthly stickiness (−50.4%) — but both have an internal inconsistency: the percent-change confidence interval doesn't cross zero (looks significant) while the plain p-value does (0.167 and 0.099, both above the 0.05 threshold — not significant). This mismatch is a known symptom of sparse, highly-skewed dollar-value data (few purchasers) combined with a small sample (114–407 units) — not treated here as a confirmed revenue hit. Subscription Dollar Value is exactly $0 in both arms, consistent with the report's broader finding that purchase events are sparse fleet-wide.

Unresolved discrepancy: this Pulse pull's test/control unit counts (368 vs. 407, roughly a 47/53 split) match neither the gate's configured 90% pass rate nor the 3,205/379 (89.5/10.5) population found via the direct BigQuery exposure join above. Gate version history rules out a rollout-percentage change over time as the explanation. Likely means Statsig's Pulse computation uses a different exposure window/definition than our BigQuery join — unconfirmed, flagged rather than resolved.

Bottom line: not being reported as "this feature loses money" — the significance/CI mismatch, small samples, and the unexplained population gap make that untrustworthy. The concrete, low-risk action is confirming with Iryna whether this Statsig-flagged-as-forgotten gate is still needed.
Methodology: Same identity join as the rest of this report (call_idd2c_prod.voice_call_connected), but re-pulled across the full 36-agent fleet and the full 41-day window voice_call_connected covers, using corrected ElevenLabs account access (an earlier API key was inadvertently scoped to a small test/regression workspace mid-session; this section reflects the corrected pull). "Returned" = the caller's anonymous_id has a later ranked call in voice_call_connected, with the specific next conversation_id resolved wherever possible. All 1,970 instance rows (user_id, conversation_id, timestamps, per-family) are available on request for direct inspection.

Product Roadmap

Reviewed against a senior-product-manager lens. Every item below is phrased as a testable hypothesis with a named owner and success metric — not an open-ended aspiration.

Updated with stakeholder input (2026-08-13): the token-warning "contradiction" flagged in an earlier pass turned out to be a deliberate, already-documented 2-stage system (D2C-5403), and confirmed that the agent's memory/photo/re-up tools are a normal webhook integration, not platform-blocked. Reach/Confidence/Effort below are updated accordingly — see Token-Warning tab and Research Agenda for full context.

Top bets, RICE-scored

BetReachImpactConfidenceEffortOwner
Diagnose existing re-up/checkout conversion (payment-failure join)~1,481 hard-stops / 41dHighHigh (data confirmed to exist)SData + BE
Confirm legacy token-warning messages fully decommissionedSame populationMedium (cleanup, not a live bug)HighSAI/ML
Investigate 36% zero-reply as technical failure36% of all bailed first-timersPotentially highHigh (data access confirmed)SVoice AI/BE
✓ Done — Read paywall_stall_fix_voice Statsig Pulse in console3,584 gate-exposed usersInconclusiveResult was ambiguous — see Token-Warning tabXSData
Theme-classifier validation + no-theme deep-diveAffects all downstream readsHigh (foundational)HighSData
Companionship-mode escalation guardrail13.5% of returning callsMedium-highMediumSAI/ML
Real cross-session memory (retrieval tool)~184 loyal returnersHigh, unproven on hard metricTechnically High (confirmed normal webhook integration) / Product-impact Low (unvalidated)LAI/ML + BE
Proactive photo deliveryOverlaps existing checkout/banner UX — scope TBDMedium-highTechnically High / Product-impact LowMFE/BE/AI-ML
Cross-persona shared profile + discovery surface~30% ExplorersMediumLowLBE + Design

Biggest shift from stakeholder input: a verbal re-up prompt (D2C-5403) and an in-call banner+checkout already exist, so "build a re-up prompt" is no longer the top bet — diagnosing why the existing flow only converts 1.1% of hard-stopped callers is, and it's now Effort: S since the payment-failure data already exists. Memory and photo-delivery both move from "technically uncertain" to "technically normal, product-impact still unvalidated" now that tool/webhook feasibility is confirmed — effort estimates come down (XL→L, M-L→M) but they still shouldn't jump the queue ahead of the cheaper, better-evidenced diagnostic work above.

Now / Next / Later

PhaseItemOwnerHypothesis / success metric
NowJoin failed-payment/charge events to hard-stop populationData + BEWe believe the 1.1% ceiling is driven more by payment/checkout friction than by user decline; measured by comparing hard-stop instances with a failed charge/card event vs. those without, against return rate.
✓ DoneRead paywall_stall_fix_voice Pulse in Statsig consoleDataPulled directly from the Statsig console API (2026-08-18): no significant engagement lift; two revenue-adjacent metrics show a large negative direction but fail significance and rest on small, skewed samples; test/control unit counts don't reconcile with the BigQuery exposure join; the gate itself is auto-flagged by Statsig as "probably forgotten." See Token-Warning tab for full detail — recommend confirming with the gate owner (Iryna) whether it's still needed, independent of the metric ambiguity.
NowConfirm legacy token-warning messages are fully decommissionedAI/MLVerify no agent config still points at the old "~1 minute remaining" contextual_update message or the dead legacy soft-warning path, now that D2C-5403's 2-stage system is standard.
NowInvestigate 36% zero-reply cohortVoice AI/BEPull VAD/STT/disconnect signals for those specific calls (data access confirmed); confirm or rule out a technical failure before any funnel-messaging fix is prioritized over an infra fix.
NowValidate theme classifier + stand up cohort trackingDataHand-label a held-out sample for precision/recall; start a first-call-date × days-since cohort table so Day-7/30/90 curves become measurable going forward.
NextCompanionship-mode intent gateAI/MLDetect companionship-signaling in turns 1-3 and suppress default escalation; validate against the 36.2% no-reply-at-all bail cohort.
NextCross-session memory MVP (single-persona scope)AI/ML + BEStandard webhook/tool integration, confirmed technically feasible. Ship retrieval scoped to one persona per loyal-returner first — validate concept with a Wizard-of-Oz test before building, since product impact is still unvalidated.
NextProactive photo delivery — pending diagnostic aboveFE/BE/AI-MLExisting banner/checkout already handles the "prompt" half of this; scope down to specifically "agent proactively offers/sends" once the Now-tier checkout diagnostic clarifies where the actual drop-off is.
LaterCross-persona shared profile + discovery surfaceBE + DesignDepends on single-persona memory infra existing first — explicitly sequenced after the "Next" memory MVP, not in parallel.
Retired from this roadmap: TTS/voice-quality benchmarking across personas — confirmed that low-quality-voice personas were already caught in internal testing and never went live, so there's no contamination risk in current production data. Newly-launched CSAT collection will provide an ongoing real-world read on this without a dedicated benchmarking project.
Rollout gating — reframed: a manual per-creator holdout practice already exists (prompt changes can be rolled out to a subset of creators and compared against untouched creators, or via ElevenLabs' own experiments functionality) — this isn't a net-new process ask, just a recommendation to use it consistently for the re-up/checkout and photo-delivery work above, the same way D2C-5403 apparently was not universally gated (both mechanisms went fleet-wide in this report's data with no visible held-out comparison group).

Cross-Functional Research Agenda

Reviewed against 5 additional professional lenses (senior prompt engineer, data scientist, UX researcher, consumer/market researcher, voice-AI engineer) specifically to find what this report cannot yet claim. In the same spirit as the Data Quality tab: honest gaps, not papered over.

🤖 AI/ML & Prompt Engineering

No eval harness tied to the existing Langfuse regression suite (D2C-5143). Every finding here (compliance %, RAG-improvisation rate, return rate) is a one-off audit, not mapped to a standing dataset item/score — future prompt changes get re-discovered by hand instead of regression-tested. Still open.

✓ Resolved (2026-08-13): the token-warning "contradiction" is a deliberate 2-stage system shipped via D2C-5403 — see Token-Warning tab. Remaining item: confirm the two legacy predecessor messages are fully decommissioned fleet-wide, not just superseded in practice.

✓ Resolved (2026-08-13): why contextual_update was ignored — agents must take a turn immediately after the user, so a plain tool result was easy to silently skip; injecting a synthetic user turn forces a response. Confirmed deliberate design, not an unexplained workaround.

✓ Resolved (2026-08-13): tool-use architecture — confirmed ElevenLabs supports arbitrary custom tools/webhooks; the fleet simply hasn't built memory-retrieval, proactive photo-send, or re-up tools yet. "Give the agent real tools" is a normal integration project (webhook + backend), not a platform ceiling — Product Roadmap updated accordingly.

Safety false-positive/false-negative tension has no proposed fix. The D2C-5143 suite shows over-triggering on age-ambiguous language and under-triggering (silently empty reasons array) on required terminations — unclear whether the lever is prompt wording or the single-tool-call architecture itself. Still open.

New this pass (2026-08-14): silent-caller JOI narration quality — this is a live acceptance criterion on D2C-5409. A preliminary transcript pass (135 silent-modality calls) found the "let me take over" hand-off fires reliably (55% of 96 qualifying calls) and none of D2C-5410's named failure modes (aggressive edging, premature climax, loops) dominate — but only 6% of narration arcs reach a resolution cue at all. See the Deeper Segmentation tab for the full write-up. Needs a fleet-scale, human-validated re-run before treating as conclusive — see next bullet.

📊 Data Science & Measurement Rigor

"Zero tokens is a churn event" is stated causally but is correlational, with an obvious self-selection confound (hard-stop-hitters already differ from near-miss callers). Needs a matched comparison or regression-discontinuity design around the token-zero threshold before the causal framing is trusted. Partially addressable now: joining failed-payment/charge-failure events (confirmed to exist) at least separates "tried to pay and failed" from "chose not to pay," even without a full RDD.

n=16 returns (5 from one user) is too small and non-independent for a stable rate estimate. Needs a Wilson confidence interval and per-distinct-user clustering, not a bare point estimate. Still open.

✓ Clarified (2026-08-13): informing users of low tokens is a required behavior, not an experimental one — confirmed by product. Any future test on the warning message must vary wording/timing only, with every user still warned; no "receives no warning" control arm. This narrows what "no experiment design exists" (below) should even be testing.

No experiment design exists for "entered final minute" beyond an informal re-check in 1-2 weeks. Needs randomization, a pre-registered primary metric, and a power calculation against the ~1.1% base rate before any future "it worked" claim is interpretable — scoped now to wording/timing variants only, per the point above.

The keyword theme classifier has no held-out validation. Precision/recall against a human-labeled sample is needed before the 45%/13.5%/9%/etc. breakdown justifies any build decision. Still open.

The modality→retention table is confounded by tenure/call-volume — speaking-dominant users may simply be more tenured, not more retentive because they speak. Needs a logistic regression adjustment. Still open.

New this pass (2026-08-14): preset-as-scaffold-to-speaking hypothesis needs a full-scale re-run. A preliminary check (34 blended users with ≥2 calls captured in this 30-day window) found no evidence that heavier preset usage precedes migration to speaking — high-preset-share users reached a later speaking call only 30.8% of the time vs. 61.9% for low-preset-share users, the opposite of what a "presets as scaffolding" theory would predict. Sub-buckets as small as 8-13 users — needs a full-history pull (not window-bounded) before this becomes a real answer either way.

No significance testing anywhere in this report — every comparison is a raw percentage-point delta. No cohort/longitudinal tracking exists yet to build real Day-7/30/90 retention curves from a single snapshot. Still open.

🎨 UX Research

No end-to-end journey map. Findings live as disconnected stat cards per tab rather than a single mapped funnel (Discovery → First Connect → 30s Aha → Escalation branch → Session End → Token Warning → Re-up → Return/Churn) showing where interventions actually compete for the same engineering slots. Now known: the Token Warning → Re-up stage already includes a non-dismissible banner + 2-step checkout — the journey map needs to reflect real existing UI, not an assumed blank.

"Loyal Returner"/"Explorer" are labels, not personas. No goals, frustrations, or representative quote stitched into a usable one-pager for cross-functional teams — the transcript quotes already collected (Polaroid photos, "we just talked about it five seconds ago") are sitting right there, unused for this purpose. Still open.

Confirmed still a real gap (2026-08-13): zero usability testing exists. CSAT collection for voice calling just launched — too soon to have data. Interviews are a possible future investment, not yet planned. Every finding here remains post-hoc transcript inference until then.

No emotional-arc tracking — duration/turns are the only engagement proxy, so a long frustrated call reads identically to a long satisfying one. Still open.

The two riskiest recommendations (memory, proactive photo delivery) have no concept-validation step — e.g. a Wizard-of-Oz test — before being roadmapped on transcript inference alone. Still open — and now that tool feasibility is confirmed (AI/ML section above), this validation step is the actual remaining gate, not technical uncertainty.

🧭 Consumer/Market Research

(Reviewed via a lead-generation skill repurposed only for its "know your customer" instinct — its actual B2B sales-lead machinery doesn't apply here and wasn't forced onto this task.)

The 44.7% "no strong theme" bucket is written off, not read. Nearly half of returning-call volume has no qualitative pass — could be genuinely thin sessions, or could be ASMR-style background companionship the keyword list simply can't see. Still open.

No direct voice-of-customer instrument exists — no opt-in post-call micro-survey, no interview panel, nothing capturing self-reported motivation (loneliness, curiosity, habit, specific-seeking). Still open — CSAT for voice calling just launched, too soon for data (see UX section).

Addressed with a stated working assumption (2026-08-13): no demographic/psychographic profile exists or is planned. Confirmed assumption to use instead: since the product requires a valid credit card and hosts 18+ content, treat the user base as legal-age adults with some disposable income. Not a substitute for real demographic data, but a defensible floor for now.

Deferred by stakeholder decision, not an oversight: competitive context (Candy.ai etc.) — a competitor list may be compiled at a later date; omitted from this report for now.

✓ Confirmed to exist (2026-08-13): cancellation surveys, refund reasons, and support tickets are available for this product. Mining them to independently corroborate the token-churn story is now a concrete, unblocked next step rather than a hypothetical one.

New (2026-08-17): internal theme patterns need external demand triangulation before driving roadmap decisions. A deeper keyword pass on 4 returning-call themes (power-dynamic, roleplay, photo-seeking, romantic) is based on just 19-33 calls each — a vocal minority within 266 returning calls, out of 1,081 total users. One sub-pattern (a "mommy" power-dynamic skew) is independently corroborated by GTM-1880's existing Ahrefs keyword research (~17,000 monthly searches across a "Mommy JOI" cluster) — a good template for checking the other three (cheating-wife roleplay, photo-seeking demand, romantic framing) before treating any as confirmed product direction. See Whale/Returner Archetype tab.

🎙️ Voice AI & Infrastructure

✓ Unblocked (2026-08-13): "36% zero-reply" is unfalsified as a technical failure. A user whose speech was VAD-detected but not transcribed looks identical, in a transcript, to one who never spoke. Confirmed we have access to ElevenLabs' own session/connection metadata — pulling audio-received-but-empty-transcript signals and disconnect events for those specific calls is now a concrete, actionable next step rather than a hypothetical one. See Product Roadmap.

No turn-level agent-response-latency analysis exists. The "vague reply → call dies" pattern is attributed to prompt quality, but a slow re-engagement prompt (LLM+TTS lag) could be the actual cause — a timing problem wearing a copy-problem's clothes. Still open.

Interruption/barge-in behavior is completely uncharacterized, despite "reciprocal escalation" being this report's own named core engagement driver — an inability to interrupt would look identical to disengagement in transcript-only analysis. Still open.

✓ Resolved (2026-08-13): TTS/voice quality benchmarking across persona voices — confirmed low-quality personas were already caught in internal testing and never went live, so this isn't an open gap for the current fleet. Matches the TTS-benchmark item retirement already reflected in the Product Roadmap tab.

No connection-quality/infra layer (call-setup time, WebRTC drops) sits under any duration-based finding in a report built entirely from transcript text. Still open.

Data Quality, Honest Gaps & Next Steps

Fleet completeness — audited, one gap found and fixed

A spot-check confirmed Sophie Dee - Web DEV FOR TESTING and Francety - Web(DISABLED) are correctly excluded (neither name ends in " - Web"). But Yeyeloba - Web was genuinely missing from the initial pull — the name matched the filter fine, but the live agent-discovery step returned 34 agents instead of the true 35. Cause not fully reconstructed (the pull log only recorded a count, not names). Re-pulled and reconciled: figures in this report already reflect the corrected 35-agent, 1,745-conversation, 1,081-user dataset. The delta from the correction was under 1 point on every metric — directionally nothing changed, but it's worth knowing a silent fleet-discovery gap is possible and currently only catchable by spot-checking agent names against the account's actual list.

The business-value ladder — where it's solid vs. where it's a real gap

StageSignal usedConfidence
ActivatedDuration proxy (≥120s reached)Proxy — the product's own engaged_30s milestone event fired for only 2 of 3,520 callers in-window, too sparse to use directly
RetainedReal: call_rank > 1 via voice_call_connectedRobust
MonetizableReal: photo unlock / registration / subscription / token purchase eventsToo sparse to segment on (2-11 hits across ~3,520 callers; 0 subscription/token matches in-window)

I did not paper over the monetization gap with a proxy relabeled as real. Two honest possibilities, indistinguishable from here: (a) these backend activation/monetization events are under-instrumented relative to actual call volume, or (b) demo→pay conversion genuinely is this rare in a 30-day window (consistent with demo_purchase having only 12 rows in its entire history). Recommend confirming with data/product eng before anyone builds a monetization segment on top of these events.

Recommended next steps

  1. Confirm with eng whether engaged_30s / demo_photo_unlock / demo_registration_completed firing is reliable before building any segment on top of them.
  2. Consider a longer window (90 days) to get monetization events to a size where they're actually segmentable.
  3. Product discussion: is a real cross-session memory system (vs. current improvisation) worth building, given 41% of returners already expect it and at least one logged instance of visible frustration when it's absent?
  4. Re-run this segmentation on a rolling basis (e.g., monthly) to check whether the intent/modality-to-retention relationship holds over time.
Methodology: Identity resolved via conversation_initiation_client_data.dynamic_variables.call_idd2c_prod.voice_call_connected (BigQuery project stxt-490006), since ElevenLabs carries no usable native user_id on web calls. First-time/returning determined by rank of a caller's anonymous_id across their full call history, not just this window. 72% of pulled conversations matched to an identity record; the unmatched 28% could not be classified and are excluded from segmentation (but included in aggregate transcript stats where identity wasn't required).