ElevenLabs · Funnel Analysis
How many ad clicks turn into a paying subscriber — the full funnel, ad click through purchase, plus what happens inside the call itself.
% shown is each stage as a share of the stage directly above it (started/landed, gated/started, registered/gated, purchased/registered) — not of the top-of-funnel total.
The shorter ~50-second demo gets more people to the gate — 50.0% of started calls vs 42.3% before — exactly as expected from a faster trigger. But registration completion once they're there roughly halved: 21.6% vs 44.2% of gate-reaches. Overall registered-per-landed fell from 4.2% to 2.1%. Read together with "Gate mechanics" above, this looks like a real tradeoff, not just a labeling artifact: a faster paywall reaches more people, but asks before as much rapport or investment has built, and fewer of them actually follow through.
Sample sizes are still modest on both sides (52 vs 88 gate-reaches) — worth re-checking as more volume accumulates post-change. The purchase comparison (4 → 0) isn't a clean read on its own: all 4 pre-change purchases already look like internal testing concentrated on launch day, so this isn't necessarily "purchases stopped after 7/3" so much as "the small test-purchase burst happened to land before it."
All 4 purchases (5 events — one person subscribed then bought a token top-up 99s later) landed within the first ~14 hours of launch on June 30. That timing pattern reads as internal/team testing rather than organic customers, though the data itself has no test flag to confirm it either way — treat the 0.3% purchase rate as unconfirmed until checked with the team directly, not as a real conversion signal yet.
Correction to the earlier estimate: the first version of this report estimated "~5% reach the registration wall" using only a 2-minute call-duration cap visible in ElevenLabs data. The real number is 47.0% of started calls (140 of 298) — nearly 10× higher. Two of the three gate triggers (tokens, 4 photos) were always faster than 2 minutes — and, as it turns out, so is the third: see "Gate mechanics" below for why the "2min" trigger itself no longer means 2 minutes.
Checking the actual time-to-gate for calls tagged "2min": median 50 seconds, tightly clustered (nearly all land at 50–51s), not anywhere near 120s. Matched against team context: on 2026-07-03 09:07:07 UTC, the default pre-signup call length on the VoiceLander was changed from 2 minutes to 50 seconds — but the reason string logged by the gate event was never updated to match. All 77 "2min"-tagged gate events in this window fall after that change (earliest: 2026-07-03 17:07 UTC, ~8 hours post-change); none fall before it.
This is a clean before/after flip. Before the change, "free tokens ran out" was overwhelmingly the dominant gate trigger (83.6%) — consistent with the original design, where a real 2-minute window gave the token pool (and occasionally 4 photo unlocks, 14.5%) time to run dry first. After it, the new ~50-second cap swamps everything else (89.9%) — there's barely enough time left for tokens to drain naturally, and none at all for anyone to unlock 4 photos (0%, down from 14.5%).
The combined 57.9% / 39.3% / 4.3% split reported earlier is real and correct as an aggregate over the full window — but it blends two genuinely different regimes together. Splitting by the 7/3 change is the more honest way to read this data going forward.
Product direction as of this update: the token-spend/countdown mechanic is being removed from the demo entirely for the next iteration — the lander's job is to get callers to earn tokens as a hook toward signing up, not spend them down during a free trial. That will also require redefining the "tokens_zero" gate trigger (and revisiting the pre-signup call-length default itself, given the above), since both are tied to mechanics being changed.
| Event | Distinct users | Total events | Avg per user |
|---|---|---|---|
| Emoji reactions | 111 | 1,892 | 17.0 |
| Photo spins | 146 | 295 | 2.0 |
| Photo unlocks | 93 | 164 | 1.8 |
| Heat milestones hit | 80 | 132 | 1.7 |
| Chip taps | 56 | 68 | 1.2 |
The 93 distinct photo-unlockers here lines up closely with the 89 calls the ElevenLabs transcript analysis flagged as having a photo-reveal moment — good independent cross-validation of that earlier finding, from a completely different data source. Emoji reactions are by far the highest-frequency action (17/user on average) — a much lighter-weight engagement signal worth keeping in view alongside the heavier photo/registration metrics.
44.4% of calls end inside 30 seconds (28.1% inside 15). This isn't the agent losing people mid-conversation — it's the caller's client disconnecting almost immediately: 90%+ of sub-30s calls end in a client-side disconnect (WebSocket codes 1000/1001/1006), not a hangup after real dialogue.
That points upstream of the script — at ad targeting quality, the lander's expectation-setting, or connection/audio-startup friction — rather than at anything Sophie says or does.
Silent — no real spoken or typed input before disconnect. 66 of the 87 never produced a single user turn at all (median time-to-disconnect: 5s, an instant bounce); the other 21 had a turn fire but ASR caught only silence/non-verbal sound (median 18s — present, but never said anything).
Engaged, then left — real content before churning: greetings, "send me a photo" (the single most common real message, 13×), emoji taps (💦🔥), short replies.
Tech issue — ElevenLabs itself logged "Connection closed before conversation started" or "Call ended without transcript"; the session never functionally started, independent of caller behavior.
Two-thirds of the early drop-off is silence, not disengagement-after-trying or infrastructure failure — more consistent with low-intent ad traffic (accidental or curiosity clicks that never meant to talk) than with a script, UX, or connectivity problem. Tech failures are real but small (2.3% of all 302 calls).
Data gap found while checking this: the backend does track mic-permission and connection-quality events (voice_mic_prompt_shown, voice_call_start_failed, voice_asr_first_audio) that could confirm or rule out a technical explanation for the silent bounces — but they're only instrumented on the in-app chat voice feature (/chat/<id>), never on the VoiceLander demo page. That's a real instrumentation gap, not a null result: this hypothesis stays unconfirmed until the same events are added to /call-sophie.
Sophie's own prompt is explicit: only end the call for one of 8 named safety conditions, never for silence or a timer — "re-engage, never wind down or say goodbye first." In practice, the end_call tool fired on 55 of 302 calls (18.2%), and none were for an authorized safety reason.
The shortest: "Call duration reached two minutes trial limit" at 67 seconds. Median duration for this bucket is 96s, range 40–126s. The model has the exact elapsed time available every single turn via {{system__call_duration_secs}} — this isn't a measurement error, it's a fabricated justification.
Checked whether the backend's registration-wall gate (2min / tokens_zero / 4photos) is what's actually triggering this: no. None of the 33 transcripts show a contextual_update or tool result carrying gate/token/photo state — the two systems are completely disconnected. The agent has no visibility into the backend gate at all.
Confirmed false positive: one call ended over a caller asking "Can you give me a recipe for pancakes?" — logged reason: "User requested non sexual, assistant restricted from giving cooking advice and conversation must end via tool." Nothing about that warrants ending a call.
Actively engaged callers mislabeled "unresponsive": several calls cite silence immediately after the caller sent an emoji burst, an explicit encouragement ("don't stop babe"), or a repeated photo request — the exact signals this report's "What's working" section flags as the strongest positive indicators. The mislabel is landing on the highest-intent segment, not just genuinely disengaged callers.
A euphemism worth checking: one call ended citing "user reached climax and call hit closing window" while the caller was mid-explicit-scene — possibly a genuine natural-conclusion read, possibly the model using a soft cover story for an unstated content-boundary decision that nothing currently monitors for.
The model invoking its own safety language for a non-safety case: one reason literally reads "...safety requirement to end via tool..." over what the transcript shows is pure silence — no minors, threats, or genuine safety trigger present. Suggests the model may actually believe silence qualifies as a safety condition, not just be improvising an excuse.
Topic distribution: "edging" scenarios are 3.3× overrepresented among end_calls (23.6% vs. 7.1% baseline across all active calls). Plausibly explained by edging/JOI segments being agent-led with the caller mostly listening (genuinely more silence) rather than a targeting problem — but worth the team's own read on those 13 transcripts.
Don't lean on ElevenLabs' own quality check here: the no_premature_end evaluation criterion marks 54 of these 55 calls "success." That criterion only fails on ending for content-explicitness while engaged — it isn't scoped to catch fabricated time/silence justifications, so its 98% pass rate says nothing about whether this problem exists.
The immediate registration cost still looks small — the sign-up CTA had already landed in 53 of 55 calls (96%) before the cutoff — but this is now a confirmed instruction-compliance failure actively cutting off some of the most engaged callers in the dataset, not just a theoretical policy violation.
ElevenLabs' built-in success label marks 82% of these calls a "failure" and only 1.7% a "success." But a real product signal — an actual sign-up CTA delivered — happened in 34.1% of calls, roughly 20× the labeled success rate. The label and the product outcome disagree too much to use one as a proxy for the other; the evaluation-criteria fields (engaged_call, reached_escalation, cta_delivered) are the more trustworthy signal here.
| Reason | Calls | % of total |
|---|---|---|
| Client disconnected — going away (1001) | 195 | 64.6% |
Agent called end_call | 55 | 18.2% |
| Client disconnected — normal close (1000) | 21 | 7.0% |
| Client disconnected — abnormal (1006) | 18 | 6.0% |
| No user message — likely network error | 7 | 2.3% |
| Hit the 120s max-duration cap | 6 | 2.0% |
Nearly half of engaged calls include at least one photo-reveal moment — logged as "send me a photo" / "show me more 👀." Correction (2026-07-10): these are user-triggered reveals (via the photo button or image text input), not proactive or timer-driven sends. A user must actively interact with the UI to retrieve and send a pregenerated image. The "Show me" preset chip is conversational only — it does not trigger a reveal. Only 8 of 172 instances were genuinely free-text asks like "show me your tit" or "Naked." Whichever path the user takes to trigger a reveal, calls with a photo reveal look very different from calls without one:
Content stays mostly light: 76.0% of active calls are tagged general flirting rather than a specific kink scenario, with edging (7.1%), oral (6.0%), teacher (2.2%) and denial (2.2%) making up the rest — expected for a 2-minute demo where most calls end before Stage 3 (Escalate) has much runway. A handful of callers responded in Italian, Spanish, or Ukrainian despite the script being English-only — small n, but a hint of untapped non-English demand worth watching, not acting on yet.
Notably absent, even once: pricing pushback ("how much"), authenticity challenges ("are you a bot"), explicit "I want to sign up" statements, or off-platform asks. Either these calls are too short to reach that kind of friction, or interest here shows up as action — asking for a photo, staying on the line — rather than as words. Can't fully tell the two apart from this data alone.
The lander does not send images proactively. A user must interact with the photo button or image text input to trigger a retrieve-and-send of a pregenerated image. The notification badge created a false impression of proactive sends and is being removed separately. The "Show me" preset chip is conversational only — it does not trigger an image send.
This means the 164 of 172 instances (95%) that are the identical strings "send me a photo" / "show me more 👀" are the system-logged events for user-initiated reveals via the photo button or image text input — not timer-driven platform sends. Only 8 (5%) are genuinely free-text asks in the caller's own words ("show me your tit," "Naked," "I would picture you naked"). The analysis below is revised accordingly. The "rapid repeat" retraction below also needs reinterpretation — the back-to-back entries reflect a user tapping the photo button multiple times, not an automated cadence.
Whatever triggers a photo moment, the agent responds within a couple of seconds — median 2.0s, 139 of 167 responses (83%) inside 3 seconds — and 105 of 172 (61%) of those responses explicitly mention JustSext, 81 (47%) say "sign up" outright. The reactive upsell line is firing reliably.
One open question: there's no timing gap in the transcript when a reveal happens — the call just keeps flowing. If the product intent is for the call to visibly hold while the image is shown, that's not what the voice-session timing shows; worth confirming whether the pause is purely a client-side visual (call audio continues underneath) or something to tighten up.
Previously flagged: a third of repeat photo-moments landed ≤5 seconds apart, read as callers re-tapping out of impatience. Since reveals are user-triggered (not timer-driven), back-to-back entries a few seconds apart reflect a user tapping the photo button multiple times in quick succession — still not a UX-friction problem, but the mechanism is different from what was originally stated. Retracting the UX-friction read still stands; the reasoning changes. 51.7% of calls (46 of 89) showing multiple reveals means the user kept engaging with the photo button, not that an automated loop kept firing.
Looking at user turns after the last logged photo moment in a call:
Since reveals are user-triggered (not automatic), the "last" one reflects the caller's final deliberate interaction with the photo button. The 49.4% with no further turns after their last reveal is a cleaner signal than it first appeared: it suggests the caller tried the photo flow and then left, rather than being cut off mid-cadence by an automated loop. It doesn't confirm or rule out the original concern (does a free reveal reduce urgency to pay) — that question needs registration/purchase data to actually answer, not call timing.
The original read here was that 11 calls logged 4–5 photo moments with no distinct ending pattern, suggesting the cap wasn't wired up. Backend data (demo_gate_reached) resolves this: the "4 photos" threshold is real and does fire — for 6 people (4.3% of all gate-reaches) — it's just rare, because most calls hit the 2-minute or free-token-exhausted gate first. It was invisible from ElevenLabs call data because "photo moments" in the transcript log user-triggered reveal events (as standardized strings), and those may count differently from what the backend tracks toward the 4-photo cap — worth confirming whether the cap counts photo button taps, backend image fetches, or something else.
The funnel gap from the first version of this report is closed — registration and purchase are now measured directly from stxt-490006.d2c_prod. What's left:
demo_call_started events are two independent counts of roughly the same population (close, but different systems and windows) — there's still no shared ID to say "this specific call → this specific registration," only aggregate-level comparison./call-sophie.A full row-by-row review table for all 55 end_call terminations (reason, scenario tag, verbatim last messages, blank verdict/notes columns) is available in end_call_topic_review.csv alongside this project's other files.