Documentation

Walkout

Every platform knows when viewers stop watching. None of them know why. Walkout is an agent that answers the second question, by putting two witnesses in the same room: the telemetry, which saw the audience but never the film, and a model that watches the film but has never seen the audience.

Or start here: curl https://walkout-production-f914.up.railway.app/api/retention/sintel

What Walkout is

A retention chart tells you that 18% of your audience left at 04:08. It does not tell you what happened at 04:08. Today the answer is an analyst scrubbing to a timestamp and guessing, and the guess is often wrong, because two completely different failures make the same shape on a chart:

The scene isn't working

A sequence drags, a subplot loses people, a joke lands wrong. The fix is an edit. Everyone leaves at roughly the same rate, and playback is perfectly healthy.

The stream is broken

One app build stutters, one CDN region degrades. The fix is an engineering ticket. The film is fine. Recutting it would be worse than doing nothing.

Get that backwards and you either recut a scene that was fine or ship a broken stream to everyone. Walkout tells them apart, and says which one it is in plain language, addressed to whoever owns the fix.

Who it is for

A company that streams video and owns its viewing data: a streaming service, a sports app, a broadcaster's catch-up platform, a training or course platform. The person using it sits on the content or operations team.

Their job today is to look at a chart, scrub to a timestamp, and form an opinion. Walkout replaces the opinion with evidence, and does it for every cliff in the title at once rather than the one someone happened to notice.

What it needs to run

Two things, and only two.

  1. Playback telemetry. The stream of events a player already emits (started, still watching at 04:08, buffered, quit), one row per viewer per ten seconds. This is the same data behind any retention chart, so a platform that can draw one already has it. Walkout reads it from ClickHouse.
  2. A link to the film. A URL the model can watch. In this demo that is a YouTube link; the film is never copied, hosted, or uploaded anywhere.

The telemetry is the part that cannot be substituted. Where an audience left is not information that exists in the picture; it exists only in the audience.

What it cannot do

It cannot take a bare video URL and tell you where viewers stopped watching. This comes up often enough to state plainly.

Finding a walk-out requires knowing what millions of viewers did. A model watching a film can tell you a scene is slow; it cannot tell you that 6,602 people quit during it, or that 4,719 of them were on one Android build. Nothing in the footage records that. If someone offers you a tool that finds retention cliffs from a URL alone, it is guessing which parts of the video seem boring, which is a different product and a much weaker one.

What Walkout does with the URL is the other half: once the telemetry has named the exact seconds worth looking at, the model goes and looks at them.

How it works

  1. Find where they left. Every play, pause and quit from every viewer, read as a survival curve. For each ten-second bucket Walkout measures the share of people still watching who quit there, and compares it against what this title normally loses. Every film has a slow leak; only the cliffs are interesting. Someone who watched to the end is not counted as having left, and the end credits are excluded entirely, because an audience leaving as the credits roll has finished the film.
  2. Work out who left. The same moment is cut nine ways: phone, TV, app build, delivery network, country, interface language, audio language, subtitle availability, and whether this was the viewer's first time. The question is not "how many left" but "was any group far more likely to leave here than the audience as a whole".
  3. Watch the moment. Gemini watches those exact seconds of the real film and reports what is happening: a synopsis, a beat-by-beat with timecodes, how it is paced, whether anyone is speaking, whether the picture itself looks broken. It is never told that anyone walked out there, which is the whole point: when it independently flags the same moment, that is corroboration rather than an echo.
  4. Say what to fix, and to whom. The numbers and the footage are reconciled into one verdict with a named owner, the evidence that decided it, and an estimate of the watch hours it would win back.

How it decides

Two measurements do most of the work, and both come out of the warehouse rather than out of a model. This matters: a model asked to adjudicate evidence is doing something it is good at, while a model asked to guess a buffering rate is not.

Concentration

How much likelier one group was to leave in this window, compared with everyone else in the same window. If every device, build, country and language quit at the same rate, the problem is on screen. If it is four times worse on one app build, the problem is whatever that build is doing.

Rebuffer lift

How much worse buffering was here than it was for the same group across the rest of the film. The comparison is what makes it mean anything. A group that buffers everywhere has a platform problem, not a cliff. Only a group that buffers here specifically explains people leaving here specifically.

What comes out

What the evidence looks likeVerdictWho owns it
Concentrated on one build or device and buffering far above that group's own baseline Technical Streaming engineer. Walkout will say in as many words do not touch the edit.
Concentrated on viewers whose interface language differs from the audio, with no subtitle track, and playback clean Localization Localization manager. A missing subtitle track, not a bad scene.
Flat across every group, playback clean, and the footage reads as slow Pacing Editor. The numbers ruled everything else out and the picture agrees.
Flat across every group, playback clean, and the footage reads as fine Unknown Nobody yet, and it says so rather than inventing a cause.

That last row is deliberate. Telemetry can prove a delivery failure and it can prove an audience skew. It cannot tell a boring scene from a well-earned quiet moment, so on its own it returns unknown and waits for the footage. A system that always produces a confident cause is not more useful; it is just harder to catch being wrong.

Reading the console

The page is one investigation, told in three steps.

1

Where they left

The dark line is how much of the audience is still watching. The filled area is the quit rate: out of everyone still watching, the share who stopped in each ten seconds. Shaded bands are the cliffs Walkout considers significant, which means they cleared both a size threshold and a statistical one, so a small spike on a handful of viewers does not qualify.

The credits are cut off the right-hand side on purpose. Leaving there means you finished, and including it would bury three real findings under one meaningless one.

2

What was on screen

The film, with a jump button per cliff. Press one and the player seeks to that exact second so you can watch what your audience was watching when they left.

Watch this moment sends that window to Gemini, which reads it blind and reports back. Windows are capped at sixty seconds, because reading video is the expensive call, and this keeps a curious visitor from spending the day's budget in one click.

The three chips are the cliffs the analysis found, which is the right default: they are the moments worth looking at. Custom is there because the claim being made is that the model reads any window of this film, and a demo that only ever reads three prepared moments invites the obvious suspicion about those three. Scrub the player anywhere, press Use playhead, pick a length, and it will read a moment nobody chose in advance.

3

Why they left

One card per cliff: the timecode, how much worse than normal it was, the verdict, and the watch hours at stake. Open a card for the evidence in sentences: which group over-indexed, by how much, and what playback was doing.

On the right, the agent works through all of it by itself and writes the summary.

The agent

Everything above can be read off the page by hand. The agent does it unprompted: it finds the cliffs, investigates each one, watches each one, and writes up what to fix, ranked by what it would win back. It takes about half a minute and roughly seven model calls.

It has three purpose-built tools (find the walk-outs, investigate one, watch one) and direct read-only access to the warehouse through the official ClickHouse MCP server, for follow-up questions the fixed tools do not answer. Every number in its write-up came from a tool call; it is instructed never to estimate one and never to invent a timecode.

Because a full run costs real quota, finished investigations are stored and replayed on the next visit, timestamped, with the button offering a fresh run. If a run is cut short, it is kept too, but flagged, and the page says it was cut short. A truncated report passed off as a finished one would be worse than an empty panel.

The demo data

A public dataset of scene-level walk-outs does not exist, so this demo generates one: 13,065,665 playback events across 250,000 viewing sessions of Sintel (Blender Foundation, CC-BY 3.0), a twelve-minute film.

Crucially, the cliffs are planted at known timecodes for known reasons. That is the only way to measure whether a diagnosis is actually right rather than merely plausible.

WindowPlanted causeWho it hit
03:40–04:10Story / pacingEveryone, playback clean
09:20–10:00TechnicalAndroid players on build 4.2.1, buffering
02:10–02:40LocalizationNon-English locales with no subtitle track
10:20–10:40DecoyMild, universal, below the significance floor, must never be reported

The timecodes are not arbitrary. Each was chosen by asking Gemini to survey the real footage first, so the telemetry and the film agree. The pacing cliff sits on the genuinely slowest passage: Sintel going to bed, static shots, no score. The localization cliff sits on the shaman scene, where the plot goal is established entirely through spoken English. An earlier draft planted the story cliff over the dragon chase, which would have produced a demo where the numbers claim a scene drags while the picture shows a kinetic action sequence.

The grader (make eval) runs the whole pipeline against this ground truth and reports pass or fail per cliff. It currently passes all four, including correctly ignoring the decoy.

What it is built on

AgentGoogle Agent Development Kit, LlmAgent with an MCPToolset
ReasoningGemini 3.5 Flash
VideoGemini 3.6 Flash, agentic video understanding
WarehouseClickHouse Cloud, read through the official mcp-clickhouse server
APIFastAPI, server-sent events for the agent stream
ServingDocker on Railway

Reasoning and video run on separate models on purpose. They want different things: one is a fast reasoner making a dozen short calls, the other reads video once and carefully. Request quota is counted per model, so separating them also stops a long investigation from starving its own video reads.

No non-Google AI model, API or agent framework is used anywhere in this project. Tool orchestration is the ADK's own MCPToolset.

The API

Everything the console does is a plain HTTP call, so anything on this page can be driven from a script, a notebook, or your own front end. There is no key and no session; the demo is read-only.

Open the interactive API reference Every endpoint with its parameters, response shape, and a Try it button that runs against this live deployment. →
EndpointWhat it returns
GET /api/titles Everything in the warehouse: runtime, where the credits start, the video link.
GET /api/retention/{title_id} The survival curve, the cliffs found on it, and the watch hours each one costs.
GET /api/investigate/{title_id} The telemetry evidence for every cliff at once: cohort breakdowns, buffering comparison, and the cause the numbers alone support. No model call, so this is fast and free.
GET /api/watch/{title_id}
?start=&end=
Gemini reads that window of the film blind: synopsis, beats with timecodes, pacing, dialogue, visible artefacts. Windows are capped at sixty seconds.
GET /api/agent/{title_id} Runs the full agent, streaming every tool call and every token as server-sent events.
GET /api/agent/{title_id}/last The most recent completed investigation, replayed from ClickHouse. Costs nothing, which is why the page opens with it.
GET /api/health Runs a real count() against ClickHouse rather than just proving the process is alive.

Try it without leaving the page: curl https://walkout-production-f914.up.railway.app/api/health

Running it yourself

The repository is public and Apache-2.0.

git clone https://github.com/martinvibes/walkout
cd walkout

make install        # virtualenv and the package
make mcp-server     # the ClickHouse MCP server, in its own environment
cp .env.example .env

make doctor         # checks the connection before anything long runs
make load           # applies the schema; safe to repeat, deletes nothing
make simulate       # 13.1M events across 250k sessions, about four minutes

make serve          # the console on http://127.0.0.1:8000
make agent          # the same investigation, in the terminal
make eval           # grade the diagnosis against the planted ground truth

The MCP server lives in its own environment because it needs version 2 of the MCP protocol library while the ADK needs version 1. They are separate processes talking over stdin and stdout, so the conflict only exists if you insist on one interpreter for both.

Limits and honesty

  • The demo data is simulated. Real telemetry is nobody's to publish. The simulation models hazard rates, buffering, device mix and locale mix, and the cliffs are planted deliberately, which makes the diagnosis gradeable but does not prove the thresholds are right for your catalogue.
  • One title. The method is per-title by construction; nothing here has been tested across a catalogue.
  • Quota. The public demo runs on a free-tier key allowing twenty requests per day per model. A full investigation costs about seven of them, so the daily budget is roughly two complete runs shared by everyone who visits. This is why finished investigations are replayed rather than re-run: opening this page costs nothing, and only the buttons spend anything.
  • If the quota does run out, the page says so in those words and keeps working. The retention curve, the cliffs, the cohort evidence and the stored investigation are all ClickHouse queries and none of them touch a model, so everything except a fresh run is unaffected. Reasoning and video also draw from separate pools, so a long investigation cannot starve its own video reads.
  • "Unknown" is a real answer and you will see it. When playback is clean and no group over-indexed, the numbers have ruled out delivery and availability and nothing more. The footage decides the rest.

Questions

Can I paste a YouTube link and have it tell me where people stopped watching?

No, and neither can anything else. That would require knowing what the audience did, and a video URL contains the film, not the audience. See what it cannot do.

Could a YouTube creator use this on their own videos?

Partly. YouTube Studio gives creators a retention curve for videos they own, so the cliffs and the footage reading would work. But YouTube only exposes an aggregate curve (no device, app build, country or subtitle data), so the story-versus-delivery split, which is the whole point, would not.

Why not just ask a model to watch the film and find the boring parts?

Because it would be guessing at what seems boring rather than measuring what actually lost people, and it would be wrong in exactly the expensive direction: a buffering failure looks perfectly fine on screen. Walkout's technical cliff is a dragon fight: frantic, well-cut, no visible artefact at all. A model watching that window alone would report a great scene. The audience left because the stream broke.

Does the model see the numbers before it watches?

No. That is deliberate and it is what makes agreement meaningful. If you tell a model that viewers left at 03:40 and then ask what is wrong at 03:40, it will find something. Asked cold, it either flags the moment on its own or it does not.

How is "watch hours recoverable" calculated?

The excess viewers who left at that cliff (the ones above what the title normally loses over the same span), multiplied by the runtime still ahead of them, adjusted for how many would have finished anyway. It is an upper bound on what fixing that one moment could win back, and it is what the findings are ranked by.