Methodology paper

Turnover-Normalized Tune-Out: a behavioral meter of track-level listener holding in continuous streams, and how it stabilizes

Scoring a track against the room it plays in: construction, validation, and settling behavior of a per-track listener-holding meter.

Permalink Cite

Abstract

A raw count of listeners lost while a track plays confounds the track with its environment: continuous streams shed listeners regardless of what is on. We describe a behavioral meter that scores per-track listener holding by normalizing tune-outs against the ambient turnover of a track's on-air environment, then correcting early-play bias, shrinking thin samples toward a station prior, and calibrating to a 0 to 100 scale. Its primary use is a standing behavioral index over a settled catalog; the early-play stages exist to make a track's first rotations readable. The evidence covers about 10.9 million listener readings and 227,000 scored plays on four active stations, plus a legacy station preserved as an archived known-groups case. The meter recovers a pre-labeled off-format group it was never shown (about 8.5 points lower on the raw, never-shrunk signal, Cohen d about 0.7; +8.2, d 0.53 at production settings), cuts the daypart and neighbor confounder correlations to about zero, estimates a track-associated residual (roughly one-fifth to one-quarter of the gross tune-out magnitude), and stabilizes in a catalog-dependent way that a variance model predicted in advance.

What this means for programmers

  • It scores how a track holds its audience relative to its environment, not a raw count of who left while it played.
  • About three-quarters or more of tune-outs are the room emptying at its normal rate; the meter estimates the share associated with the track itself, roughly one-fifth to one-quarter.
  • It sees a known off-format group it was never told about (about 8.5 points lower) and is not fooled by when a track airs (after correction, scheduling explains well under 0.1% of the score on the flagship station).
  • A track's broad tier holds from about 12 plays: among tracks that went on to reach 76 plays, four in five are still in the same band 64 plays later. The score never stops moving, by design, and it moves as much late in a track's life as early; once the tier holds, that movement is play-to-play variation plus any real change in audience response, not the meter still converging. The middle of the catalog is tightly bunched, so fine distinctions there carry less weight than the tiers themselves.
  • For scale: a weekly callout wave of about a hundred respondents carries a computed margin of about ±5 points on a comparable 0 to 100 scale, and telling two waves apart requires a true change of roughly 10 points (Section 6). The meter's reading updates with every play, at no additional fieldwork.

1 Introduction

A program director's core question about a track is behavioral: does it hold the audience, or send it away? On a continuous stream the obvious measurement, counting how many listeners leave while the track plays, answers the wrong question. Listeners leave constantly, on schedules driven by the daypart and by ordinary attention, whatever is playing. Two tracks can lose the same number of listeners while one is genuinely driving them off and the other merely airs when the room is emptying. A raw tune-out count cannot tell them apart; it measures the room as much as the track.

This confounding is the central obstacle to a per-track behavioral score. The meter described here addresses it directly: it measures a track's tune-outs against the ambient turnover of the track's surrounding environment, so a track is credited only for the listeners it loses beyond what the room was shedding anyway. What remains is a track-associated residual, a smaller and cleaner signal than the raw count. Around that core correction sit the practical steps a live, play-by-play score needs: a correction for an early-play bias, shrinkage of thin early samples toward a station prior, and calibration to a bounded 0 to 100 scale. The meter is used in two modes: evaluating a new track over its first rotations, where the early-play corrections matter most, and, once the score has settled, as an independent behavioral index across the catalog.

The contribution is not only the instrument but its validation. We show, on a large behavioral record rather than a survey panel, that the meter recovers a group known to differ before the meter existed, that it removes named confounders, that it estimates the track-associated residual in listener terms, that a suspected early-play artifact is an artifact and is removed, and that the score's stabilization is understood well enough to have been predicted in advance. Sections 4.5 and 7 state what the meter does not measure and where it cannot resolve tracks at all.

2 Background

The unit of measurement is the listener count, polled continuously (nominally every five seconds; the recorded cadence is about six seconds, ten to twelve readings per minute, around the clock). Over the full database, archived stations included, this is about 10.9 million readings, of which roughly 26.5% record a fall in the audience from the previous reading. Radio and streaming audiences are known to be high-turnover: about 45% of US radio panel listening occasions are shorter than five minutes [10], and streaming measurement standards already discount very short sessions as non-listening [9]. That turnover is not noise to be removed; it is the baseline a per-track signal must be read against. The meter's premise is that the informative quantity is not a track's gross tune-outs but its tune-outs in excess of the ambient rate of its environment.

Two framings recur in what follows. Raw tune-out (or "gross drop") is the one-sided sum of listeners lost during a play, unchanged in meaning from a naive count. The corrected signal is the signed excess of that loss over the ambient rate of the environment, and it can be negative: a track that holds better than its surroundings has a negative excess, and about 58% of plays do. Play-level populations differ slightly across the analyses that follow, because each applies its own eligibility screen (Table 1):

Table 1. Analysis populations.

PopulationScreenPlaysWhere used
Scored plays, four active stationsaudience and duration quality gatesabout 227,000the score itself; Section 5
Daypart and neighbor analysisits own qualifying screenabout 207,000Section 4.2
Turnover decompositionturnover-scored plays, frozen data snapshot222,217Section 4.3

Note. The archived known-groups station of Section 4.1 is scored separately and is not part of these counts.

The corpus is also screened for integrity: one measured third-party channel was excluded because its reported listener counts, though high, lacked the natural turnover structure every genuine audience in the corpus displays.

3 The meter, in principle

The meter transforms a track's plays into a 0 to 100 score through a sequence of named stages. The mechanisms and their effects are given; the window over which the ambient rate is measured, the quality gates on the underlying plays, the prior strength, and the transform constants are proprietary and not disclosed.

1. Gross tune-out. For each play, the one-sided count of listeners lost during the play.

2. Turnover-floor normalization. Subtract the ambient tune-out rate of the track's surrounding on-air environment, estimated from the environment itself, so the track is charged only for its own excess. The result is signed.

3. Stabilizing transform. Map the signed excess through a monotone transform that preserves ordering across the full range, including among the best-holding plays, where a naive floor at zero would destroy the ordering.

4. Early-play bias correction. Remove a systematic offset that inflates a track's first plays (Section 4.4 shows this offset is an ambient artifact, not a real early-life effect).

5. Shrinkage toward a station prior. Pull a track with few plays toward the station's typical value, so a thin sample is not shown as an extreme [3, 4]. The pull is heaviest at the first play and eases as plays accrue.

6. Regression blend. Combine the corrected sample with a station-level regression estimate, weighted to minimize held-out error [7].

7. Calibration to 0 to 100. Map the result onto a bounded scale against the station's own catalog, so the endpoints are always populated and the score is interpretable as a within-station standing.

The paper's claims concern the behavior of this pipeline: that stage 2 removes the named confounders, that stages 4 to 6 are conservative and understood, and that the calibrated output stabilizes predictably.

4 Validation

Statistical conventions used throughout: bootstrap confidence intervals are two-sided 90% percentile intervals from 1,000 resamples drawn over tracks, and their bounds correspond to 5% one-sided limits for the directional claims they accompany; the settle-symmetry differences in Appendix B.4 are reported with conventional 95% intervals. Group comparisons use Welch's unequal-variance t with Cohen d on the pooled standard deviation. No multiple-comparison correction is applied; results are read as a converging pattern, not isolated tests. Resampling is over tracks within a station; stations are analyzed separately and are not treated as independent replicates of one another.

4.1 Known-groups: the meter recovers a group it was never told about

Known-groups validation carries particular weight for a behavioral meter, because the group label cannot leak into a measurement that is blind to it [2]. On a legacy station whose archived schedule provides a known-groups case, the programmer placed one deliberate off-format slot, a slow ballad at the end of each hour, in an otherwise uptempo format whose audience was understood to reject it. The group was labeled before the meter existed, from the station's own programming labels. The setting is also the meter's normal operating regime: a settled catalog with long airplay histories, not fresh tracks in their first rotations.

The meter scores the ballads lower without being told which tracks they are. At production settings the gap, defined throughout as the native-group mean minus the outsider-group mean, is +8.2 points on the 0 to 100 scale (Welch t = 5.32, df = 111, Cohen d = 0.53 [8], p < 10-6; the unit of analysis is the track: outsiders n = 82, natives n = 721). The correction is conservative: before turnover normalization the gap is +19.7 (Welch t = 11.28, df = 115, d = 1.14) and it compresses monotonically to +8.2, halving a real, pre-labeled difference without erasing it. Computed on the raw, never-shrunk signal, as validation requires, the gap is +8.53 (Welch t = 4.00, d = 0.69, 90% bootstrap CI [5.01, 12.03]; tracks with at least five qualifying plays, outsiders n = 34, natives n = 548); removing the display shrinkage raises the effect size. The bootstrap interval is for the point gap, not the effect size. Two independent labelings agree: the station's own format buckets place the native sub-styles within 2.0 points of each other with only the schedule-defined outsider dropping, and an independent third-party genre labeling shows the same ballad subset scoring 3 to 6 points below its own tag mean. A third check is acoustic rather than editorial: an audio-derived arousal value, computed from the station's own library files by a published music-emotion regressor over loudness-normalized audio [15-19] and blind to the meter, the schedule and the listeners, separates the outsider group from the native catalog at d = 1.86. That establishes the programming label as an acoustically real property of the recordings rather than a bookkeeping category; it is not evidence about the meter, which never saw the audio. Across the catalog as a whole the same audio value is only weakly associated with the score (Spearman rho = 0.11, 95% bootstrap CI [0.04, 0.18] over tracks, n = 758, measured on the raw never-shrunk signal with no play-count gate), which is the pattern expected if arousal marks a category boundary rather than a performance gradient. The slot-position confounder is ruled out: 97% of the outsider's plays air in the final 15 minutes of the hour, but the ambient tune-out rate of that slot is indistinguishable from the station average (difference under 0.1%).

Figure 1. Score distributions of the pre-labeled off-format group and the native catalog on the 0 to 100 scale, with the gap (+8.2 points, d = 0.53) annotated. The labels predate the meter, which never saw them. Source: legacy-station programming labels vs meter scores; retrospective analysis of archived data.

4.2 Confounder removal: the score is not the clock or the neighbors

The pre-specified confounder is a track's exposure to the naturally leaky overnight hours. Measured across all four active stations (track-level correlations of each track's overnight-exposure share against its value; the unit of analysis is the track, n = 1,888 / 2,504 / 352 / 592 tracks in the station order below, drawn from about 207,000 qualifying plays), the raw count correlates with overnight exposure at r = 0.22 to 0.38 on every station, while the corrected score correlates at about zero (raw to corrected: Dance UK 0.384 to -0.027, Box UK 0.220 to -0.013, Play 90s 0.308 to -0.142, The Big 80s 0.328 to 0.002). In variance terms, after correction when a track airs explains well under 0.1% of a track's score on three of the four stations, and about 2% on the fourth, where the correction slightly overshoots (the conservative direction). A second confounder, contamination from neighboring tracks, is likewise removed: raw tune-out is strongly coupled to adjacent tracks (r = 0.51 to 0.67, adjacent plays share the same ambient environment), and the corrected score's neighbor correlation collapses to zero or slightly negative, with no positive quality leak. A downstream consequence supports the removal: on a shared-anchor scale, whose endpoints are pooled across the fleet rather than set per station, the four station means compress from a 29.5-point spread (raw) into a 5-point band (corrected). Because that scale applies no per-station calibration, the compression cannot be an artifact of each station being rescaled to its own catalog. It is necessary evidence of fleet-level alignment, though it does not by itself establish per-track cross-station comparability (Section 7).

Figure 2. Daypart confounder before and after turnover correction, per station: raw correlation with overnight exposure beside the corrected near-zero values. Source: track-level correlations over about 207,000 qualifying plays, four active stations.

4.3 What the meter separates: the room versus the track

Decomposing every play's gross tune-outs into the room (the ambient rate the environment was shedding regardless) and the track-associated residual (the track's "own" signal in what follows, an association rather than a proven cause), over 222,217 turnover-scored plays, shows directly what the score captures. As a net share the room accounts for essentially all of the gross tune-out and the track's own net contribution is about zero (-2.9%), forced by the meter measuring excess over the floor. As a magnitude, however, the track's own residual is about a quarter of the raw drop: roughly three-quarters of a play's gross tune-outs are the room and one-quarter is the track's own movement, in both directions. (The 22 to 28% span is across weighting methods; the median play sits at 21.1%.) The split is stable across audience size (19.8% at 200 to 500 listeners, 20.7% above 500). This should not be read as "a quarter of listeners left because of the track"; the accurate statement is that the track's own swing is about a quarter of the raw drop in magnitude. Note that the 26.5% tune-out frequency of Section 2 and this roughly 20 to 25% own-signal magnitude are different quantities that merely happen to land near each other.

Figure 3. The room and the track: the average track's first twelve plays, gross tune-outs decomposed into the ambient room rate and the track's own residual. The own share is a magnitude, not a net cause. Source: decomposition over 222,217 turnover-scored plays, frozen data snapshot.

4.4 The early-play bias is an artifact the correction removes

A track's raw score is systematically better on its first plays, which could be a genuine novelty effect or an artifact. A forward-versus-backward mirror test separates the two: run forward on fresh raw data, the first-play offset reproduces on both calibrated stations (+12.0% and +13.4%); run backward, which serves as the null, the offset does not behave like a single early-life phenomenon. After turnover correction the offset vanishes on one station (to -2.1%, interval spanning zero) and is small with flipped sign on the other. The raw early-play effect is largely an ambient-turnover artifact, and the correction removes it; the full test, with bootstrap intervals and the backward-tail decomposition, is in Appendix B.1.

Figure 4. The raw early-play curve beside the corrected one, with the backward mirror overlaid: the forward entry offset reproduces on raw data and vanishes after correction. Source: forward-versus-backward mirror test on raw play data, two stations.

4.5 Negative controls: what the meter does not, and should not, measure

Discriminant validity requires that the meter be flat where no signal should exist. Rotation burn is undetectable: intensity, weekly trend, and intro-tune-out correlations are all about zero. At these sample sizes the analysis had the sensitivity to detect harm larger than about 0.5 points per play at the rotation levels these stations run (well below chart-radio exposure), and no effect of that magnitude was observed; a formal equivalence test was not run. Measuring burn remains survey research's ground (Section 6). A genre-pair control across five pairs, with null pairs included, separates only where a real difference exists (Appendix B.2).

5 Stabilization

"Settling" here does not mean mathematical stability. The score is never mathematically stable and is not designed to be: it deliberately keeps weighting recent plays, and the thing it measures moves in the world. The operational question is narrower, and it is the only one a programmer needs answered: after how many plays has a track's reading firmed up enough to support a confident rotation decision. About 12 plays is the ranking threshold: from that point a track's broad tier mostly holds, though its displayed value is still about 11 points from its long-run level (about 7 by play 50). This is measured directly rather than inferred from a separation statistic: of the tracks that reach 76 plays, 81% (Dance UK), 80% (Box UK) and 87% (Big 80s) are still in the same broad band at play 76 as at play 12. Early plays carry a calibrated expectation, not a forecast: a high score can only be earned from accumulated listening. About 44 plays is the raw out-of-sample point-settle, where the corrected running mean comes within 15% of its held-out asymptote on a 1/sqrt(n) curve. (The "flattens by play 10" shape a naive convergence chart draws is an arithmetic identity of averaging, not a settling number at all.) The 12-play threshold is the one the product refers to when it marks a track as settled; the point-settle is a different, later measurement of the raw running mean, not the product's marking.

The tier-persistence claim holds at the width a rotation decision is actually made at, and only at that width: cut the same catalog into thirds instead, and agreement falls to 42%, 37% and 35% against a 33% chance rate. That is the permanently tied middle of this section restated, and it is why the reportable unit is a band and not a digit.

Reliability was predicted before it was measured. It is measured model-free as a split-half correlation [5, 6] (split each track's plays at random, score each half with no shrinkage or prior, correlate the halves), so it cannot be inflated by the model agreeing with itself. A variance decomposition predicted these half correlations, per catalog, before measurement: predicted 0.375, 0.101, and about zero for the wide-spread, current-leaning, and survivorship-flat catalogs; measured 0.378, 0.108, and 0.005. Stepped up to the full track, reliability climbs with plays but never nears 1 and stays strongly catalog-dependent: about 0.49 to 0.64 on a wide-spread library, 0.16 to 0.33 on a current-leaning station, and about zero on a survivorship-flat one, at any number of plays. The model's group-level consequences and a note on prior-inflated reliability figures are in Appendix B.3.

After the point-settle the score does not freeze. A structural residual of about 8 points remains and does not shrink with more plays: the meter keeps weighting recent plays so it can follow real change. Its size was predicted from the design before measurement, and it is audience-independent (audience-to-residual correlation 0.00 across four stations). It is movement of the instrument, not of the track, and it can be shown directly: measured on the same tracks over equal-width windows, the displayed score moves as much between plays 68 and 76 as between plays 12 and 20, and no late-life slowdown is detectable on any station (the window comparisons and their intervals are in Appendix B.4). The advice to a programmer is not to wait until the score stops moving, because it never does; this is the resolution of the meter, and the resolution is the same at every age. Because this floor is about 0.6 times the between-track spread, roughly half the catalog sits within one residual of the median, permanently tied, not for lack of plays. Ties of this size are not particular to this meter: differences of that magnitude sit inside a fielded survey wave's margin as well (Section 6).

It also bounds what the movement can be. A settled track's score can move for exactly three reasons. The meter can move it, which is legitimate while new plays are arriving and is an error only if a score changes with no new play of its own; by construction it does not, since a track's value is its last play's value. The scale can move under it, because the 0 to 100 endpoints are anchored to the catalog and a better-holding track arriving changes what 100 means; this moves an idle track's displayed number without any new play of its own, and it is a property of a relative index, not a defect. Or the world can move, because a track played again months later meets an audience whose interest in it has changed. The first is ruled out without a new play, the second is a declared consequence of scoring on a within-station scale, and the third is the quantity the meter exists to track. Sustained movement after the tier holds is therefore a mix of the first and the third: each new play carries the structural residual of Section 4.5, and any real change in a track's standing with the audience rides on top of it. What it is not is the instrument failing to converge.

Figure 5. Displayed-scale spread by play: outcome bands start within about 7 points at play 1 and fan to about 32 (Dance UK) and 37 (Box UK) by play 12; the typical band stays flat, the permanently tied middle. Source: stored per-play score curves, heavily played tracks on two production catalogs.

Extreme-band reproducibility, and what a band is worth. Reliability states how much of the score is signal; it does not answer whether a track flagged at one end of the catalog stays there. On the same prior-free split-half protocol, tracks that one half of the data places in the worst-holding tenth reappear in the worst tenth of the other half 32 to 35% of the time against a 10% chance rate where the catalog has spread, and no better than chance on the survivorship-flat one. The recovery is symmetric: the best band is identified no worse than the worst, so the meter carries no special sensitivity to poor performers. The meter's contribution is to direct the individual rotation decisions where the differences are; whether corrected rotations compound into a larger station-level TSL effect is the outcome study of Section 6. All of these figures are measured out of sample deliberately, because ranking tracks and costing the list on the same plays inflates a worst-list deficit by 55 to 104% (Section 7); the discrimination statistics and the listener-unit accounting are in Appendix B.5.

6 Discussion

The meter occupies a specific position in radio music research: it measures behavioral hold, a construct the industry's survey instruments do not observe. Callout asks a recruited sample whether they like a track and whether they have tired of it; the meter asks what the actual audience did, reading by reading at the polling cadence, once the track's environment is accounted for. The two measure different things, and each covers ground the other cannot: survey research can test unaired candidates from hooks, resolves demographics, and measures burn directly (a dimension Section 4.5 shows the meter does not detect at these exposure levels), while the meter observes behavior at a cadence and a marginal cost no fielded survey can approach. A formal method-agreement comparison between the meter and a survey instrument on shared tracks [1] has not been run.

The magnitudes of Section 5 should be read against what those standard instruments deliver, for scale rather than criticism. Published practitioner guidance commonly illustrates callout waves with samples of roughly 100 to 150 respondents [11]; at those sizes a single song score on a 0 to 100 scale carries a computed 95% margin of about ±4 to 6 points, and distinguishing two waves statistically requires a true change of roughly 10 points (computed from standard rating-scale dispersions [12]; to our knowledge no callout or music-test vendor publishes a per-song margin of error or a test-retest figure). Practitioners already work at this resolution: programming guidance reads weekly callout on trends rather than single readings [13], and one music-testing firm reports seeing a 20 to 30% change among a station's top 200 songs from one library test to the next [14]. Against these anchors the meter's structural residual of about 8 points is of the same order as a survey wave's computed margin, and the tier-persistence figures of Section 5 are, to our knowledge, the only published stability numbers for a per-song radio music metric of any kind. The comparison is offered for scale, not superiority: the instruments measure different constructs, on different economics. A callout program prices its fieldwork per wave; the meter's readings accrue with every play, at no marginal cost, on the station's own audience rather than a recruited sample.

Steady state is the meter's primary regime: most of a catalog has aired for months or years, and the strongest validation evidence, the known-groups recovery of Section 4.1 and the reliability profile of Section 5, is settled-catalog evidence. The meter's natural outputs are the ones a programmer can act on: which tracks hold better or worse than their room, which tier a track is in, and whether a settled track's recent movement is real change or the structural residual of Section 5. Because the confounder removal substantially improves fleet-level alignment of station means, the design goal is one metric across a network: the same track's holding read, not inferred, on every station that plays it; per-track cross-station validation remains outstanding (below). The design choices most open to question, the shrinkage and the blend, are conservative by construction and, in the known-groups test, err toward understating a real difference rather than inventing one.

A note on the score's name and its construct. The product displays the score as a POP Score, but the construct is not popularity in the chart or fashion sense: it is listener holding, and for a music director it is the more actionable of the two. Every tune-out ends a listening occasion, so the score estimates the track-associated component of time spent listening (TSL) relative to ambient turnover. Whether score-guided rotation changes raise station-level TSL is an outcome study we have not run.

7 Limitations

8 Conclusion

Measuring how a track holds listeners on a continuous stream requires separating the track from its environment. Normalizing tune-outs against ambient turnover substantially reduces the measured confounding: the meter recovers a pre-labeled off-format group, its confounder correlations fall to about zero, and its stabilization follows a variance model specified in advance. It separates the ends of a catalog about three times better than chance where the catalog has real spread and not at all where it does not; the tightly packed middle is consistent with restriction of range in these catalogs combined with the meter's finite resolution, and differences of under one listener departure per week sit inside a fielded survey wave's margin as well. What distinguishes the meter among the industry's music-research instruments is not freedom from uncertainty but its treatment: the margins, reliability and stability reported here are measured, published, and updated with every play the audience hears.


Appendix A. Description of methodology: disclosed effects, reserved parameters

This appendix names each correction, states its purpose and the magnitude of its effect, and reserves the numeric internals, in the manner of an audience-measurement currency's public description of methodology.

measurement
Listener counts are polled continuously (nominally every five seconds; recorded cadence about six seconds), and per-minute aggregates feed the pipeline. The exact within-minute count (ten, eleven, or twelve readings) does not affect the metric, which aggregates per minute; denser sampling would not make it more accurate.
turnover-floor normalization
The ambient tune-out rate is estimated from the track's surrounding on-air environment and subtracted from its gross tune-out. The window length and the floor's cap are reserved; the effect (daypart correlation: raw 0.22 to 0.38, corrected about zero; neighbor coupling: raw 0.51 to 0.67, corrected about zero; own residual roughly one-fifth to one-quarter of gross) is disclosed.
quality gates
Plays below audience and duration thresholds are excluded to keep the floor well estimated. The thresholds are reserved; they are set as noise safety-nets, not functional cutoffs, and touch no ordinary play. Population counts per analysis are given in Section 2.
outlier and outage handling
A single polling interval's tune-out is capped relative to the prevailing audience, and a play's total tune-out is capped relative to its audience and length; the cap thresholds are reserved, and the smallest audiences are left outside the per-play cap. A sudden audience surge, and its settling back toward the prior level, is filtered so that the settling is not recorded as tune-out. A gap in the polling stream resets the ambient estimate, so no tune-out is attributed across an outage; a play whose surrounding window falls below a minimum of covered minutes receives no score at all.
stabilizing transform
A monotone transform on the signed excess, chosen to preserve ordering across the range. Its constants are reserved.
early-play bias correction
A decaying offset removed from a track's first plays; validated as an ambient artifact by the forward-backward mirror test (Section 4.4).
shrinkage and blend
A track's thin early sample is pulled toward the station prior and blended with a station-level regression, with weight easing as plays accrue; the prior strength and blend weight are reserved, the effect (an about 11-point gap at play 12, closing to about 7 by play 50) is disclosed.
calibration
The result maps onto 0 to 100 against the station's own catalog percentiles; the bounds are computed from winsorized percentile estimates, with the winsorization limits reserved.
computational traceability and internal reproducibility
Every reported number is regenerated by a named internal diagnostic; the known-groups and confounder analyses are read-only measurements against archived data. This is internal reproducibility, not independent replication, since the parameters and raw data are not public.
scope of the score
The score is a within-station standing on a behavioral signal, not a precise headcount of who left because of a track, and not a survey of preference; the middle of the catalog is a statistical tie and should be read as such.

Appendix B. Extended results

Each entry states the question, the method, and the result, so a technically oriented reader can verify that no main-text result rests on a single analysis. The main text is meant to stand without this appendix.

B.1 Early-play mirror test (Section 4.4). Question: is the raw first-play advantage a real early-life effect or an ambient artifact? Method: run the play-index curve forward on fresh raw data, and backward as the null. Result: forward, the first-play offset reproduces (Dance UK +12.0%, 90% bootstrap CI [8.0, 22.4]; Box UK +13.4%, 90% CI [5.2, 25.8]); backward, one station's curve is flat (an entry-only effect) while the other's shows a non-flat tail (a whole-lifecycle trend), so the raw "early illusion" is not a single early-life phenomenon. After turnover correction, on the station with the lifecycle tail both the entry offset and the tail deviation vanish (+12.0% to -2.1%, 90% bootstrap CI [-4.9, 0.8], spanning zero); on the other the entry residual is small with flipped sign. The offset behaves as an ambient-turnover artifact.

B.2 Genre-pair negative control (Section 4.5). Question: does the meter separate genres only where a real difference exists? Method: five pre-specified genre pairs across four stations, null pairs included, compared on the corrected signal with analytic 95% intervals. Result: only the pairs with a real, externally grounded difference separate; the null pairs do not. Fig. B1 is the forest plot.

Figure B1. Genre-pair forest plot with confidence intervals: pairs that should differ land off zero, null pairs cross it. Source: five genre pairs across four stations, null pairs included.

B.3 Reliability model detail (Section 5). Question: what does the variance model imply beyond the split-half check? Method: the same decomposition, stepped to group level. Result: the play count at which a group's ordering is expected to become reliable is about 22 plays for a distinct set of tracks and about 348 for a tightly clustered one. A reliability computed on the displayed, shrunk score can read near 0.6, but that figure is inflated by both halves sharing the prior and holds only for the wide-spread library; the main text reports the prior-free figure throughout.

B.4 Settle symmetry (Section 5). Question: does the score's movement slow down late in a track's life? Method: take only tracks that reach 76 plays, so the same tracks appear in every comparison, and measure displayed-score movement over equal-width early and late windows. Result: over an eight-play window the movement is the same at both ends (7.5 points at plays 12 to 20 against 7.2 at plays 68 to 76 on Dance UK, 7.6 against 7.6 on Box UK, 7.4 against 6.3 on Big 80s; paired late-minus-early differences -0.33 points, 95% CI -1.24 to +0.66; -0.04, -1.75 to +1.66; -1.18, -3.50 to +1.20). Over a thirty-two-play window the movement falls by about a quarter (13.4 to 10.5, 14.8 to 11.8, 14.3 to 11.1). A track's distance to its long-run level shrinks early, but its movement from one handful of plays to the next does not shrink at all: the meter's resolution is the same at every age. Real external drift is separate and slow, with a half-life beyond 200 days on the largest station.

B.5 Extreme-band discrimination in listener units (Section 5). Question: what is a band worth in listeners, and how much does a ranked list flatter itself? Method: the prior-free split-half band recovery of Section 5, expressed as discrimination statistics and in departures per week, with ranked lists re-costed on held-out plays. Result: worst-tenth recovery is 32 to 35% against a 10% chance rate (AUC 0.67 to 0.74) on the catalogs with spread. The extreme-to-extreme difference in departures per week is five to twelve times the middle-to-middle one, and adjacent tracks in the middle differ by under one listener departure per week, so the middle of a catalog is not merely hard to resolve but nearly indifferent to resolve. In absolute terms the worst tenth accounts for about half a percent of during-music departures on the wide-spread library (counting only the immediate, during-play departures it explains). Costing a ranked worst-list on the plays that ranked it inflates its deficit by 55 to 104%, the share rising as reliability falls, because selection captures noise as well as signal and the luck does not repeat.

Disclosure

Author contributions. Airplay Live Research was responsible for the conceptualization, methodology, software development, data curation, formal analysis, validation, visualization, and preparation of the manuscript.

Competing interests. The meter is a commercial product of Airplay Live. The research was conducted by the founder and developer of Airplay Live, who has a commercial interest in the methodology described. No external funding was received. The validation design relies on pre-labeled groups, negative controls, and out-of-sample tests partly for that reason.

Ethics and privacy. The measurement consists of aggregate concurrent listener counts read from the stations' streaming servers. No individual listener is identified or identifiable: the data contain no user identifiers, demographic attributes, or per-person listening histories, and no intervention in any station's programming was made for this study. The analyses are retrospective computations over these aggregate operational records; on that basis, the work was not treated as human-subjects research and no ethics review was sought.

Data availability. The underlying listener logs are operational data of the measured stations and are not public. Every reported number is regenerated by a named internal diagnostic against frozen data snapshots (Appendix A), and the published figures carry the aggregate data behind each claim. Requests for closed review access can be made via the correspondence address.

References

Statistical and measurement foundations (1-8); industry standards and practitioner context (9-14); audio analysis and music emotion recognition (15-19).

Statistical and measurement foundations

  1. Bland, J. M., & Altman, D. G. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 327(8476), 307-310.
  2. Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302.
  3. James, W., & Stein, C. (1961). Estimation with quadratic loss. Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, 1, 361-379.
  4. Efron, B., & Morris, C. (1977). Stein's paradox in statistics. Scientific American, 236(5), 119-127.
  5. Spearman, C. (1910). Correlation calculated from faulty data. British Journal of Psychology, 3(3), 271-295.
  6. Brown, W. (1910). Some experimental results in the correlation of mental abilities. British Journal of Psychology, 3(3), 296-322.
  7. Bates, J. M., & Granger, C. W. J. (1969). The combination of forecasts. Operational Research Quarterly, 20(4), 451-468.
  8. Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.

Industry standards and practitioner context

  1. Triton Digital. Description of Methodology for Webcast Metrics and Webcast Metrics Local (MRC-accredited). https://help.tritondigital.com/docs/description-of-methodology-wcm-wcml-overview (accessed 2026-08-05).
  2. Inside Audio Marketing. (2024, November 13). What Nielsen's three-minute listening qualifier means for radio. https://www.insideaudiomarketing.com/post/what-nielsen-s-three-minute-listening-qualifier-means-for-radio (accessed 2026-08-05).
  3. Ryan, S. (2016, December 8). Callout research tips (part 1): Music core matters. Broadcast Dialogue. https://broadcastdialogue.com/callout-research-tips-part-1-music-core-matters/ (accessed 2026-08-06).
  4. Sauro, J., & Lewis, J. R. (2023, April 25). What are typical standard deviations for rating scales? MeasuringU. https://measuringu.com/rating-scale-standard-deviations/ (accessed 2026-08-06).
  5. Benson, K. (2018, April 22). The BURN factor: How much is too much? P1 Media Group. https://p1mediagroup.com/the-burn-factor/ (accessed 2026-08-06).
  6. Hollander, D. (2024, May). Key timing considerations for more effective music testing. MusicMaster newsletter. https://musicmaster.com/newsletters/0524.php (accessed 2026-08-06).

Audio analysis and music emotion recognition

  1. Bogdanov, D., Wack, N., Gómez, E., Gulati, S., Herrera, P., Mayor, O., Roma, G., Salamon, J., Zapata, J. R., & Serra, X. (2013). Essentia: An audio analysis library for music information retrieval. In Proceedings of the 14th International Society for Music Information Retrieval Conference (ISMIR 2013) (pp. 493-498). Curitiba, Brazil.
  2. Alonso-Jiménez, P., Bogdanov, D., Pons, J., & Serra, X. (2020). Tensorflow audio models in Essentia. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 266-270). IEEE. https://doi.org/10.1109/ICASSP40776.2020.9054688
  3. Pons, J., & Serra, X. (2019). musicnn: Pre-trained convolutional neural networks for music audio tagging (arXiv:1909.06654). arXiv. https://doi.org/10.48550/arXiv.1909.06654
  4. Soleymani, M., Caro, M. N., Schmidt, E. M., Sha, C.-Y., & Yang, Y.-H. (2013). 1000 songs for emotional analysis of music. In Proceedings of the 2nd ACM International Workshop on Crowdsourcing for Multimedia (CrowdMM '13) (pp. 1-6). ACM. https://doi.org/10.1145/2506364.2506365
  5. Aljanaki, A., Yang, Y.-H., & Soleymani, M. (2017). Developing a benchmark for emotional analysis of music. PLoS ONE, 12(3), e0173392. https://doi.org/10.1371/journal.pone.0173392

Cite this work

APA

Airplay Live Research. (2026). Turnover-normalized tune-out: A behavioral meter of track-level listener holding in continuous streams, and how it stabilizes (Version 2.4). Airplay Live. https://www.airplaylive.com/research/

BibTeX

@techreport{airplaylive2026meter, author = {{Airplay Live Research}}, title = {Turnover-Normalized Tune-Out: a behavioral meter of track-level listener holding in continuous streams, and how it stabilizes}, institution = {Airplay Live}, year = {2026}, note = {Version 2.4}, url = {https://www.airplaylive.com/research/} }

Version history

Licensed CC BY 4.0.

Correspondence: research@airplaylive.com