Building Smart Playlists Without an LLM
September 21, 2026


“Make me a two-hour playlist for Pilates” sounds like a Tuesday-afternoon feature. Filter by tempo, string some songs together, done by lunch. We spent a couple of weeks on it, and almost none of that time went into the part we expected to be hard — the AI. It went into deciding where a measurement should run, what a “smart playlist” actually is under the hood, and — eventually — two bugs that only showed up once real data hit a real database.
This is the version of that story with the wrong turns left in, because the wrong turns are the useful part.
What “smart” actually requires
A playlist app can’t filter by tempo or energy if nothing has ever measured them. Song files carry a title and an artist; they don’t carry a BPM or a loudness curve. MusicBrainz — the usual place to look for music metadata — doesn’t have this either; it’s a database of identity (who recorded what, on which release), not of signal. AcousticBrainz used to have it, stopped collecting years ago, and its confident-sounding “mood” and “genre” tags were trained on datasets too small to trust anyway.
So the numbers have to come from decoding the actual audio. We looked at doing this with one existing library instead of building it ourselves, and struck out twice for different reasons. One candidate returns an opaque similarity vector — great for “play me something like this,” useless for “give me something between 95 and 125 BPM.” The other, the one every AcousticBrainz-style pipeline actually used, is AGPL-licensed — which matters specifically because we don’t ship this software for someone to run themselves; we operate it as the service behind the app, and AGPL’s obligations are written for exactly that shape of deployment.
What we settled on is boring on purpose: decode with ffmpeg, take loudness from its EBU R128 filter, take tempo from aubio, and compute the rest — spectral centroid, onset rate, dynamic range — directly over the decoded samples. A few hundred lines over an FFT, not a dependency. One extractor, one version number stamped on every track it measures.
Where should the measuring happen?
Our first instinct was to measure on upload, client-side. The desktop apps already touch the file on its way to storage; decoding it once more while it’s sitting right there costs nothing, and it means the server never has to fetch a library’s worth of audio back down just to look at it.
We reversed that, and the reason isn’t obvious until you hit it. “Energy” isn’t a fixed number — it’s a percentile computed across a whole library, because there’s no universal ground truth for “high-energy” and there doesn’t need to be; the scale only has to be consistent within one collection. That holds as long as one thing measures every track. It stops holding the moment two apps, on two platforms, each run their own copy of the extractor. Nothing crashes. The scale just quietly stops meaning anything, at whatever pace the two codebases drift — and there’s no log line for “your energy filter is now subtly wrong.”
So: one extractor, one place, server-side. A household opts in explicitly, since measuring a whole library is real CPU and real storage traffic, not something that should happen silently. Once it’s on, the most-played tracks get measured first, so the feature is useful at five percent progress instead of waiting on a hundred.
A playlist is a query, not a list
The part we were most stubborn about: nothing in this system is allowed to hand back a list of track IDs directly. A playlist request compiles into a small, closed rule instead — something closer to this, simplified:
{
"filters": { "bpm": [95, 125], "energy": [0.35, 0.7] },
"limit": { "kind": "duration", "minutes": 120 },
"arc": "warmup-sustain-cooldown",
"sort": "weighted-random"
}
The rule is what gets stored, shared, and re-run — not its result. Three reasons, and any one of them would have been enough on its own. A generator asked to name actual tracks will invent ones that aren’t in your library — this is exactly the failure mode people mean by “hallucination,” and it’s much easier to produce than to notice. A library of any real size doesn’t fit in a context window, so something retrieval-shaped has to happen regardless of how clever the front end is. And when a playlist looks wrong, a rule is something you can read, argue with, and fix; a list of forty UUIDs is not.
Skip isn’t dislike
A family library complicates a part of this a personal one never has to deal with: the taste doing the filtering isn’t one person’s. A household playlist that’s actually just whoever plays music the most, wearing everyone else’s name, isn’t personalization — it’s noise with good production values.
So preference is tracked per person, per track, as two signals that are deliberately kept apart. A skip decays — a half-life, not a strike. Skip something once on a bad morning and its weight recovers over the following weeks; skip it constantly and it quietly falls out of rotation without anyone having declared war on it. An explicit “don’t play this to me again” is a different, much slower-decaying signal, and it stays separate on purpose: a skip is a noisy, contextual thing — wrong mood, mid-conversation, wrong moment — and a dislike isn’t. Collapse the two into one counter and you get the annoying behavior everyone’s hit on some other app: skip a track once because your phone rang, and it vanishes from rotation for a month as if you’d banned it.
Play counts pull double duty. They’re what orders the opt-in measurement pass — the tracks a household actually plays are the first ones measured, mentioned earlier — and they weight which candidates get picked during weighted-random selection, so a genuine favorite surfaces more often than a plain filter would produce, without turning into a fixed top-ten that never changes.
The part that actually does the work
Once you have a set of candidates matching a rule, turning that into something that feels like a playlist rather than a filtered SQL result is a sequencing problem, and it’s almost entirely unglamorous. No artist twice in a row. A cap per artist, so one prolific band doesn’t swallow a small library’s worth of “energetic.” Weighted-random selection instead of a deterministic top-N, because the same ranked list run twice should not produce the same playlist twice — that’s not smart, that’s a cache. And a target energy arc across the whole request, not a flat band: a two-hour session should spend ten minutes warming up, most of the middle sustained, and the last stretch coming down, which is what actually makes “for Pilates” mean something instead of just “songs in a tempo range, shuffled.”
None of this needed a model. It needed someone to sit down and write the constraint-satisfaction loop carefully, which is a less exciting sentence than “we used AI,” and also the reason the feature works at all.
We tried the LLM anyway. It lost.
The genuinely tempting version of “ask for a playlist in plain English” is an on-device model that turns your sentence directly into the rule above. We tried it — a real device, a real prompt, the real vocabulary of genres in our own test library handed over as context, so it had no excuse to guess.
Two failures were enough to kill it as the default. It invented genre names that weren’t anywhere in the vocabulary we gave it — twice, in a ten-prompt run. And it correctly pulled a stated duration out of an English sentence, then missed the same information in an otherwise identical Czech one. What replaced it is almost insultingly plain: a deterministic parser that matches a handful of activities — pilates, a workout, winding down — against durations spelled out in two languages, and returns a clean “I couldn’t make sense of that” instead of a confident guess when it doesn’t recognize the request. It has never invented a genre, because it physically can’t; it only ever emits values it already knows are real.
The model stays on the table for later. But it has to actually beat that plain baseline first, on a request it’s never seen, in both languages we test in — not just run and look impressive in a demo.
Three bugs a database found that no test did
Even after all of that, three real bugs made it into a staging environment before anyone caught them, and all three are worth admitting to.
The preference weight from the section above had its own, found earliest: a track played four times somehow scored below one skipped four times. The cause was folding “how recently this was played” into the number stored for a track’s weight, instead of keeping it as a separate, query-time multiplier — recent plays and recent skips distorted the stored score in opposite directions, and unevenly, so the ranking quietly inverted. The unit tests didn’t catch it because they’d been written against the same wrong design. Only seeding real listening sessions and looking at the actual output caught it.
The energy scale, being a percentile, has to be rebuilt whenever new tracks get measured — and for a while, that only happened once a night. Turn the feature on, ask for a playlist an hour later, and every energy-filtered request matched nothing. The error was technically accurate and completely misleading: “nothing in the library fits this” is the right sentence for a genuinely narrow request, and the wrong explanation for “the index hasn’t caught up yet.” Fixed by rebuilding the moment a measurement pass finishes, not on a nightly schedule.
The third bug is smaller and more instructive. A day-count filter — “not played in the last 60 days” — got bound to the database as a floating-point number. Postgres 17, on our development machine, accepted that without complaint. Postgres 16, which is what we actually deploy to, does not, and returned a server error on four of the five built-in presets. It passed every unit test, since none of them touch a real database — and it passed locally for a less excusable reason: “locally” and “the server” were quietly two different major versions of the same database.
The fix was one line. Finding it took someone actually running the feature against something that behaves like production, not reading the diff and trusting it.
What we’d tell a team building the next one
Most of what makes a “smart” feature feel smart is the boring 80%: careful filters and a sequencing pass that respects real constraints. That’s true whether or not a model is anywhere near it. A closed, validated schema for anything a model — or any client — produces is what makes trusting that output cheap later; skip it early and you’ll be retrofitting it under pressure. Pick a scale that’s internally consistent for your own data rather than chasing a universal number that doesn’t exist. And test an LLM against the language and the edge cases it wasn’t tuned on before you let it be a default, not after.
The last one is the cheapest lesson available and the easiest to skip: if your development database and your production database are different major versions, that gap will eventually be the bug, and it will look exactly like a green build lying to you.
Smart playlists are still working their way from our own test library out to everyone else. Register on the front page if you’d like to be one of the first to get a playlist from it — and tell us when it gets one wrong.