A listening galaxy · formed May 21, 2022 · age 4.4 years
Every song I've ever played is a star.
Four years of my Spotify history, condensed into a galaxy: every song a star, every artist a constellation, every play a little more light. Built on the data warehouse I engineered from my raw listening export.
91,160listens
4,647stars
1,230constellations
4,395hours of light
LR
Four years of my listening, as a universe.
Headphones recommended.
Replay
May 2022
0 stars lit · drag to spin · shift-drag to move · scroll to zoom
May 20222023202420252026
in my head(phones)
Now playing
in my head(phones) · end of Side A
Side A is over. Lift the needle.
Formation complete · 4.4 years · one galaxy
Now, slowly, track by track.
You just watched 4,647 songs arrive. The album stops at the 19 moments that changed me: the night I played one song 119 times, the October that broke me, the morning I left home. Every number comes straight from my data.
Field guide
How to read the sky
A starOne song, born the first time I listened for at least 30 seconds, Spotify's own threshold for a stream.
Size is luminosityHow many times I've listened, with the radius growing by the square root so a star's area tracks its listens.
Color: era of formationThe year a star formed: periwinkle 2022, lilac 2023, mauve 2024, dusty rose 2025, champagne 2026.
Color: emotional spectrumHappy pink love, yellow party, orange confident, sour lilac bittersweet, burgundy heartbreak and deep purple dark, tagged by me or inferred from what I play a song with.
CoolingA star I've stopped playing cools and dims in the year after its last listen.
ConstellationOne artist, whose lines are the shortest path through their songs (a minimum spanning tree).
OrbitDistance from the core is when an artist entered my life, so the galaxy expands outward over four years, and inside each constellation the most-played songs orbit closest to the middle.
Gold arcs: gravitySongs that pull each other along: ones I play back-to-back in the same listening session.
The signalListens per month, pink when my #1 artist was the brightest source and plum when someone else was, with dots for story chapters.
The profile
What four years of listening says about me
Spotify never asked my age, my gender, where I live, when I sleep or where my family is from. It didn't need to. Every guess below uses only timestamps and song choices, the same signals any platform holds, with the signal it reads and the evidence from my data.
The point
A playlist is a diary nobody meant to keep
None of this came from anything I typed. That's why the pipeline drops IP addresses, countries and devices in its first step, leaves out private sessions, and why this site publishes summaries per song and never the time of a single play.
Case study
How I built Heavy Rotation
Spotify Wrapped sums up a year in a few cards, so I built a data warehouse from my raw listening history and drew all four years as a sky you can move through, to see how artists entered my life, took over, faded and came back.
The data
Four years, one export
182,293raw records
91,160real listens
5,053sessions
4,647songs
1,230artists
1,887song links
Spotify's extended streaming history is one record per play (timestamp, track, milliseconds listened, why it started and ended), and about half of them last under 30 seconds, so here's how the raw file narrows down to the stars you see:
What the data revealed
More than a playlist
A listening history is a diary nobody meant to keep. Timestamps, skips and sessions were enough to find my moods, my sleep, my routines, the big days of my life, my trips home and my cultural identity, none of which Spotify ever asked me about.
Every one of these is computed by the pipeline and checked against what I remember; the tour, in my head(phones), tells them as an album: Side A in Dallas, Side B in Austin, and 5 from the vault.
The pipeline
From a zip file to a sky
Extract · PythonEvery Streaming_History_Audio_*.json is read straight out of Spotify's zip, and the raw files never leave my laptop.
↓
Transform · PythonSongs only, no private sessions, no IP, country or device, Austin time, de-duplicated, overnight loops removed and the many IDs Spotify gives one song merged into one.
↓
Model · PostgreSQL star schemaOne fact table of plays and five dimensions (track, artist, album, date, session), checked for integrity before anything loads.
↓
Analyze · SQLWindow functions and CTEs for eras, streaks and obsessions, served by FastAPI at listening-history.
↓
Export · Pythongalaxy_export.py turns the same tables into one small JSON of per-song and per-artist summaries, links, moods and story chapters, never single plays.
↓
Render · Canvas + JavaScriptNo framework and no build step, just one canvas projecting every star, line and particle into 3D at 60 frames per second.
The hard parts
What took the most thinking
One song, many IDs. Spotify gives the single, the album cut and the deluxe edition separate IDs, so I merge them on a normalized title plus artist and every song keeps one history.
Bad data is still data. A song left on repeat overnight logged 145 plays while I slept, so the pipeline drops each loop by date and title with the reason written right next to it.
Which songs belong together? I count how often two songs play back-to-back inside one session and score each pair by cosine similarity, together ÷ √(listens A × listens B), so my #1 artist doesn't link to everything and half the links cross artists.
Moods for songs I never tagged. I hand-tagged 389 songs, and every other song borrows the mood that wins 60% of its back-to-back links (label propagation), which adds 617 more and colors 82% of my listening without guessing the rest.
What two songs share. Pick any two stars and a breadth-first search finds the shortest chain of back-to-back plays between them, next to everything else they have in common.
Thousands of stars in 3D. Glows come from cached sprites, hit-testing reuses the projected positions and constellation lines are computed once as minimum spanning trees, so 4,647 stars spin at 60 fps.
Time as state. Every star knows when it was born and when it went quiet, so the timeline, the replay and the tour all just move one number.
Running on your device right now: – frames per second, 4,647 stars, data file – KB.
Data and privacy
What I left out on purpose
IP addresses, countries and devices are dropped in the first transform, private sessions are excluded, the raw export never touches GitHub, and the published file holds summaries per song and artist instead of a log of when I listened.
Built by Suhani Tiwari
Data engineering · Data modeling · Visualization · Front-end development. I'm an MIS student at McCombs, UT Austin, and I made Heavy Rotation to show that personal behavioral data can be explored as a story, not just read off a dashboard.