Hi all — I’ve built something on top of MusicBrainz and I’d like the people who know this data best to tell me what they think and if there’s any issues/mistakes. It’s not comprehensive but I’m added in manageable chunks to avoid big data bills!
It’s called Antecedent. The idea: most music-history tools connect artists who sound alike. This one connects them through the people who actually carried sounds between them — the session players, producers and engineers in the relationship data. You can pick two artists and it finds the shortest real path between them through the credits.
Rudy Van Gelder alone sits at over 1,100 connections in the graph. An enormous amount of what we think of as “the Blue Note sound” is one engineer in a room in New Jersey.
Where it comes from: MusicBrainz core (CC0) for artists, releases and — crucially — the relationship data, which is the part nobody else models as connective tissue. Wikidata and Discogs dumps for extra info. Everything is CC0 or attributed; I’ve deliberately kept out the NC-licensed tables. Full sources are on the attribution page. It’s a bit of fun but I really interested in making it useful.
What I’d genuinely like input on:
- Spot-check some credits. Where has my mapping from MB relationship types to my own edge types gone wrong?
- Any ideas for improvements.
- Any placeholder or special-purpose entities I should be excluding that I haven’t?
- Double ups or glaring holes in the data (there are plenty..).
Thanks for the database. It’s remarkable, and this wouldn’t exist without it.
Joel.