That album, you know, "½Ð¸ò§Ú¨Ó" by "¤pè". It's a very popular CD!

Hi everyone! Remember me? I haven’t logged into here for four years! Anyway, I have a puzzle for anyone who can speak, or at least read and write Mandarin or Cantonese or one of the many Chinese languages I don’t even have the slightest knowledge of.

I impulsively looked at Top CD stubs - MusicBrainz and the third most commonly looked up CD Stub is “½Ð¸ò§Ú¨Ó” by “¤p­è” with 1871 lookups. This one, if you want to check it out:

So… I used to decipher codes for a living, and I took one look at that CD Stub and thought: “I bet I can decode this.” …and I did! The original metadata (when originally entered into CDDB or whatever) was encoded in CP-950 or BIG5-HKSCS or similar, and then at some time between 2005 to 2011, it got imported (transitively) into MusicBrainz as if it was CP-1252 or ISO8859-1 or similar… which was then converted, as if it was LATIN1 and not BIG5, into UTF-8, which is how you see the data right now in that MusicBrainz CD Stub.

This is what you get decoding and reëncoding it correctly:
請跟我來
CD stub by 小剛

  1. 翠湖寒
  2. 忘了我是誰
  3. 千言萬語
  4. 我祗在乎你
  5. 愛的代價
  6. 風聲鶴唳
  7. 被遺忘的時光
  8. 情人的眼略
  9. 筑d多珍重
  10. 請跟我來
  11. 你的眼神
  12. 兌現
  13. 該如何面對你

I can’t read a single character of Chinese, I can’t speak a word of Cantonese and Mandarin, after a lot of Google Searches, I believe that this CD Stub is this album:

… but … as far as I can tell, Google Search and Google Translate treat all of these as synonyms:

《小刚》
《小剛》
《鍾志剛》
《钟志刚》

… but are these “the same”? I don’t know enough Chinese to understand exactly what is going on with the way these are written. Likewise, are the following titles “the same”?

《請跟我來》
《请跟我来》

The same writing difference happens with the name of track ten: 《請跟我來》vs. 《請跟我來》I don’t know if this is a meaningful difference.

I also discovered that there’s a 2005 and a 2007 release of this album with different covert art, so I wouldn’t attach this CD Stub anyway (even if I could read Chinese) until I knew which release this CD Stub is actually for. (MusicBrainz only has the 2005 release at the moment.)

I haven’t added the 2007 release… because I can’t read Chinese and I have no idea if I’m doing it correctly. I have so little clue, that I can’t even figure out if these two artists are in fact the same artist and should be merged:

I can’t figure out if these are the same, or if this is a popular book an ancient Chinese poetry:

(There’s a 1983 album by another, different, Chinese artist with the same title.)

Also, I have no idea if this third MusicBrainz artist is also the same singer:

Is this the same guy, too? Spotify – Web Player

Anyway, so, if anyone who understands… “mandopop” (is that really what this is called?) wants to do some editing, you can give the third most popular CD Stub on MusicBrainz a real home attached to an actual MusicBrainz release.

Postscript:

If you want to try decoding this 文字化け CD Stub yourself, pipe the characters through something like:

iconv -f utf8 -t latin1 | iconv -f big5 -t utf8

Appendix:

Discogs doesn’t have any information about 請跟我來 by 小剛. When you Google Search for those EAN/UPC codes for the two albums above, you get three results and MusicBrainz is one of them. There are other EAN/UPC codes listed on these other sites (below), which may or may not be the same album. I’m not sure because I don’t understand enough Chinese to know for certain.

I found a bunch of these possibly related clues while Google Searching the decoded strings from that CD Stub. I’ve inserted spaces here to stop “discourse .org” from turning these into hyperlinks. I hate text editors like this one that turn anything with a “.” between two words into a random hyperlink to who-knows-where, and no ability to turn this misfeature off. grrr.

“baike.baidu. com/item/请跟我来/8828469”

“www.eslite. com/product/1004245682563534”

“www.yesasia. com/us/請跟我來-vinyl-lp-中國版/1045607431-0-0-0-zh_TW/info.html”

“www.ccr. com.tw/goods/122950”

“www.ccr. com.tw/goods/316198”

“www.aiyiny. com/thread-3401-1-1.html”

10 Likes

Me too. Surround the URL with < > as <my.url> the < > disappear and leave you with a real URL without being messed around by the GUI. Example https://community.metabrainz.org/t/that-album-you-know-d-o-u-o-by-pe-its-a-very-popular-cd/815409 has < > around it to leave it unmolested by the GUI

Amazing detective work you have done there. And I totally recognise why :smiley: Just can’t help read the language

Edit: Also [text you want](url.you.want) also works. (Yeah, had to escape some characters to make that appear on screen… but square brackets round human text and normal brackets round the url it goes to.

1 Like

(Also not really an expert but) the character differences you are seeing look like the difference between simplified characters (used e.g. in modern Mandarin, in most of mainland China) and traditional characters (used e.g. in Cantonese, in Hong Kong and Taiwan), but with the same (approximately?) underlying meaning. It looks like you had to guess at the original encoding – both CP-950 and BIG5-HKSCS are for traditional characters, but possibly if it was in an encoding for simplified characters that would explain the discrepancy.

I forgot to mention, I know some popular artists with non-Latin-alphabet names have “search hint” aliases for the common 文字化け encoding error versions on their names. I attempted to add “½Ð¸ò§Ú¨Ó” and “¤p­è” aliases to that album and that artist respectively… and then I tried to search for “½Ð¸ò§Ú¨Ó” and “¤p­è”, and it didn’t work. By “didn’t work” I mean it was zero results with a direct search, and 300,000 random results using an indirect search with and without search operators.

So, I tried turning the aliases from a “Search Hint” into a regular alias, but without setting a type or locale. Once again… these strings are not found when searched for.

So… I should probably just make a new topic about this topic, but…

  1. Should popular misencodings of popular names be stored as aliases?
  2. Should the Search Engine itself just do this transformation itself automatically, without a human editor needing to manually add each individual aliases to individual database objects?
  3. If a human is adding each 文字化け alias manually, individually, by hand… Should there be a new Type and Locale specifically for the 文字化け alias. (I think there really should be.)
  4. Why are there no results right now when you search for “½Ð¸ò§Ú¨Ó” in releases, or “¤p­è” in artists?
1 Like

Yeah, I was guessing at the encoding. I mean, I guessed the first transformation correctly, but I wrote a script to just brute-force attempt to encode/decode the second half of the transformation. (I was manually trying JIS and GB because I guessed it was probably one of those CJK encodings, I just didn’t guess BIG5 right off the bat, so I just tried every encoding iconv knows and looked through the results for anything that seemed to decode/encode without errors.) I was a little bit worried that I would need to resort to doing a whole N by N matrix of every possible iconv encoding vs. every other one, but fortunately it didn’t come to that.

Once upon a time, I could just look at a bunch of octets in a hex dump, and could tell you immediately what the encoding (probably) is, but I’ve forgotten most of the ancient CJK ones now. It’s been decades since I last saw one.

Those with more strokes are traditional Han, used outside mainland China. The ones with less strokes are simplified Han, used in mainland China.

More than synonyms, they are the same words, but with script variation.

So, the printed CD tracklist you have depends of the country of release.

These look obviously different, one a 13 track album, the other a 3 track single with the last two tracks from that album (according to Spotify – Web Player though the dates are weird, album from 2007, single from 2019). The track name differences seem the already mentioned traditional vs simplified Chinese.

There’s a handy tool for guessing at these: Universal online Cyrillic decoder - recover your texts - it says it’s for Cyrillic, but it can do Chinese too. It works by pasting in the text, clicking on the “Select one” dropdown to show the decodings it tried, and scrolling through the options until you come to one that looks right.

1 Like

I tested this out with:

And after a few false guesses, the (manually set by me) encoding is GB2312 which produces:

金刚般若波罗蜜经
~ CD stub by 国语佛经传统课诵
# Title Length
1 金刚经 1:14:13

… which is in fact the real name of a CD according to Google Search. This is the time consuming part really… copy & pasting the output into Google Search if you don’t know how to read 漢字

Here are some more CD Stubs to try this with:

Hey, I think there’s something terribly broken with whatever search engine you’re using on musicbrainz.org It will say something like Found 2345 results for "whatever", and that’s like 23 pages if you view 100 results per page… and then… start reading through the first half dozen pages in sequential order… and, at least for me just now, after about 700 results I started to notice that the results are repeating themselves… but in a randomized order. And if you return to the “previous page”, after going to the “next page”, so you’re on page 8, and then go to page 9, and then go to page 8, you get a completely different set of results on page 8 the second time you load it after returning from page 9 (as opposed to advancing from page 7 the first time).

It’s not like the offset has shifted, like there’s just a problem with the cursor offset. The listed search results are different!

This happened with both “Indexed Search” and “Direct Search” which behaved surprisingly similar to the indexed search in giving me 9000 random results that did not contain the string I searched for.

1 Like

Indeed, it’s an awful known bug: https://tickets.metabrainz.org/browse/MBS-12154