python - Facebook JSON badly encoded

Question

Welcome To Ask or Share your Answers For Others

python - Facebook JSON badly encoded

posted Oct 17, 2021 in Technique[技术] by 深蓝 (71.8m points)

python - Facebook JSON badly encoded

I downloaded my Facebook messenger data (in your Facebook account, go to settings, then to Your Facebook information, then Download your information, then create a file with at least the Messages box checked) to do some cool statistics

However there is a small problem with encoding. I'm not sure, but it looks like Facebook used bad encoding for this data. When I open it with text editor I see something like this: Radosu00c5u0082aw. When I try to open it with python (UTF-8) I get Rados?x82aw. However I should get: Rados?aw.

My python script:

text = open(os.path.join(subdir, file), encoding='utf-8')
conversations.append(json.load(text))

I tried a few most common encodings. Example data is:

{
  "sender_name": "Radosu00c5u0082aw",
  "timestamp": 1524558089,
  "content": "No to trzeba ostatnie treningi zrobiu00c4u0087 xD",
  "type": "Generic"
}

Question&Answers:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Reply

深蓝 · Answer 1 · 2021-10-16T22:23:08+0000

I can indeed confirm that the Facebook download data is incorrectly encoded; a Mojibake. The original data is UTF-8 encoded but was decoded as Latin -1 instead. I’ll make sure to file a bug report.

In the meantime, you can repair the damage in two ways:

Decode the data as JSON, then re-encode any strings as Latin-1, decode again as UTF-8:

>>> import json
>>> data = r'"Radosu00c5u0082aw"'
>>> json.loads(data).encode('latin1').decode('utf8')
'Rados?aw'

Load the data as binary, replace all u00hh sequences with the byte the last two hex digits represent, decode as UTF-8 and then decode as JSON:

import re
from functools import partial

fix_mojibake_escapes = partial(
     re.compile(rb'\u00([da-f]{2})').sub,
     lambda m: bytes.fromhex(m.group(1).decode()))

with open(os.path.join(subdir, file), 'rb') as binary_data:
    repaired = fix_mojibake_escapes(binary_data.read())
data = json.loads(repaired.decode('utf8'))

From your sample data this produces:

{'content': 'No to trzeba ostatnie treningi zrobi? xD',
 'sender_name': 'Rados?aw',
 'timestamp': 1524558089,
 'type': 'Generic'}

Categories

python - Facebook JSON badly encoded

python - Facebook JSON badly encoded

Please log in or register to add a comment.

Please log in or register to reply this article.

1 Reply

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags