Hello dear Python gods :)
I'm new at Python and was going over the documentation .
provides info for compressing text but IMHO some important additional info is needed, without which I expect nearly all users would get stuck just as I have. Moreover, while looking for a solution I noticed a lot of articles with outdated examples which no longer work - this confused me even more to the point I started suspecting something was amiss on my end...
Anyhow, adding this info would do the trick :
zlib works only with bytes and in Python 3.x strings are no longer bytes.
Example :
>>> zlib.compress('this is a sample string which does not work in zlib because it requires bytes and not unicode')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: a bytes-like object is required, not 'str'
>>>
Luckily Python has a very simple solution for this : encode to utf-8, compress, decode back to unicode like so :
>>> text = 'this is a sample string which does not work in zlib because it requires bytes and not unicode with chars like this »שלום '
>>> encoded = text.encode()
>>> text = 'this is a sample with unicode chars »שלום '
>>> encoded = text.encode()
>>> compressed = zlib.compress(encoded)
>>> decompressed = zlib.decompress(compressed)
>>> decoded = decompressed.decode()
>>> len(text),len(encoded),len(compressed) #notice that compression of short strings can be larger than the original
(42, 47, 54)
>>> print(text,'\n',encoded,'\n',compressed)
this is a sample with unicode chars »שלום
b'this is a sample with unicode chars \xc2\xbb\xd7\xa9\xd7\x9c\xd7\x95\xd7\x9d '
b'x\x9c+\xc9\xc8,V\x00\xa2D\x85\xe2\xc4\xdc\x82\x9cT\x85\xf2\xcc\x92\x0c\x85\xd2\xbc\xcc\xe4\xfc\x94T\x85\xe4\x8c\xc4\xa2b\x85C\xbb\xaf\xaf\xbc>\xe7\xfa\xd4\xebs\x15\x00\xb2K\x14|'
>>> text == decoded
True
Thank you ! You are doing a great thing !
B.R.,
Guy Goldner