Strings, Unicode, Bytes & Text Encodings

In Python 3, text (str) and binary data (bytes) are strictly decoupled. A str is an abstract sequence of Unicode code points, while a bytes object is an immutable sequence of raw 8-bit integers ($0..255$). Understanding PEP 393 Flexible String Representation, UTF-8 encoding mechanics, and buffer protocols is essential for systems engineering.

This chapter details CPython’s string memory layout, compact Unicode objects, bytes vs bytearray memory buffers, and text encoding boundaries.


1. PEP 393: Flexible String Representation

Prior to Python 3.3, CPython allocated all strings using either 2-byte (UCS-2) or 4-byte (UCS-4) arrays, wasting vast amounts of RAM for simple ASCII text. PEP 393 introduced a dynamic, compact internal representation:

CPython inspects the maximum code point character in a string at creation time and selects one of three compact C array representations:

  1. 1-Byte Representation (Latin-1 / ASCII): Used if max code point $\le U+00FF$ (1 byte per char).
  2. 2-Byte Representation (UCS-2): Used if max code point $\le U+FFFF$ (2 bytes per char).
  3. 4-Byte Representation (UCS-4): Used if max code point $> U+FFFF$ (e.g. Emojis, 4 bytes per char).
CPython PEP 393 String Memory Layout:

ASCII String "hello" (Max char <= U+00FF):
[ PyASCIIObject Header (48B) ] -> [ 'h' | 'e' | 'l' | 'l' | 'o' | '\0' ] (1 byte/char)

Emoji String "hello Python 🐍" (Max char U+1F40D > U+FFFF):
[ PyCompactUnicodeObject Header (72B) ] -> [ 4 bytes per character array ]

Because of PEP 393, pure ASCII strings consume 1 byte per character payload, while strings containing a single emoji expand the internal C-array to 4 bytes per character for all characters in that string.


2. Text (str) vs. Binary (bytes & bytearray)

  • str: Immutable sequence of Unicode code points. Cannot be sent directly over network sockets or written directly to disk without encoding.
  • bytes: Immutable sequence of raw bytes ($0..255$). Represents encoded text, image data, or network payloads.
  • bytearray: Mutable version of bytes. Allows in-place byte modifications without re-allocating new objects on the heap.
# Encoding: str -> bytes (Unicode Code Points -> UTF-8 Bytes)
text = "Python 🐍"
encoded_bytes = text.encode("utf-8")
print(encoded_bytes)  # b'Python \xf0\x9f\x90\x8d' (Emoji takes 4 UTF-8 bytes)

# Decoding: bytes -> str (UTF-8 Bytes -> Unicode Code Points)
decoded_text = encoded_bytes.decode("utf-8")

3. Unicode Normalization Forms (NFC vs. NFD)

Visual equality does not guarantee binary equality in Unicode. A character like é can be represented as a single precomposed code point (U+00E9) or a base letter e plus a combining accent (U+0065 + U+0301).

import unicodedata

str1 = "café"       # Precomposed (NFC)
str2 = "cafe\u0301"  # Decomposed (NFD)

print(str1 == str2)  # False! (Binary code points differ)

# Solution: Normalize strings before storing or comparing
norm1 = unicodedata.normalize("NFC", str1)
norm2 = unicodedata.normalize("NFC", str2)
print(norm1 == norm2)  # True!

4. Production Trade-offs & In-Place Buffers

  • String Concatenation in Loops: Strings are immutable. Executing s += char inside a loop creates an $O(N^2)$ allocation catastrophe. Use ''.join(list_of_strings) or a bytearray buffer.
  • bytearray for Sockets: When reading binary streams from sockets, pre-allocate a bytearray and read directly into its buffer to eliminate per-chunk heap allocations.
Display Options
Appearance
Text Size
100%