Python Module 2 – Values, variables, numbers and strings
Strings: literals, indexing and slicing
Write Python strings with quotes, escapes, triple quotes and raw strings, pick out characters and slices by position, and see why a string never changes.
What you will learn
- Create strings with quotes, triple quotes and raw strings
- Index and slice strings with negative indices and steps
- Explain string immutability and its consequences
Before you start
On this page
Every name, message, file line and web page a program handles is text, and in Python text is a string, a value
of type str. A string is a sequence: it holds characters in order, each at a numbered position, so you can pick out
one character or a run of them. And once a string exists it never changes. Those two facts explain almost everything
in this lesson.
Writing strings in your code
A string written directly in a program is a string literal. Put the text between single quotes or double quotes; both make exactly the same string, so choose whichever means less escaping:
# Single and double quotes make the same kind of string: the quotes are not part of the text.
print('chai' == "chai")
# Pick the quote that is not inside the text, or put a backslash before it.
print("It's ready")
print('She said "two cups"')
print('It\'s ready' == "It's ready")
# A backslash starts an escape sequence: \n is a newline, \t a tab, \\ one backslash.
menu = "tea\t20\ncoffee\t30"
print(menu)
print(repr(menu))
print(len("\n"), len("\\"))
# \u followed by four hex digits is a character by its Unicode number.
print("\u0905", "\u03c0")
# Triple quotes keep line breaks exactly as typed.
note = """Shop opens at 9.
Closed on Sundays."""
print(note)
# Literals written next to each other are joined into one string.
title = ("Namaste, "
"world")
print(title) Output
True It's ready She said "two cups" True tea 20 coffee 30 'tea\t20\ncoffee\t30' 1 1 अ π Shop opens at 9. Closed on Sundays. Namaste, world
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 literals.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
What each part shows:
'chai' == "chai"isTrue: the quotes only mark where the text starts and ends.- Text that contains one kind of quote is easiest to write with the other kind around it. Otherwise put a backslash
in front of the quote, as in
'It\'s ready'. - A backslash starts an escape sequence, a way to write a character that you cannot type directly or that would
end the string:
\nis a line break,\ta tab and\\a single backslash. Each escape is one character, which is whylen("\n")is 1. print()shows the text itself, with real line breaks and tabs.repr()shows the string the way you would write it in code, quotes and escapes included, which makes invisible characters visible. The interactive shell usesrepr()when it echoes a value.\ufollowed by four hexadecimal digits is a character by its Unicode number:\u0905is the Devanagari letter अ and\u03c0is π.\N{GREEK SMALL LETTER PI}names the same character in words.- Triple quotes (
"""or''') let a string run over several lines and keep the line breaks as typed. You will meet them again as docstrings, the descriptions at the top of functions. - Two literals next to each other, with nothing but spaces between them (or line breaks, inside brackets), become
one string. That is how the last part of the program splits a long text over two lines without a
+.
Raw strings, and the Windows path mistake
Backslashes are where most string bugs start. A Windows path typed as an ordinary string silently changes:
# A Windows path typed as an ordinary string: \n and \t are escape sequences.
print("C:\new\table.csv")
# Two ways to keep every backslash: a raw string, or doubled backslashes.
print(r"C:\new\table.csv")
print("C:\\new\\table.csv")
# Forward slashes also work for paths on Windows, and need no escaping.
print("C:/new/table.csv") Output
C: ew able.csv C:\new\table.csv C:\new\table.csv C:/new/table.csv
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 windows_path.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
\n in C:\new became a line break and \t in \table a tab, so the first line prints a broken path and Python
reports no error. A raw string, written with r before the opening quote, treats every backslash as a plain
character. Doubling each backslash works too, and so do forward slashes, which Windows accepts in paths.
Raw strings are also the usual way to write regular expressions, which use backslashes a lot. A backslash followed by
a letter that is not an escape, such as \d (a digit, in a regular expression), is kept as it is, but Python warns
about it:
pattern = "\d+"
print(pattern, len(pattern))
fixed = r"\d+"
print(fixed, len(fixed), pattern == fixed) Output
\d+ 3 \d+ 3 True
Printed as an error (standard error)
regex_escape.py:1: SyntaxWarning: "\d" is an invalid escape sequence. Such sequences will not work in the future. Did you mean "\\d"? A raw string is also an option. pattern = "\d+"
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 regex_escape.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The program still worked: both strings are a backslash, a d and a +. The warning, printed on standard error
when the file is compiled, matters because the language reference says a future version will raise a SyntaxError
instead, so write r"\d+" (or "\\d+") now.
Version note
Unknown escapes have produced a SyntaxWarning since Python 3.12; Python 3.11 and older stay silent by default.
The wording above, with the suggestions "\\d" and a raw string, is new in Python 3.14.
A raw string has one limit: it cannot end with a single backslash, because that backslash still keeps the closing quote from ending the string:
folder = r"C:\Users\Asha\"
print(folder) Output (exit status 1)
Printed as an error (standard error)
File "raw_ending.py", line 1
folder = r"C:\Users\Asha\"
^
SyntaxError: unterminated string literal (detected at line 1); perhaps you escaped the end quote?
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 raw_ending.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Nothing runs, as with every syntax error, and the message (since Python 3.13) even guesses the cause. When a path
must end with a separator, write "C:\\Users\\Asha\\" or "C:/Users/Asha/", or leave the separator off and add it
later.
Picking out characters by position
Each character of a string has a position, its index, and word[i] gives the character at index i. Python
counts from 0, so the first character is word[0]. Negative indexes count from the end: word[-1] is the last
character and word[-2] the one before it.
Positions and slices of the string "PLANET"
Text description of the diagram
The diagram shows the string "PLANET" as six boxes in a row, one letter in each: P, L, A, N, E and T.
- Under each box is its position counted from the left, starting at 0: P is 0, L is 1, A is 2, N is 3, E is 4 and T is 5.
- Under that is its position counted from the right, starting at -1: T is -1, E is -2, N is -3, A is -4, L is -5 and P is -6.
- A bracket above the boxes covers L, A, N and E. It is labelled word[1:5], which gives 'LANE': the slice starts at position 1 and stops just before position 5, so T is not included.
- A bracket below the boxes covers N, E and T. It is labelled word[-3:], which gives 'NET': the slice starts three from the end and, with nothing after the colon, runs to the end.
There is no separate character type in Python: word[0] is itself a string, one character long. Asking for a
position the string does not have stops the program:
word = "PLANET"
print(word[5])
print(word[6]) Output (exit status 1)
T
Printed as an error (standard error)
Traceback (most recent call last):
File "past_the_end.py", line 3, in <module>
print(word[6])
~~~~^^^
IndexError: string index out of range
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 past_the_end.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
"PLANET" has six characters, so its indexes run from 0 to 5 (or -6 to -1). Index 6 would be the seventh character.
The last valid index is always len(word) - 1, and a loop or calculation that ends one step too far is one of the
most common causes of an IndexError.
Slices: a run of characters
A slice takes several characters at once. word[start:stop] begins at index start and stops before index
stop, so word[1:5] takes indexes 1, 2, 3 and 4. A third number, the step, says how far to move each time.
This program prints each expression next to its value (the = inside the braces of an f-string does that; the
f-strings lesson explains it):
word = "PLANET"
# One character: positions count from 0 at the left, or from -1 at the right.
print(f"{word[0]=} {word[5]=} {word[-1]=} {word[-6]=}")
# A slice [start:stop] starts at start and stops just before stop.
print(f"{word[0:4]=}")
print(f"{word[1:5]=}")
print(f"{word[3:]=}")
print(f"{word[-3:]=}")
print(f"{word[:2]=}")
# A third number is the step: every second character, or backwards.
print(f"{word[::2]=}")
print(f"{word[::-1]=}")
# Out-of-range positions never fail in a slice: they are cut back, or the slice is empty.
print(f"{word[4:100]=}")
print(f"{word[10:]=}")
print(f"{word[4:2]=}") Output
word[0]='P' word[5]='T' word[-1]='T' word[-6]='P' word[0:4]='PLAN' word[1:5]='LANE' word[3:]='NET' word[-3:]='NET' word[:2]='PL' word[::2]='PAE' word[::-1]='TENALP' word[4:100]='ET' word[10:]='' word[4:2]=''
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 slices.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Some rules to take from it:
- Leave out
startto begin at the start, and leave outstopto run to the end:word[:2]is the first two characters andword[3:]everything from index 3 on. - Negative numbers work in slices as they do in indexes.
word[-3:], the last three characters, is the usual way to get the end of a string whatever its length. - Because the stop is excluded,
word[:i] + word[i:]is always the whole string again, andword[i:j]hasj - icharacters when0 <= i <= j <= len(word). - A step of
-1walks backwards, soword[::-1]is the string reversed. - Slices are forgiving where indexes are not. A stop past the end is cut back to the end, a start past the end gives
the empty string
'', and a start after the stop gives''too. None of them raises an error (only a step of 0 does), so a slice that silently comes back empty usually means one of its numbers is wrong.
Strings never change
Strings are immutable: once a string exists, nothing can change its characters. You can build a new string from pieces of an old one, but you cannot assign to one of its positions:
word = "PLANET"
# Build a new string from pieces of the old one.
changed = "J" + word[1:]
print(changed, word)
# Changing a character in place is not allowed.
word[0] = "J" Output (exit status 1)
JLANET PLANET
Printed as an error (standard error)
Traceback (most recent call last):
File "immutable.py", line 8, in <module>
word[0] = "J"
~~~~^^^
TypeError: 'str' object does not support item assignment
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 immutable.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
The new string "JLANET" was printed next to the old one, which is still "PLANET". Then word[0] = "J" failed.
This has some everyday consequences:
- To “change” a string, build a new one and bind a name to it:
word = "J" + word[1:]makeswordrefer to the new string, while the old one is untouched (and is thrown away once nothing refers to it). - Every string method, such as
upper()orreplace(), returns a new string and leaves the original alone, so you must keep the result:name.upper()on a line of its own does nothing useful. The next lesson covers the methods. - When a program builds one long text from many small pieces, adding them one at a time with
+=creates a new string each time. The Python documentation warns that the total time can then grow with the square of the text’s length, and suggests putting the pieces in a list and calling"".join(pieces)once at the end. - Because a string cannot change, it is safe to share: two names bound to the same string can never surprise each other. That is also why strings can be dictionary keys.
What len() counts
len() returns the number of items in a string, and each item is a Unicode code point: a number that the
Unicode standard gives to a letter, a digit, a sign or an emoji. Most of the time one code point is one letter you
see, but not always:
import unicodedata
greeting = "नमस्ते" # "namaste" in Hindi
print(len(greeting))
# Each item of a string is one Unicode code point: a number with a name.
for ch in greeting:
print(f"U+{ord(ch):04X} {unicodedata.name(ch)}")
# Slices count code points too, so they can split what looks like one letter.
print(greeting[:3], greeting[:4])
# Two ways to write é: one code point, or e followed by a combining accent.
one = "\u00e9"
two = "e\u0301"
print(one, two, len(one), len(two), one == two) Output
6 U+0928 DEVANAGARI LETTER NA U+092E DEVANAGARI LETTER MA U+0938 DEVANAGARI LETTER SA U+094D DEVANAGARI SIGN VIRAMA U+0924 DEVANAGARI LETTER TA U+0947 DEVANAGARI VOWEL SIGN E नमस नमस् é é 1 2 False
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 code_points.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
"नमस्ते" looks like three or four letters, yet it is six code points. In Devanagari the vowel sign े and the
virama ्, which removes the vowel from स so that it joins the next letter, are code points of their own that
combine with the letters around them. Slices count code points as well, so greeting[:4] ends with the virama and
shows half of a joined letter. The last line shows the same effect in a Latin script: an é can be one code point or
two (an e followed by a combining accent). Both look the same, but Python compares code points, so they are not
equal.
Note
Treat len() as “how many code points”, never as “how wide it looks” or “how many bytes it takes in a file”. Bytes
depend on the encoding, which a later lesson on text and bytes covers. To compare text that may use either form
of é, unicodedata.normalize() first turns both into the same code points.
Key takeaways
- Single and double quotes make the same string; escapes such as
\n,\t,\\and\u0905write characters you cannot type, and triple quotes keep line breaks. - A raw string (
r"…") keeps every backslash, which suits Windows paths and regular expressions, but it cannot end with a single backslash. word[i]counts from 0 at the left and from -1 at the right; a position past the end raisesIndexError.word[start:stop:step]includes the start and excludes the stop, and positions out of range never raise an error;word[::-1]reverses.- Strings are immutable: build new strings instead of changing them, keep the results of string methods, and join
many pieces with
"".join(). len()counts Unicode code points, which are not always the letters you see.
Exercise
Exercise · Easy · Python
Mask an account number for a receipt
Receipts and banking screens show only the end of an account number, such as XXXXXXXX9012. Write mask_account(number), which takes the number as a string and returns the masked version.
- Replace every digit except the last four with
X, so the result has the same length asnumber:mask_account("123456789012")returns"XXXXXXXX9012", andmask_account("12345")returns"X2345". - A number of four digits or fewer would be shown in full that way, so mask all of it:
mask_account("1234")returns"XXXX"andmask_account("7")returns"X". - Raise
ValueErrorwhennumberis empty or holds anything but the digits 0 to 9, such as a space, a hyphen, a letter or a digit from another script.
The sample tests import your function from mask.py.
Starter code · mask.py
def mask_account(number):
"""Return number with every digit but the last four replaced by X."""
# Replace this line with your code.
return number The sample tests · test_mask.py
from mask import mask_account
def raises_value_error(number):
"""True when mask_account(number) raises ValueError."""
try:
mask_account(number)
except ValueError:
return True
return False
def test_long_numbers():
"""keeps the last four digits of a long number"""
assert mask_account("123456789012") == "XXXXXXXX9012"
assert mask_account("9876543210") == "XXXXXX3210"
def test_five_digits():
"""masks one digit of a five-digit number"""
assert mask_account("12345") == "X2345"
def test_short_numbers():
"""masks every digit of a number with four digits or fewer"""
assert mask_account("1234") == "XXXX"
assert mask_account("987") == "XXX"
assert mask_account("7") == "X"
def test_same_length():
"""returns a string as long as the number"""
for number in ["0000111122223333", "55555", "42"]:
assert len(mask_account(number)) == len(number)
def test_rejects_other_characters():
"""raises ValueError for spaces, hyphens, letters and other scripts' digits"""
assert raises_value_error("1234 5678")
assert raises_value_error("12-34-56")
assert raises_value_error("12a456")
assert raises_value_error("١٢٣٤٥٦")
def test_rejects_empty():
"""raises ValueError for an empty string"""
assert raises_value_error("") A hint
number[-4:] is the last four characters, and "X" * 8 is eight X's. Try both on a string of three characters before you trust them: a slice past the start gives the whole string, and a negative count gives "". For the check, number.isdigit() is True when every character is a digit of any script ("١٢".isdigit() is True too), and number.isascii() is True only for plain ASCII text, so together they accept exactly 0 to 9.
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- The Python Tutorial: An Informal Introduction to Python (Text) (Python Software Foundation)
- Text Sequence Type, str (The Python Standard Library) (Python Software Foundation)
- Common Sequence Operations (The Python Standard Library) (Python Software Foundation)
- String and Bytes literals (The Python Language Reference) (Python Software Foundation)
- unicodedata, Unicode Database (Python Software Foundation)
- File path formats on Windows systems: canonicalize separators (Microsoft)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress