Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Python Module 2 – Values, variables, numbers and strings

Strings: literals, indexing and slicing

Write Python strings with quotes, escapes, triple quotes and raw strings, pick out characters and slices by position, and see why a string never changes.

  • Beginner
  • 25 minutes
  • Examples run with Python 3.14.8 and Pyodide 314.0.7
  • By MySmartCoPilot

What you will learn

  • Create strings with quotes, triple quotes and raw strings
  • Index and slice strings with negative indices and steps
  • Explain string immutability and its consequences

Before you start

On this page

Every name, message, file line and web page a program handles is text, and in Python text is a string, a value of type str. A string is a sequence: it holds characters in order, each at a numbered position, so you can pick out one character or a run of them. And once a string exists it never changes. Those two facts explain almost everything in this lesson.

Writing strings in your code

A string written directly in a program is a string literal. Put the text between single quotes or double quotes; both make exactly the same string, so choose whichever means less escaping:

Quotes, escapes and triple quotes Python · literals.py
# Single and double quotes make the same kind of string: the quotes are not part of the text.
print('chai' == "chai")

# Pick the quote that is not inside the text, or put a backslash before it.
print("It's ready")
print('She said "two cups"')
print('It\'s ready' == "It's ready")

# A backslash starts an escape sequence: \n is a newline, \t a tab, \\ one backslash.
menu = "tea\t20\ncoffee\t30"
print(menu)
print(repr(menu))
print(len("\n"), len("\\"))

# \u followed by four hex digits is a character by its Unicode number.
print("\u0905", "\u03c0")

# Triple quotes keep line breaks exactly as typed.
note = """Shop opens at 9.
Closed on Sundays."""
print(note)

# Literals written next to each other are joined into one string.
title = ("Namaste, "
         "world")
print(title)

Output

True
It's ready
She said "two cups"
True
tea	20
coffee	30
'tea\t20\ncoffee\t30'
1 1
अ π
Shop opens at 9.
Closed on Sundays.
Namaste, world

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 literals.py

What each part shows:

  • 'chai' == "chai" is True: the quotes only mark where the text starts and ends.
  • Text that contains one kind of quote is easiest to write with the other kind around it. Otherwise put a backslash in front of the quote, as in 'It\'s ready'.
  • A backslash starts an escape sequence, a way to write a character that you cannot type directly or that would end the string: \n is a line break, \t a tab and \\ a single backslash. Each escape is one character, which is why len("\n") is 1.
  • print() shows the text itself, with real line breaks and tabs. repr() shows the string the way you would write it in code, quotes and escapes included, which makes invisible characters visible. The interactive shell uses repr() when it echoes a value.
  • \u followed by four hexadecimal digits is a character by its Unicode number: \u0905 is the Devanagari letter अ and \u03c0 is π. \N{GREEK SMALL LETTER PI} names the same character in words.
  • Triple quotes (""" or ''') let a string run over several lines and keep the line breaks as typed. You will meet them again as docstrings, the descriptions at the top of functions.
  • Two literals next to each other, with nothing but spaces between them (or line breaks, inside brackets), become one string. That is how the last part of the program splits a long text over two lines without a +.

Raw strings, and the Windows path mistake

Backslashes are where most string bugs start. A Windows path typed as an ordinary string silently changes:

A path with backslashes Python · windows_path.py
# A Windows path typed as an ordinary string: \n and \t are escape sequences.
print("C:\new\table.csv")

# Two ways to keep every backslash: a raw string, or doubled backslashes.
print(r"C:\new\table.csv")
print("C:\\new\\table.csv")

# Forward slashes also work for paths on Windows, and need no escaping.
print("C:/new/table.csv")

Output

C:
ew	able.csv
C:\new\table.csv
C:\new\table.csv
C:/new/table.csv

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 windows_path.py

\n in C:\new became a line break and \t in \table a tab, so the first line prints a broken path and Python reports no error. A raw string, written with r before the opening quote, treats every backslash as a plain character. Doubling each backslash works too, and so do forward slashes, which Windows accepts in paths.

Raw strings are also the usual way to write regular expressions, which use backslashes a lot. A backslash followed by a letter that is not an escape, such as \d (a digit, in a regular expression), is kept as it is, but Python warns about it:

An escape Python does not know Python · regex_escape.py
pattern = "\d+"
print(pattern, len(pattern))

fixed = r"\d+"
print(fixed, len(fixed), pattern == fixed)

Output

\d+ 3
\d+ 3 True

Printed as an error (standard error)

regex_escape.py:1: SyntaxWarning: "\d" is an invalid escape sequence. Such sequences will not work in the future. Did you mean "\\d"? A raw string is also an option.
  pattern = "\d+"

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 regex_escape.py

The program still worked: both strings are a backslash, a d and a +. The warning, printed on standard error when the file is compiled, matters because the language reference says a future version will raise a SyntaxError instead, so write r"\d+" (or "\\d+") now.

Version note

Unknown escapes have produced a SyntaxWarning since Python 3.12; Python 3.11 and older stay silent by default. The wording above, with the suggestions "\\d" and a raw string, is new in Python 3.14.

A raw string has one limit: it cannot end with a single backslash, because that backslash still keeps the closing quote from ending the string:

A raw string that ends in a backslash Python · raw_ending.py
folder = r"C:\Users\Asha\"
print(folder)

Output (exit status 1)

Printed as an error (standard error)

  File "raw_ending.py", line 1
    folder = r"C:\Users\Asha\"
             ^
SyntaxError: unterminated string literal (detected at line 1); perhaps you escaped the end quote?

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 raw_ending.py

Nothing runs, as with every syntax error, and the message (since Python 3.13) even guesses the cause. When a path must end with a separator, write "C:\\Users\\Asha\\" or "C:/Users/Asha/", or leave the separator off and add it later.

String Escape / Unescape Paste text to see it as a Python string literal with every escape written out, or turn escapes back into text.

Picking out characters by position

Each character of a string has a position, its index, and word[i] gives the character at index i. Python counts from 0, so the first character is word[0]. Negative indexes count from the end: word[-1] is the last character and word[-2] the one before it.

The six letters of PLANET with their positions 0 to 5 from the left and -6 to -1 from the right, and the slices [1:5] and [-3:].word[1:5] → 'LANE'P0-6L1-5A2-4N3-3E4-2T5-1from the leftfrom the rightword[-3:] → 'NET'

Positions and slices of the string "PLANET"

Text description of the diagram

The diagram shows the string "PLANET" as six boxes in a row, one letter in each: P, L, A, N, E and T.

  • Under each box is its position counted from the left, starting at 0: P is 0, L is 1, A is 2, N is 3, E is 4 and T is 5.
  • Under that is its position counted from the right, starting at -1: T is -1, E is -2, N is -3, A is -4, L is -5 and P is -6.
  • A bracket above the boxes covers L, A, N and E. It is labelled word[1:5], which gives 'LANE': the slice starts at position 1 and stops just before position 5, so T is not included.
  • A bracket below the boxes covers N, E and T. It is labelled word[-3:], which gives 'NET': the slice starts three from the end and, with nothing after the colon, runs to the end.

There is no separate character type in Python: word[0] is itself a string, one character long. Asking for a position the string does not have stops the program:

One step past the end Python · past_the_end.py
word = "PLANET"
print(word[5])
print(word[6])

Output (exit status 1)

T

Printed as an error (standard error)

Traceback (most recent call last):
  File "past_the_end.py", line 3, in <module>
    print(word[6])
          ~~~~^^^
IndexError: string index out of range

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 past_the_end.py

"PLANET" has six characters, so its indexes run from 0 to 5 (or -6 to -1). Index 6 would be the seventh character. The last valid index is always len(word) - 1, and a loop or calculation that ends one step too far is one of the most common causes of an IndexError.

Slices: a run of characters

A slice takes several characters at once. word[start:stop] begins at index start and stops before index stop, so word[1:5] takes indexes 1, 2, 3 and 4. A third number, the step, says how far to move each time. This program prints each expression next to its value (the = inside the braces of an f-string does that; the f-strings lesson explains it):

Slices of one word Python · slices.py
word = "PLANET"

# One character: positions count from 0 at the left, or from -1 at the right.
print(f"{word[0]=}  {word[5]=}  {word[-1]=}  {word[-6]=}")

# A slice [start:stop] starts at start and stops just before stop.
print(f"{word[0:4]=}")
print(f"{word[1:5]=}")
print(f"{word[3:]=}")
print(f"{word[-3:]=}")
print(f"{word[:2]=}")

# A third number is the step: every second character, or backwards.
print(f"{word[::2]=}")
print(f"{word[::-1]=}")

# Out-of-range positions never fail in a slice: they are cut back, or the slice is empty.
print(f"{word[4:100]=}")
print(f"{word[10:]=}")
print(f"{word[4:2]=}")

Output

word[0]='P'  word[5]='T'  word[-1]='T'  word[-6]='P'
word[0:4]='PLAN'
word[1:5]='LANE'
word[3:]='NET'
word[-3:]='NET'
word[:2]='PL'
word[::2]='PAE'
word[::-1]='TENALP'
word[4:100]='ET'
word[10:]=''
word[4:2]=''

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 slices.py

Some rules to take from it:

  • Leave out start to begin at the start, and leave out stop to run to the end: word[:2] is the first two characters and word[3:] everything from index 3 on.
  • Negative numbers work in slices as they do in indexes. word[-3:], the last three characters, is the usual way to get the end of a string whatever its length.
  • Because the stop is excluded, word[:i] + word[i:] is always the whole string again, and word[i:j] has j - i characters when 0 <= i <= j <= len(word).
  • A step of -1 walks backwards, so word[::-1] is the string reversed.
  • Slices are forgiving where indexes are not. A stop past the end is cut back to the end, a start past the end gives the empty string '', and a start after the stop gives '' too. None of them raises an error (only a step of 0 does), so a slice that silently comes back empty usually means one of its numbers is wrong.

Strings never change

Strings are immutable: once a string exists, nothing can change its characters. You can build a new string from pieces of an old one, but you cannot assign to one of its positions:

Changing one character Python · immutable.py
word = "PLANET"

# Build a new string from pieces of the old one.
changed = "J" + word[1:]
print(changed, word)

# Changing a character in place is not allowed.
word[0] = "J"

Output (exit status 1)

JLANET PLANET

Printed as an error (standard error)

Traceback (most recent call last):
  File "immutable.py", line 8, in <module>
    word[0] = "J"
    ~~~~^^^
TypeError: 'str' object does not support item assignment

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 immutable.py

The new string "JLANET" was printed next to the old one, which is still "PLANET". Then word[0] = "J" failed. This has some everyday consequences:

  • To “change” a string, build a new one and bind a name to it: word = "J" + word[1:] makes word refer to the new string, while the old one is untouched (and is thrown away once nothing refers to it).
  • Every string method, such as upper() or replace(), returns a new string and leaves the original alone, so you must keep the result: name.upper() on a line of its own does nothing useful. The next lesson covers the methods.
  • When a program builds one long text from many small pieces, adding them one at a time with += creates a new string each time. The Python documentation warns that the total time can then grow with the square of the text’s length, and suggests putting the pieces in a list and calling "".join(pieces) once at the end.
  • Because a string cannot change, it is safe to share: two names bound to the same string can never surprise each other. That is also why strings can be dictionary keys.

What len() counts

len() returns the number of items in a string, and each item is a Unicode code point: a number that the Unicode standard gives to a letter, a digit, a sign or an emoji. Most of the time one code point is one letter you see, but not always:

Code points in a Hindi word Python · code_points.py
import unicodedata

greeting = "नमस्ते"  # "namaste" in Hindi
print(len(greeting))

# Each item of a string is one Unicode code point: a number with a name.
for ch in greeting:
    print(f"U+{ord(ch):04X} {unicodedata.name(ch)}")

# Slices count code points too, so they can split what looks like one letter.
print(greeting[:3], greeting[:4])

# Two ways to write é: one code point, or e followed by a combining accent.
one = "\u00e9"
two = "e\u0301"
print(one, two, len(one), len(two), one == two)

Output

6
U+0928 DEVANAGARI LETTER NA
U+092E DEVANAGARI LETTER MA
U+0938 DEVANAGARI LETTER SA
U+094D DEVANAGARI SIGN VIRAMA
U+0924 DEVANAGARI LETTER TA
U+0947 DEVANAGARI VOWEL SIGN E
नमस नमस्
é é 1 2 False

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 code_points.py

"नमस्ते" looks like three or four letters, yet it is six code points. In Devanagari the vowel sign े and the virama ्, which removes the vowel from स so that it joins the next letter, are code points of their own that combine with the letters around them. Slices count code points as well, so greeting[:4] ends with the virama and shows half of a joined letter. The last line shows the same effect in a Latin script: an é can be one code point or two (an e followed by a combining accent). Both look the same, but Python compares code points, so they are not equal.

Note

Treat len() as “how many code points”, never as “how wide it looks” or “how many bytes it takes in a file”. Bytes depend on the encoding, which a later lesson on text and bytes covers. To compare text that may use either form of é, unicodedata.normalize() first turns both into the same code points.

Unicode Character Inspector Paste नमस्ते or any other text to see each code point it is made of, with its number and name.

Key takeaways

  • Single and double quotes make the same string; escapes such as \n, \t, \\ and \u0905 write characters you cannot type, and triple quotes keep line breaks.
  • A raw string (r"…") keeps every backslash, which suits Windows paths and regular expressions, but it cannot end with a single backslash.
  • word[i] counts from 0 at the left and from -1 at the right; a position past the end raises IndexError.
  • word[start:stop:step] includes the start and excludes the stop, and positions out of range never raise an error; word[::-1] reverses.
  • Strings are immutable: build new strings instead of changing them, keep the results of string methods, and join many pieces with "".join().
  • len() counts Unicode code points, which are not always the letters you see.

Exercise

Exercise · Easy · Python

Mask an account number for a receipt

Receipts and banking screens show only the end of an account number, such as XXXXXXXX9012. Write mask_account(number), which takes the number as a string and returns the masked version.

  • Replace every digit except the last four with X, so the result has the same length as number: mask_account("123456789012") returns "XXXXXXXX9012", and mask_account("12345") returns "X2345".
  • A number of four digits or fewer would be shown in full that way, so mask all of it: mask_account("1234") returns "XXXX" and mask_account("7") returns "X".
  • Raise ValueError when number is empty or holds anything but the digits 0 to 9, such as a space, a hyphen, a letter or a digit from another script.

The sample tests import your function from mask.py.

Starter code · mask.py

def mask_account(number):
    """Return number with every digit but the last four replaced by X."""
    # Replace this line with your code.
    return number
The sample tests · test_mask.py
from mask import mask_account


def raises_value_error(number):
    """True when mask_account(number) raises ValueError."""
    try:
        mask_account(number)
    except ValueError:
        return True
    return False


def test_long_numbers():
    """keeps the last four digits of a long number"""
    assert mask_account("123456789012") == "XXXXXXXX9012"
    assert mask_account("9876543210") == "XXXXXX3210"


def test_five_digits():
    """masks one digit of a five-digit number"""
    assert mask_account("12345") == "X2345"


def test_short_numbers():
    """masks every digit of a number with four digits or fewer"""
    assert mask_account("1234") == "XXXX"
    assert mask_account("987") == "XXX"
    assert mask_account("7") == "X"


def test_same_length():
    """returns a string as long as the number"""
    for number in ["0000111122223333", "55555", "42"]:
        assert len(mask_account(number)) == len(number)


def test_rejects_other_characters():
    """raises ValueError for spaces, hyphens, letters and other scripts' digits"""
    assert raises_value_error("1234 5678")
    assert raises_value_error("12-34-56")
    assert raises_value_error("12a456")
    assert raises_value_error("١٢٣٤٥٦")


def test_rejects_empty():
    """raises ValueError for an empty string"""
    assert raises_value_error("")
A hint

number[-4:] is the last four characters, and "X" * 8 is eight X's. Try both on a string of three characters before you trust them: a slice past the start gives the whole string, and a negative count gives "". For the check, number.isdigit() is True when every character is a digit of any script ("١٢".isdigit() is True too), and number.isascii() is True only for plain ASCII text, so together they accept exactly 0 to 9.

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 5 immutable.py stops with a TypeError on its last line. What does it print before that?

    What does this program print? Choose one answer.

    word = "PLANET"
    
    # Build a new string from pieces of the old one.
    changed = "J" + word[1:]
    print(changed, word)
    
    # Changing a character in place is not allowed.
    word[0] = "J"
    Show the answer to question 1

    Answer: it prints

    JLANET PLANET

    "J" + word[1:] builds a new string and leaves word as it was, so the program prints both. The error happens only when the last line runs, because changing one character of a string in place is not allowed.

  2. Question 2 of 5 What does this print?

    Read the code, then choose one answer.

    word = "PLANET"
    print(word[2:4] + word[-1])
    Show the answer to question 2

    Answer: ANT

    word[2:4] starts at position 2 (A) and stops before position 4, so it is "AN". word[-1] is the last character, "T". Joined with + they make "ANT".

  3. Question 3 of 5 word is "PLANET". Which expressions give "NET"?

    Choose every answer that is right.

    Show the answer to question 3

    Answer:

    • word[3:]
    • word[-3:]
    • word[3:6]

    N is at position 3 (or -3), and leaving out the stop, or writing 6, runs to the end. word[3:5] and word[-3:-1] both stop before the T and give "NE".

  4. Question 4 of 5 Which of these lines is a syntax error?

    Choose one answer.

    Show the answer to question 4

    Answer: folder = r"C:\temp\"

    In a raw string the backslash before the closing quote still stops that quote from ending the string, so r"C:\temp\" never ends. Double the backslashes ("C:\\temp\\") or use forward slashes when a path must end with a separator.

  5. Question 5 of 5 len("नमस्ते") counts the code points of the Hindi word for "namaste". What does it return?

    Type a number.

    Show the answer to question 5

    Answer: 6 code points

    The word looks like three or four letters, but it is six code points: न, म, स, the virama that joins स to the next letter, त, and the vowel sign े. len() counts code points, not the letters you see.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.