Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Python Module 2 – Values, variables, numbers and strings

Essential string methods for cleaning text

Clean and search text with Python's string methods (strip, split, join, replace, find, startswith, casefold) and avoid the traps in strip() and isdigit().

  • Beginner
  • 25 minutes
  • Examples run with Python 3.14.8 and Pyodide 314.0.7
  • By MySmartCoPilot

What you will learn

  • Clean and split text with strip, split, join and replace
  • Search with find, index, startswith, endswith and in
  • Normalise case correctly with lower, upper and casefold

Before you start

On this page

Text that comes from people, files and web pages is rarely tidy. It has spaces at the ends, capitals in odd places, commas with gaps after them and the occasional empty field. Strings come with dozens of methods, functions you call with a dot after a string (text.strip()), and a handful of them do most of the cleaning and searching that real programs need.

One rule from the previous lesson shapes how you use all of them: strings never change, so a method gives you a new string and leaves the original alone. Keep the result (name = name.strip()), and chain methods when you need several steps: raw.strip().casefold() strips first, then folds the case of the stripped copy.

Clean and split: strip, split, join and replace

Here is a line of comma-separated fruit names typed by a person, cleaned step by step:

Cleaning a messy line Python · clean_line.py
raw = "  Mango, apple ,BANANA,, kiwi \n"

# strip() removes spaces, tabs and line breaks from both ends.
print(repr(raw.strip()))

# split(",") cuts at every comma and keeps the empty field between ",,".
fields = raw.strip().split(",")
print(fields)

# Clean each field in turn and keep the ones that are not empty.
names = []
for field in fields:
    name = field.strip().casefold()
    if name:
        names.append(name)
print(names)

# join() is the opposite of split(): one string, with the separator between the parts.
print(" | ".join(names))

# split() with no argument splits on any run of whitespace and drops empty parts.
print(" ".join("  too   many \t spaces ".split()))

Output

'Mango, apple ,BANANA,, kiwi'
['Mango', ' apple ', 'BANANA', '', ' kiwi']
['mango', 'apple', 'banana', 'kiwi']
mango | apple | banana | kiwi
too many spaces

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 clean_line.py

  • strip() removes whitespace (spaces, tabs, line breaks) from both ends; lstrip() and rstrip() do only the left or the right end. A line read from a file usually ends with \n, so line.rstrip("\n") or line.strip() is often the first thing a program does with it.
  • split(",") cuts the string at every comma and returns a list of the pieces. It keeps empty pieces: two commas in a row give an empty string between them, which is right for a spreadsheet row with an empty cell, and which is why the loop skips fields that are empty after stripping.
  • split() with no argument behaves differently: it splits on any run of whitespace and never returns empty pieces. " ".join(text.split()) is a quick way to collapse several spaces and tabs into one space.
  • join() is called on the separator and given the pieces: " | ".join(names) puts " | " between them. The pieces must all be strings; ", ".join([1, 2, 3]) raises TypeError, so convert numbers with str() first.
  • replace(old, new) replaces every match; replace(old, new, 1) replaces only the first.

The strip() trap

The argument of strip() is not a word to remove. It is a set of characters, and strip() keeps removing any of them from both ends until it meets a character that is not in the set:

strip() or removesuffix()? Python · strip_or_removesuffix.py
site = "mocha.com"
print(site.strip(".com"))
print(site.removesuffix(".com"))
print("comic.com".strip(".com"), "comic.com".removesuffix(".com"))
print("notes.txt".removesuffix(".pdf"))

Output

ha
mocha
i comic
notes.txt

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 strip_or_removesuffix.py

"mocha.com".strip(".com") removes m, o and c from the front and .com from the end, and comic.com loses almost everything. To remove an exact beginning or ending, use removeprefix() or removesuffix() (added in Python 3.9 by PEP 616). They remove it at most once, and return the string unchanged when it does not start or end that way.

Search: in, find, index, startswith and endswith

Searching a log line Python · search.py
line = "ERROR: disk full on /dev/sda1"

# in answers yes or no; it is the clearest test.
print("disk" in line, "DISK" in line)

# find() gives the position of the first match, or -1 when there is none.
print(line.find("disk"), line.find("cpu"))

# startswith() and endswith() take one text or a tuple of texts to try.
print(line.startswith("ERROR"), line.endswith(("sda1", "sdb1")))
print(line.count("s"), line.replace("disk", "drive"))

# A trap: a match at the very start is position 0, and 0 counts as false.
if line.find("ERROR"):
    print("found it")
else:
    print("not found? find() returned", line.find("ERROR"))

# index() is like find() but raises ValueError when there is no match.
print(line.index("cpu"))

Output (exit status 1)

True False
7 -1
True True
2 ERROR: drive full on /dev/sda1
not found? find() returned 0

Printed as an error (standard error)

Traceback (most recent call last):
  File "search.py", line 20, in <module>
    print(line.index("cpu"))
          ~~~~~~~~~~^^^^^^^
ValueError: substring not found

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 search.py

  • "disk" in line is the clearest way to ask whether one string contains another. Like every comparison of strings it is case-sensitive, so "DISK" in line is False.
  • find() returns the position where the match starts, or -1 when there is none. Use it when you need the position, for example to slice the text after it.
  • index() does the same but raises ValueError when there is no match, as the last line of the program shows. Use it when a missing match means something is wrong.
  • startswith() and endswith() check the beginning and the end, and accept a tuple of possibilities: name.endswith((".jpg", ".png")).
  • count() counts matches that do not overlap.

Warning

Never use the result of find() as a yes-or-no test. A match at the very start is position 0, which counts as false, and “no match” is -1, which counts as true, so if line.find("ERROR"): gets both cases wrong: the program above says “not found?” about a line that starts with ERROR. Write if "ERROR" in line: instead.

Case: lower, upper, casefold and title

Changing case Python · case.py
# lower() and casefold() both make lower case; casefold() also removes other case differences.
print("Straße".lower(), "Straße".casefold())
print("STRASSE".lower() == "Straße".lower())
print("STRASSE".casefold() == "Straße".casefold())

print("ravi KUMAR".upper(), "ravi KUMAR".title(), "ravi KUMAR".capitalize())

# title() starts a new word after every character that is not a letter, apostrophes included.
print("o'brien's cafe".title())

Output

straße strasse
False
True
RAVI KUMAR Ravi Kumar Ravi kumar
O'Brien'S Cafe

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 case.py

lower() and upper() put every letter in lower or upper case. To compare text whatever its case, use casefold() on both sides: it removes more differences than lower(). The German street name in the program shows one of them. lower() turns "Straße" into "straße", keeping the ß because that letter is lower case already, so the result does not match "STRASSE".lower(), which is "strasse". casefold() writes the ß as ss, so both sides fold to "strasse" and the comparison is True. Many languages have letters like this one, and that is why caseless comparison needs casefold().

title() starts each word with a capital and puts the rest of the word in lower case ("ravi KUMAR" becomes "Ravi Kumar"), but to title() a word is simply a run of letters. So an apostrophe starts a new word, and "o'brien's cafe" becomes "O'Brien'S Cafe". That suits many names and few sentences. capitalize() gives only the first character of the whole string a capital and makes the rest lower case.

Case Converter Paste a sentence to compare upper, lower, title and other case styles side by side.

Which characters count as digits

Three methods test for digits, and they disagree about some characters:

isdecimal(), isdigit() and isnumeric() Python · digits.py
import unicodedata

print("text  isdecimal  isdigit  isnumeric  int(text)  name")
for text in ["7", "٧", "७", "²", "½"]:
    try:
        number = int(text)
    except ValueError:
        number = "ValueError"
    print(f"{text:4}  {text.isdecimal()!s:9}  {text.isdigit()!s:7}  {text.isnumeric()!s:9}  {number!s:10} {unicodedata.name(text)}")

Output

text  isdecimal  isdigit  isnumeric  int(text)  name
7     True       True     True       7          DIGIT SEVEN
٧     True       True     True       7          ARABIC-INDIC DIGIT SEVEN
७     True       True     True       7          DEVANAGARI DIGIT SEVEN
²     False      True     True       ValueError SUPERSCRIPT TWO
½     False      False    True       ValueError VULGAR FRACTION ONE HALF

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 digits.py

  • isdecimal() is the strictest: it accepts only characters that can make up a base-10 number. That includes the digits of other scripts, such as the Arabic-Indic ٧ and the Devanagari ७, and int() reads those too.
  • isdigit() also accepts digits that cannot form a number, such as the superscript ².
  • isnumeric() accepts anything with a numeric value, even the fraction ½.

None of them accepts a minus sign, a decimal point or spaces at the ends, all of which int() or float() handle. So to check that text is a number, try to convert it and catch the error: put int(text) in a try block and handle ValueError, as the program does. The input and output lesson uses this to read numbers safely.

Many replacements at once: translate

replace() handles one substitution per call. When you need to change or delete many different characters, build a table once with str.maketrans() and apply it with translate(), which goes through the string a single time:

Character tables with maketrans() Python · translate.py
# maketrans() with a third argument lists characters to delete.
phone = "(022) 2345-6789"
digits_only = str.maketrans("", "", "() -")
print(phone.translate(digits_only))

# With two strings of the same length, each character of the first
# becomes the character at the same position in the second.
devanagari_digits = str.maketrans("०१२३४५६७८९", "0123456789")
print("Flat ४०५, floor ३".translate(devanagari_digits))

# A dictionary can map a character to a longer string, or to None to delete it.
safe = str.maketrans({"&": "and", "#": None})
print("Salt & pepper #1".translate(safe))

Output

02223456789
Flat 405, floor 3
Salt and pepper 1

Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 translate.py

Whitespace Remover Trim the lines of pasted text and collapse runs of spaces and tabs, the way strip() and split() tidy text in code.

Key takeaways

  • String methods return new strings, so assign the result; chaining (raw.strip().casefold()) applies them in order.
  • strip() cleans the ends; split(",") keeps empty fields while split() splits on runs of whitespace; join() glues string pieces with a separator.
  • strip(".com") removes any of those characters; removeprefix() and removesuffix() remove an exact text once.
  • Test containment with in. find() returns -1 for no match and 0 for a match at the start, so it is not a yes-or-no test; index() raises ValueError instead.
  • Compare text without case with casefold() on both sides.
  • isdecimal(), isdigit() and isnumeric() accept different characters; to validate a number, convert it with int() and catch ValueError.

Exercise

Exercise · Easy · Python

Tidy up a name that someone typed

Names typed into a form arrive with extra spaces, tabs and capitals in odd places. Write normalise_name(raw), which returns the name tidied up by these rules:

  • No spaces, tabs or line breaks at either end, and exactly one space between words.
  • Each word written with str.title(), so "aNNE-marie" becomes "Anne-Marie" and "o'neil" becomes "O'Neil".
  • The particles listed in PARTICLES (da, de, del, der, di, la, le, van and von) in lower case however they were typed, except when one is the first word: "ludwig VAN beethoven" becomes "Ludwig van Beethoven", but "van gogh" becomes "Van Gogh".
  • A name with no words at all, such as "" or " ", gives "".

Names in scripts without capital letters, such as "राम कुमार", only have their spaces tidied: title() leaves their letters as they are. The sample tests import your function from names.py.

Starter code · names.py

PARTICLES = ("da", "de", "del", "der", "di", "la", "le", "van", "von")


def normalise_name(raw):
    """Return raw as a tidy name: single spaces, each word capitalised, particles such as 'van' in lower case."""
    # Replace this line with your code.
    return raw
The sample tests · test_names.py
from names import normalise_name


def test_spaces():
    """removes spaces at the ends and keeps one space between words"""
    assert normalise_name("  ada   lovelace ") == "Ada Lovelace"
    assert normalise_name("grace\thopper\n") == "Grace Hopper"


def test_capitals():
    """writes each word with a capital letter first and the rest in lower case"""
    assert normalise_name("aLAN tURING") == "Alan Turing"
    assert normalise_name("anne-marie o'neil") == "Anne-Marie O'Neil"


def test_particles():
    """keeps particles after the first word in lower case"""
    assert normalise_name("ludwig VAN beethoven") == "Ludwig van Beethoven"
    assert normalise_name("maría DEL carmen") == "María del Carmen"
    assert normalise_name("Vincent van Gogh") == "Vincent van Gogh"


def test_first_word():
    """capitalises a particle that is the first word"""
    assert normalise_name("van gogh") == "Van Gogh"
    assert normalise_name("de la cruz") == "De la Cruz"


def test_other_scripts():
    """tidies the spaces of names in other scripts"""
    assert normalise_name("  राम   कुमार  ") == "राम कुमार"
    assert normalise_name("ÉMILE zola") == "Émile Zola"


def test_blank():
    """gives an empty string for a name with no words"""
    assert normalise_name("") == ""
    assert normalise_name(" \t ") == ""
A hint

raw.split() with no argument already deals with the spaces: it drops them at both ends and treats any run of spaces, tabs and line breaks as one separator, so " ".join(raw.split()) puts single spaces back. Then go through the words one at a time and decide for each: is it the first word, and is word.casefold() one of PARTICLES?

The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.

Check yourself

5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.

  1. Question 1 of 5 What does strip_or_removesuffix.py print?

    What does this program print? Choose one answer.

    site = "mocha.com"
    print(site.strip(".com"))
    print(site.removesuffix(".com"))
    print("comic.com".strip(".com"), "comic.com".removesuffix(".com"))
    print("notes.txt".removesuffix(".pdf"))
    Show the answer to question 1

    Answer: it prints

    ha
    mocha
    i comic
    notes.txt

    strip(".com") treats its argument as a set of characters and removes any of ., c, o and m from both ends, so it eats into the words themselves. removesuffix(".com") removes that exact ending once, and leaves the string unchanged when it does not end that way, which is why notes.txt keeps its .txt.

  2. Question 2 of 5 What does this print?

    Read the code, then choose one answer.

    print("  red,  green  ".split())
    Show the answer to question 2

    Answer: ['red,', 'green']

    With no argument, split() splits on runs of whitespace and drops the empty pieces at the ends. It does not know about commas, so the comma stays attached to red. split(",") would split on the comma instead and keep the spaces.

  3. Question 3 of 5 Which test is true exactly when the text "error" appears somewhere in line?

    Choose one answer.

    Show the answer to question 3

    Answer: "error" in line

    in gives True or False directly. find() returns -1, which is true, when there is no match and 0, which is false, for a match at the start, so line.find("error") is wrong both ways, and > 0 still misses a match at position 0. index() raises ValueError instead of returning a false value.

  4. Question 4 of 5 A user types a city name and you want to compare it with a stored one, ignoring case. Which is the most reliable?

    Choose one answer.

    Show the answer to question 4

    Answer: typed.casefold() == stored.casefold()

    casefold() is meant for caseless comparison: it removes more case differences than lower(), for example it turns ß into ss, so "STRASSE" and "Straße" match. Both sides must be folded; comparing a changed copy of one side with the unchanged other side fails whenever the stored text has different capitals.

  5. Question 5 of 5 text is the superscript two, "²". Which of these are true?

    Choose every answer that is right.

    Show the answer to question 5

    Answer:

    • text.isnumeric() is True
    • text.isdigit() is True

    A superscript digit counts as a digit and as numeric, but it is not a decimal character, the kind that can make up a base-10 number, so isdecimal() is False and int("²") raises ValueError. To check that text is a whole number, try int() and catch the ValueError.

References

Related tools

Report a problem with this lesson

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.