Python Module 2 – Values, variables, numbers and strings
Essential string methods for cleaning text
Clean and search text with Python's string methods (strip, split, join, replace, find, startswith, casefold) and avoid the traps in strip() and isdigit().
What you will learn
- Clean and split text with strip, split, join and replace
- Search with find, index, startswith, endswith and in
- Normalise case correctly with lower, upper and casefold
Before you start
On this page
Text that comes from people, files and web pages is rarely tidy. It has spaces at the ends, capitals in odd places,
commas with gaps after them and the occasional empty field. Strings come with dozens of methods, functions you
call with a dot after a string (text.strip()), and a handful of them do most of the cleaning and searching that
real programs need.
One rule from the previous lesson shapes how you use all of them: strings never change, so a method gives you a new
string and leaves the original alone. Keep the result (name = name.strip()), and chain methods when you need
several steps: raw.strip().casefold() strips first, then folds the case of the stripped copy.
Clean and split: strip, split, join and replace
Here is a line of comma-separated fruit names typed by a person, cleaned step by step:
raw = " Mango, apple ,BANANA,, kiwi \n"
# strip() removes spaces, tabs and line breaks from both ends.
print(repr(raw.strip()))
# split(",") cuts at every comma and keeps the empty field between ",,".
fields = raw.strip().split(",")
print(fields)
# Clean each field in turn and keep the ones that are not empty.
names = []
for field in fields:
name = field.strip().casefold()
if name:
names.append(name)
print(names)
# join() is the opposite of split(): one string, with the separator between the parts.
print(" | ".join(names))
# split() with no argument splits on any run of whitespace and drops empty parts.
print(" ".join(" too many \t spaces ".split())) Output
'Mango, apple ,BANANA,, kiwi' ['Mango', ' apple ', 'BANANA', '', ' kiwi'] ['mango', 'apple', 'banana', 'kiwi'] mango | apple | banana | kiwi too many spaces
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 clean_line.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
strip()removes whitespace (spaces, tabs, line breaks) from both ends;lstrip()andrstrip()do only the left or the right end. A line read from a file usually ends with\n, soline.rstrip("\n")orline.strip()is often the first thing a program does with it.split(",")cuts the string at every comma and returns a list of the pieces. It keeps empty pieces: two commas in a row give an empty string between them, which is right for a spreadsheet row with an empty cell, and which is why the loop skips fields that are empty after stripping.split()with no argument behaves differently: it splits on any run of whitespace and never returns empty pieces." ".join(text.split())is a quick way to collapse several spaces and tabs into one space.join()is called on the separator and given the pieces:" | ".join(names)puts" | "between them. The pieces must all be strings;", ".join([1, 2, 3])raisesTypeError, so convert numbers withstr()first.replace(old, new)replaces every match;replace(old, new, 1)replaces only the first.
The strip() trap
The argument of strip() is not a word to remove. It is a set of characters, and strip() keeps removing any of
them from both ends until it meets a character that is not in the set:
site = "mocha.com"
print(site.strip(".com"))
print(site.removesuffix(".com"))
print("comic.com".strip(".com"), "comic.com".removesuffix(".com"))
print("notes.txt".removesuffix(".pdf")) Output
ha mocha i comic notes.txt
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 strip_or_removesuffix.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
"mocha.com".strip(".com") removes m, o and c from the front and .com from the end, and comic.com loses
almost everything. To remove an exact beginning or ending, use removeprefix() or removesuffix() (added in Python
3.9 by PEP 616). They remove it at most once, and return the string unchanged when it does not start or end that way.
Search: in, find, index, startswith and endswith
line = "ERROR: disk full on /dev/sda1"
# in answers yes or no; it is the clearest test.
print("disk" in line, "DISK" in line)
# find() gives the position of the first match, or -1 when there is none.
print(line.find("disk"), line.find("cpu"))
# startswith() and endswith() take one text or a tuple of texts to try.
print(line.startswith("ERROR"), line.endswith(("sda1", "sdb1")))
print(line.count("s"), line.replace("disk", "drive"))
# A trap: a match at the very start is position 0, and 0 counts as false.
if line.find("ERROR"):
print("found it")
else:
print("not found? find() returned", line.find("ERROR"))
# index() is like find() but raises ValueError when there is no match.
print(line.index("cpu")) Output (exit status 1)
True False 7 -1 True True 2 ERROR: drive full on /dev/sda1 not found? find() returned 0
Printed as an error (standard error)
Traceback (most recent call last):
File "search.py", line 20, in <module>
print(line.index("cpu"))
~~~~~~~~~~^^^^^^^
ValueError: substring not found
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 search.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
"disk" in lineis the clearest way to ask whether one string contains another. Like every comparison of strings it is case-sensitive, so"DISK" in lineisFalse.find()returns the position where the match starts, or -1 when there is none. Use it when you need the position, for example to slice the text after it.index()does the same but raisesValueErrorwhen there is no match, as the last line of the program shows. Use it when a missing match means something is wrong.startswith()andendswith()check the beginning and the end, and accept a tuple of possibilities:name.endswith((".jpg", ".png")).count()counts matches that do not overlap.
Warning
Never use the result of find() as a yes-or-no test. A match at the very start is position 0, which counts as
false, and “no match” is -1, which counts as true, so if line.find("ERROR"): gets both cases wrong: the program
above says “not found?” about a line that starts with ERROR. Write if "ERROR" in line: instead.
Case: lower, upper, casefold and title
# lower() and casefold() both make lower case; casefold() also removes other case differences.
print("Straße".lower(), "Straße".casefold())
print("STRASSE".lower() == "Straße".lower())
print("STRASSE".casefold() == "Straße".casefold())
print("ravi KUMAR".upper(), "ravi KUMAR".title(), "ravi KUMAR".capitalize())
# title() starts a new word after every character that is not a letter, apostrophes included.
print("o'brien's cafe".title()) Output
straße strasse False True RAVI KUMAR Ravi Kumar Ravi kumar O'Brien'S Cafe
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 case.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
lower() and upper() put every letter in lower or upper case. To compare text whatever its case, use casefold()
on both sides: it removes more differences than lower(). The German street name in the program shows one of them.
lower() turns "Straße" into "straße", keeping the ß because that letter is lower case already, so the result
does not match "STRASSE".lower(), which is "strasse". casefold() writes the ß as ss, so both sides fold to
"strasse" and the comparison is True. Many languages have letters like this one, and that is why caseless
comparison needs casefold().
title() starts each word with a capital and puts the rest of the word in lower case ("ravi KUMAR" becomes
"Ravi Kumar"), but to title() a word is simply a run of letters. So an apostrophe starts a new word, and
"o'brien's cafe" becomes "O'Brien'S Cafe". That suits many names and few sentences. capitalize() gives only the
first character of the whole string a capital and makes the rest lower case.
Which characters count as digits
Three methods test for digits, and they disagree about some characters:
import unicodedata
print("text isdecimal isdigit isnumeric int(text) name")
for text in ["7", "٧", "७", "²", "½"]:
try:
number = int(text)
except ValueError:
number = "ValueError"
print(f"{text:4} {text.isdecimal()!s:9} {text.isdigit()!s:7} {text.isnumeric()!s:9} {number!s:10} {unicodedata.name(text)}") Output
text isdecimal isdigit isnumeric int(text) name 7 True True True 7 DIGIT SEVEN ٧ True True True 7 ARABIC-INDIC DIGIT SEVEN ७ True True True 7 DEVANAGARI DIGIT SEVEN ² False True True ValueError SUPERSCRIPT TWO ½ False False True ValueError VULGAR FRACTION ONE HALF
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 digits.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
isdecimal()is the strictest: it accepts only characters that can make up a base-10 number. That includes the digits of other scripts, such as the Arabic-Indic٧and the Devanagari७, andint()reads those too.isdigit()also accepts digits that cannot form a number, such as the superscript².isnumeric()accepts anything with a numeric value, even the fraction½.
None of them accepts a minus sign, a decimal point or spaces at the ends, all of which int() or float() handle.
So to check that text is a number, try to convert it and catch the error: put int(text) in a try block and handle
ValueError, as the program does. The input and output lesson uses this to read numbers safely.
Many replacements at once: translate
replace() handles one substitution per call. When you need to change or delete many different characters, build a
table once with str.maketrans() and apply it with translate(), which goes through the string a single time:
# maketrans() with a third argument lists characters to delete.
phone = "(022) 2345-6789"
digits_only = str.maketrans("", "", "() -")
print(phone.translate(digits_only))
# With two strings of the same length, each character of the first
# becomes the character at the same position in the second.
devanagari_digits = str.maketrans("०१२३४५६७८९", "0123456789")
print("Flat ४०५, floor ३".translate(devanagari_digits))
# A dictionary can map a character to a longer string, or to None to delete it.
safe = str.maketrans({"&": "and", "#": None})
print("Salt & pepper #1".translate(safe)) Output
02223456789 Flat 405, floor 3 Salt and pepper 1
Recorded with Python 3.14.8 on macOS 26 arm64. To run it yourself: mise exec python@3.14.8 -- python3 translate.py
Runs on this device, in your browser. The first run downloads Python (about 13.5 MB), which is kept for the next runs.
Your run, in this browser
Key takeaways
- String methods return new strings, so assign the result; chaining (
raw.strip().casefold()) applies them in order. strip()cleans the ends;split(",")keeps empty fields whilesplit()splits on runs of whitespace;join()glues string pieces with a separator.strip(".com")removes any of those characters;removeprefix()andremovesuffix()remove an exact text once.- Test containment with
in.find()returns -1 for no match and 0 for a match at the start, so it is not a yes-or-no test;index()raisesValueErrorinstead. - Compare text without case with
casefold()on both sides. isdecimal(),isdigit()andisnumeric()accept different characters; to validate a number, convert it withint()and catchValueError.
Exercise
Exercise · Easy · Python
Tidy up a name that someone typed
Names typed into a form arrive with extra spaces, tabs and capitals in odd places. Write normalise_name(raw), which returns the name tidied up by these rules:
- No spaces, tabs or line breaks at either end, and exactly one space between words.
- Each word written with
str.title(), so"aNNE-marie"becomes"Anne-Marie"and"o'neil"becomes"O'Neil". - The particles listed in
PARTICLES(da, de, del, der, di, la, le, van and von) in lower case however they were typed, except when one is the first word:"ludwig VAN beethoven"becomes"Ludwig van Beethoven", but"van gogh"becomes"Van Gogh". - A name with no words at all, such as
""or" ", gives"".
Names in scripts without capital letters, such as "राम कुमार", only have their spaces tidied: title() leaves their letters as they are. The sample tests import your function from names.py.
Starter code · names.py
PARTICLES = ("da", "de", "del", "der", "di", "la", "le", "van", "von")
def normalise_name(raw):
"""Return raw as a tidy name: single spaces, each word capitalised, particles such as 'van' in lower case."""
# Replace this line with your code.
return raw The sample tests · test_names.py
from names import normalise_name
def test_spaces():
"""removes spaces at the ends and keeps one space between words"""
assert normalise_name(" ada lovelace ") == "Ada Lovelace"
assert normalise_name("grace\thopper\n") == "Grace Hopper"
def test_capitals():
"""writes each word with a capital letter first and the rest in lower case"""
assert normalise_name("aLAN tURING") == "Alan Turing"
assert normalise_name("anne-marie o'neil") == "Anne-Marie O'Neil"
def test_particles():
"""keeps particles after the first word in lower case"""
assert normalise_name("ludwig VAN beethoven") == "Ludwig van Beethoven"
assert normalise_name("maría DEL carmen") == "María del Carmen"
assert normalise_name("Vincent van Gogh") == "Vincent van Gogh"
def test_first_word():
"""capitalises a particle that is the first word"""
assert normalise_name("van gogh") == "Van Gogh"
assert normalise_name("de la cruz") == "De la Cruz"
def test_other_scripts():
"""tidies the spaces of names in other scripts"""
assert normalise_name(" राम कुमार ") == "राम कुमार"
assert normalise_name("ÉMILE zola") == "Émile Zola"
def test_blank():
"""gives an empty string for a name with no words"""
assert normalise_name("") == ""
assert normalise_name(" \t ") == "" A hint
raw.split() with no argument already deals with the spaces: it drops them at both ends and treats any run of spaces, tabs and line breaks as one separator, so " ".join(raw.split()) puts single spaces back. Then go through the words one at a time and decide for each: is it the first word, and is word.casefold() one of PARTICLES?
Results of the sample tests
| Test | Result | Details |
|---|
What your code printed
The sample tests run on this device, in your browser (Pyodide): nothing is sent to mysmartcopilot.com. The first run downloads Python (about 13.5 MB), which is kept for the next runs. A check in your browser is feedback for you, not proof that the code is right for every input.
Check yourself
5 questions about this lesson. Every answer and why it is right is on the page, behind “Show the answer”. Your score stays in this browser.
References
- String Methods (The Python Standard Library) (Python Software Foundation)
- str.casefold() (The Python Standard Library) (Python Software Foundation)
- str.removeprefix() and str.removesuffix() (The Python Standard Library) (Python Software Foundation)
- PEP 616: String methods to remove prefixes and suffixes (Python Software Foundation)
- int() (Built-in Functions) (Python Software Foundation)
Related tools
Report a problem with this lesson
Kept only in this browser. Your Learn progress