Programming and IT
Python: reading a file
Working with a file takes three steps: open it, read or write, close it. You are better off not closing it by hand — that is what the `with` statement is for. And the parameter people forget most is the encoding: without it, the same code reads a file differently on different machines.
In this article
Opening a file properly: the with statement#
=
with closes the file for you — both when the block finishes normally and when
an exception is raised inside it. The manual version, f = open(...) followed by
f.close(), works too, but if an error happens between those lines the file
stays open and written data may never reach the disk.
The mode is the second argument:
| Mode | What it does |
|---|---|
r |
read, the default; no file — FileNotFoundError |
w |
write; an existing file is truncated |
a |
append to the end; the file is created if missing |
x |
create; file already exists — FileExistsError |
rb wb |
binary mode, works with bytes |
The most expensive beginner mistake is opening a file in w mode "just to have a
look": the contents vanish the moment it opens, before the first write.
Three ways to read#
= # everything as one string
= # a list of lines, newlines included
# one line at a time, any file size
The third way is the one to use by default. A file object is itself an iterator and hands out lines one by one without loading the whole file into memory: that is how you read a hundred lines and a gigabyte log alike.
Each line arrives with its trailing \n, so print(line) gives double line
breaks. Strip them with rstrip("\n") — with the argument, otherwise rstrip()
also cuts meaningful spaces at the end of the line.
=
This expression adds up the numbers in a file without building a list — a trick from the article on Python generators.
Writing and appending#
=
write does not add a newline — you add it yourself. Despite its name,
writelines does not add one either: it just writes the items back to back. So
for a list of strings, "\n".join(...) is handier, as in the third example — the
method itself is covered in the article on Python strings.
print can write straight to a file: print("text", file=f) — and it adds the
newline for you.
Encoding#
Without the encoding argument, Python uses the system's default encoding. On
Linux and macOS that is usually UTF-8, but on Windows it has historically been a
legacy code page such as cp1252. The same code then reads the file differently:
text on one machine, a UnicodeDecodeError or garbled characters on another.
The rule is simple: pass encoding="utf-8" in every open call for text. If a
file came from an old program and does not decode as UTF-8, try
encoding="cp1252" or encoding="latin-1". When there are a few broken bytes and
you do not want to lose the whole file, errors="replace" helps — undecodable
spots become a replacement character instead of a crash.
Binary mode has no encoding at all: open(path, "rb") returns bytes, and the
encoding argument cannot be combined with it.
Errors and the short form with pathlib#
=
=
FileNotFoundError is a normal situation that you catch; how catching works is
explained in the article on Python exceptions. Its
neighbours are PermissionError (no access rights) and IsADirectoryError (you
passed a folder).
For short operations the standard library has pathlib:
=
# hello
# True .txt notes
Path joins paths with /: Path("data") / "2026" / "log.txt" works the same on
every system, unlike gluing strings together with slashes.
Practice: create a file with ten lines, read it line by line and write only the lines longer than five characters to a new file, putting a line number at the start of each. Check that the original file is untouched.
Step-by-step plan
- Create and readWrite three lines to a file in w mode, then read it whole using with.
- Line-by-line loopGo through the file with a for loop and strip the trailing newline with rstrip("\n").
- AppendOpen the same file in a mode, add a line and make sure the old contents are still there.
- Catch a missing fileOpen a name that does not exist and handle FileNotFoundError.
- Repeat with pathlibDo the same with Path.read_text and Path.write_text, passing the encoding explicitly.
Start learning this in your own space
The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.
Check yourself
1.Which mode truncates an existing file when it is opened?
2.What does the with statement do when working with a file?
3.Which exception does open("no-such-file.txt") raise in r mode?
4.A file contains the lines "10", "20", "30". What is sum(int(line) for line in open("nums.txt", encoding="utf-8"))?
Sources
-
Reading and Writing Files in the Python tutorialThe Reading and Writing Files section: with, modes, methodsfree
-
The open function in the referenceEvery argument of open: mode, encoding, errors, newlinefree
-
The pathlib modulePaths as objects, read_text and write_textfree
Was this helpful?