Programming and IT

Python: reading a file

Working with a file takes three steps: open it, read or write, close it. You are better off not closing it by hand — that is what the `with` statement is for. And the parameter people forget most is the encoding: without it, the same code reads a file differently on different machines.

Updated
In this article

Opening a file properly: the with statement#

with open("notes.txt", encoding="utf-8") as f:
    content = f.read()
print(content)

with closes the file for you — both when the block finishes normally and when an exception is raised inside it. The manual version, f = open(...) followed by f.close(), works too, but if an error happens between those lines the file stays open and written data may never reach the disk.

The mode is the second argument:

Mode What it does
r read, the default; no file — FileNotFoundError
w write; an existing file is truncated
a append to the end; the file is created if missing
x create; file already exists — FileExistsError
rb wb binary mode, works with bytes

The most expensive beginner mistake is opening a file in w mode "just to have a look": the contents vanish the moment it opens, before the first write.

Three ways to read#

with open("notes.txt", encoding="utf-8") as f:
    whole = f.read()          # everything as one string

with open("notes.txt", encoding="utf-8") as f:
    lines = f.readlines()     # a list of lines, newlines included

with open("notes.txt", encoding="utf-8") as f:
    for line in f:            # one line at a time, any file size
        print(line.rstrip("\n"))

The third way is the one to use by default. A file object is itself an iterator and hands out lines one by one without loading the whole file into memory: that is how you read a hundred lines and a gigabyte log alike.

Each line arrives with its trailing \n, so print(line) gives double line breaks. Strip them with rstrip("\n") — with the argument, otherwise rstrip() also cuts meaningful spaces at the end of the line.

with open("nums.txt", encoding="utf-8") as f:
    total = sum(int(line) for line in f if line.strip())
print(total)

This expression adds up the numbers in a file without building a list — a trick from the article on Python generators.

Writing and appending#

with open("out.txt", "w", encoding="utf-8") as f:
    f.write("first line\n")
    f.write("second line\n")

with open("out.txt", "a", encoding="utf-8") as f:
    f.write("added later\n")

rows = ["one", "two", "three"]
with open("rows.txt", "w", encoding="utf-8") as f:
    f.write("\n".join(rows))

write does not add a newline — you add it yourself. Despite its name, writelines does not add one either: it just writes the items back to back. So for a list of strings, "\n".join(...) is handier, as in the third example — the method itself is covered in the article on Python strings.

print can write straight to a file: print("text", file=f) — and it adds the newline for you.

Encoding#

Without the encoding argument, Python uses the system's default encoding. On Linux and macOS that is usually UTF-8, but on Windows it has historically been a legacy code page such as cp1252. The same code then reads the file differently: text on one machine, a UnicodeDecodeError or garbled characters on another.

The rule is simple: pass encoding="utf-8" in every open call for text. If a file came from an old program and does not decode as UTF-8, try encoding="cp1252" or encoding="latin-1". When there are a few broken bytes and you do not want to lose the whole file, errors="replace" helps — undecodable spots become a replacement character instead of a crash.

with open("legacy.txt", encoding="cp1252") as f:
    print(f.read())

Binary mode has no encoding at all: open(path, "rb") returns bytes, and the encoding argument cannot be combined with it.

Errors and the short form with pathlib#

try:
    with open("no-such-file.txt", encoding="utf-8") as f:
        data = f.read()
except FileNotFoundError:
    data = ""
    print("no file, carrying on with empty data")

FileNotFoundError is a normal situation that you catch; how catching works is explained in the article on Python exceptions. Its neighbours are PermissionError (no access rights) and IsADirectoryError (you passed a folder).

For short operations the standard library has pathlib:

from pathlib import Path

p = Path("notes.txt")
p.write_text("hello\n", encoding="utf-8")
print(p.read_text(encoding="utf-8"))   # hello
print(p.exists(), p.suffix, p.stem)    # True .txt notes

Path joins paths with /: Path("data") / "2026" / "log.txt" works the same on every system, unlike gluing strings together with slashes.

Practice: create a file with ten lines, read it line by line and write only the lines longer than five characters to a new file, putting a line number at the start of each. Check that the original file is untouched.

Step-by-step plan

  1. Create and readWrite three lines to a file in w mode, then read it whole using with.
  2. Line-by-line loopGo through the file with a for loop and strip the trailing newline with rstrip("\n").
  3. AppendOpen the same file in a mode, add a line and make sure the old contents are still there.
  4. Catch a missing fileOpen a name that does not exist and handle FileNotFoundError.
  5. Repeat with pathlibDo the same with Path.read_text and Path.write_text, passing the encoding explicitly.

Start learning this in your own space

The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.

Start the plan

Check yourself

1.Which mode truncates an existing file when it is opened?

2.What does the with statement do when working with a file?

3.Which exception does open("no-such-file.txt") raise in r mode?

4.A file contains the lines "10", "20", "30". What is sum(int(line) for line in open("nums.txt", encoding="utf-8"))?

Sources

Was this helpful?

More articles

Programming and IT Python: decorators A decorator is a function that takes another function and returns a new one with extra behaviour. The `@` sign above a definition is just a shorthand for an assignment. Keep that in mind and the whole topic takes one evening. Programming and IT git checkout Historically `git checkout` does several unrelated things at once: it moves your working copy to another branch, creates a branch and pulls a single file out of a commit. Because it was so overloaded, Git 2.23 added two separate commands — `switch` and `restore`. Here are both sides: the old way and its modern replacement. Programming and IT SQL queries: examples explained An SQL query describes which rows you need, not how to find them. Below, the same two tables go through all the main constructs of the language — from simple filtering to joins and subqueries — and every query comes with its result down to the last row. Programming and IT How to learn Python from scratch Python is a good first programming language: code reads almost like text, and the standard library covers most everyday tasks. This plan takes you from installing the interpreter to your own scripts covered by tests in about four months, at roughly an hour a day. Programming and IT Markdown tables A Markdown table is built from pipes and a separator line under the header. Below is the table syntax with column alignment, plus a short cheat sheet for the rest of the markup — headings, lists, links, images, code, quotes and task lists. Programming and IT Python: classes A class describes what data an object is made of and what it can do. The `__init__` constructor fills in a new object, and `self` is a reference to the object itself. Below: a minimal working class, the difference between class and instance attributes, and ways to make the code shorter.

More solutions