%%{init: {"theme": "base", "themeVariables": {"fontFamily": "Arial"}}}%%
flowchart TD
Repo["pandas-dev / pandas"] --> Core["pandas/core/ (Core Logic & Algorithms)"]
Repo --> Tests["pandas/tests/ (Massive Automated Test Suite)"]
Repo --> CI[".github/workflows/ (Continuous Integration & Pre-commit)"]
classDef dir fill:#e0e7ff,stroke:#4338ca,stroke-width:2px,color:#000
class Repo,Core,Tests,CI dir
Code Quality
Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it.
— Brian W. Kernighan
In research computing, code often grows organically: a script starts as a quick experiment, expands into several functions, and eventually produces figures for a manuscript. But how do we ensure that changes made months later don’t subtly corrupt scientific calculations? How do we maintain readability when collaborating with laboratory colleagues?
In this session, we establish habits for long-term code health in our bookstats project: 1. Automated formatting and linting with Ruff to keep styling uniform and catch syntax bugs instantly. 2. Git pre-commit hooks with pre-commit to automatically intervene before non-compliant code enters version control. 3. Structured interface documentation using NumPy-style docstrings. 4. Test-Driven Development (TDD) using pytest to build confidence through the red-green-refactor cycle. 5. Descriptive Zipf analysis implemented via TDD and integrated into our interactive Marimo visualization.
The Pandas Repository Microscope
Before writing tests for our own project, let’s look at how mature, world-class scientific software is organized.
Open the open-source repository for pandas-dev/pandas on GitHub in your browser:
Inspect the pandas/tests/ directory. Notice: - Nearly every single module in pandas/core/ has a corresponding test module in pandas/tests/. - There are thousands of tests exercising edge cases, NaN values, empty series, and unexpected datatypes. - Automated tests are not an afterthought written when a paper is submitted; they are the architectural scaffolding that allows hundreds of contributors across the world to modify pandas without breaking existing user code.
Our bookstats package follows this exact software engineering principle.
Create the Feature Branch
Following our collaborative Git workflow, we isolate our code quality additions on a dedicated feature branch.
In VS Code or your terminal:
git switch main
git pull
git switch -c add-pre-commitVerify your branch with git branch --show-current. It should display add-pre-commit.
Code Quality with Ruff
Consistency in formatting and linting saves countless hours of code review debates. Instead of manual style checking, we use Ruff, an extremely fast Python linter and formatter written in Rust.
Add Ruff as a development dependency using uv:
uv add --dev ruffConfiguring Ruff and NumPy Docstrings
We configure Ruff in pyproject.toml. Open pyproject.toml and configure the linter to enforce standard formatting and NumPy-style docstrings (pydocstyle convention):
[tool.ruff]
line-length = 88
target-version = "py312"
[tool.ruff.lint]
extend-select = [
"E", # pycodestyle errors
"F", # pyflakes
"I", # isort import sorting
"D", # pydocstyle docstring conventions
"UP", # pyupgrade modern Python syntax
]
[tool.ruff.lint.pydocstyle]
convention = "numpy"In scientific computing, functions often accept matrices, dataframes, arrays, and numerical thresholds. NumPy-style docstrings provide dedicated, human-readable sections (Parameters, Returns, Notes, Examples) that make function interfaces immediately understandable.
Check your codebase with Ruff:
# Check code for linting issues
uv run ruff check .
# Check formatting without modifying files
uv run ruff format --check .To automatically format files, run:
uv run ruff format .Automated Guardrails: Pre-commit Hooks
Linters and formatters only protect a codebase if contributors remember to run them before committing. A pre-commit hook is a script that Git executes automatically whenever you run git commit. If the check fails or modifies code, Git aborts the commit, giving you an immediate chance to review and fix the issue.
Add pre-commit to development dependencies:
uv add --dev pre-commitCreate a configuration file named .pre-commit-config.yaml in the root of your repository:
repos:
- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.9.9
hooks:
- id: ruff
args: [--fix]
- id: ruff-formatInstall the Git hook into your local .git/hooks/ directory:
uv run pre-commit installYou should see: pre-commit installed at .git/hooks/pre-commit.
Experiencing a Pre-commit Intervention
Let’s test our automated guardrail. Intentionally introduce poorly formatted code with missing documentation.
In src/bookstats/counts.py, add this unformatted dummy function at the bottom:
def messy_sample( x,y ):
z=x+y
return zNow stage and attempt to commit the change:
git add src/bookstats/counts.py
git commit -m "feat: add sample function"Watch what happens in your terminal:
ruff.....................................................................Failed
- hook id: ruff
- files were modified by this hook
src/bookstats/counts.py:210:1: D103 Missing docstring in public function
ruff-format..............................................................Failed
- hook id: ruff-format
- files were modified by this hook
1 file reformatted
Git blocked the commit! 1. ruff-format automatically fixed the spacing: def messy_sample(x, y):. 2. ruff flagged that our public function is missing a docstring (D103).
Inspect the diff in VS Code Source Control. You will see the reformatted layout. Now write a proper NumPy-style docstring explaining what the function does:
def messy_sample(x: int, y: int) -> int:
"""Compute the sum of two integers for testing pre-commit.
Parameters
----------
x : int
First integer addend.
y : int
Second integer addend.
Returns
-------
int
Sum of x and y.
"""
z = x + y
return zRemove or keep the sample function, stage the clean changes, and commit again:
git add src/bookstats/counts.py
git commit -m "feat: configure ruff and pre-commit hooks"This time, all hooks pass with green checkmarks, and Git completes the commit!
Push your branch and merge via Pull Request into main:
git push -u origin add-pre-commitMerge the PR on GitHub, switch back to main, and pull:
git switch main
git pullTest-Driven Development (TDD) and Unit Testing
Now that our code style is automated, how do we verify that our logic behaves as intended?
“Smash first, analyze later.”
— Advice from a CERN particle physicist
In Test-Driven Development (TDD), we invert the traditional habit of writing code first and testing later:
%%{init: {"flowchart": {"curve": "basis"}, "theme": "base", "themeVariables": {"fontFamily": "Arial"}}}%%
flowchart LR
subgraph Red["1. RED"]
T1["Write failing test"]
end
subgraph Green["2. GREEN"]
T2["Write minimal code to pass"]
end
subgraph Refactor["3. REFACTOR"]
T3["Clean up & improve structure"]
end
T1 ==> T2 ==> T3 ==> T1
classDef redBox fill:#fee2e2,stroke:#ef4444,stroke-width:2px,color:#000
classDef greenBox fill:#dcfce7,stroke:#22c55e,stroke-width:2px,color:#000
classDef blueBox fill:#e0e7ff,stroke:#6366f1,stroke-width:2px,color:#000
class T1 redBox
class T2 greenBox
class T3 blueBox
- Red: Write a small, clear test specifying the desired behavior. Run it and watch it fail.
- Green: Write the simplest possible code to make the test pass.
- Refactor: Clean up the implementation while using the passing test as a safety harness.
Setting Up Pytest
Create a new feature branch:
git switch -c add-testsAdd pytest as a development dependency:
uv add --dev pytestConfigure pytest in pyproject.toml so it recognizes src/ and your tests directory:
[tool.pytest.ini_options]
testpaths = ["tests"]
pythonpath = ["src"]Writing Your First Unit Test (Red Phase)
A unit test verifies one isolated function with known inputs and expected outputs.
Create tests/test_counts.py with this initial test:
"""Unit tests for word extraction and counting functions."""
from bookstats.counts import extract_words
def test_extract_words_normalizes_text():
assert extract_words("Hello, HELLO! World?") == [
"hello",
"hello",
"world",
]Run pytest:
uv run pytest -vIf your extraction function does not yet strip punctuation or convert to lower case, pytest will print a clear failure:
FAILED tests/test_counts.py::test_extract_words_normalizes_text - AssertionError
Commit this failing test to record the test-driven requirement:
git add tests/test_counts.py
git commit -m "test: add failing unit test for text normalization"Making the Test Pass (Green Phase)
Now modify src/bookstats/counts.py to ensure extract_words handles case and punctuation cleanly:
import re
def extract_words(text: str) -> list[str]:
"""Extract lowercase words from text, stripping punctuation.
Parameters
----------
text : str
Input document text.
Returns
-------
list of str
List of lowercase alphabetic tokens.
"""
return re.findall(r"\b[a-zA-Z]+\b", text.lower())Run pytest again:
uv run pytest -vtests/test_counts.py::test_extract_words_normalizes_text PASSED [100%]
It passes! Commit the passing implementation:
git add src/bookstats/counts.py
git commit -m "feat: normalize words to lowercase and strip punctuation"Adapting Additional Unit Tests
Now extend tests/test_counts.py with tests for empty strings and word counting:
from bookstats.counts import count_words, extract_words
def test_extract_words_empty_string():
assert extract_words("") == []
def test_count_words():
words = ["apple", "banana", "apple", "cherry", "apple", "banana"]
df = count_words(words)
assert df["word"].to_list() == ["apple", "banana", "cherry"]
assert df["count"].to_list() == [3, 2, 1]
def test_count_words_empty():
df = count_words([])
assert len(df) == 0
assert "word" in df.columns
assert "count" in df.columnsRun pytest to ensure all unit tests pass:
uv run pytestIntegration Testing: File Boundaries
Unit tests verify functions in isolation using synthetic Python lists. But research workflows read and write files on disk! An integration test checks that multiple components work together across system boundaries (such as filesystems and I/O).
Pytest provides a built-in fixture called tmp_path that gives each test a clean, isolated temporary directory that is automatically deleted after the test finishes.
Create tests/test_integration.py:
"""Integration test for end-to-end file processing boundary."""
from pathlib import Path
import polars as pl
from bookstats.counts import process_book_file
def test_process_book_file_creates_expected_csv(tmp_path: Path):
# 1. Arrange: Create a minimal mock Project Gutenberg book
sample_text = (
"*** START OF THE PROJECT GUTENBERG EBOOK TEST ***\n"
"The quick brown fox jumps over the lazy dog.\n"
"The fox was quick!\n"
"*** END OF THE PROJECT GUTENBERG EBOOK TEST ***\n"
)
input_file = tmp_path / "sample.txt"
input_file.write_text(sample_text, encoding="utf-8")
output_file = tmp_path / "counts.csv"
# 2. Act: Run the real file-processing boundary
process_book_file(input_file, output_file)
# 3. Assert: Verify the output CSV exists and contains expected counts
assert output_file.exists()
loaded_df = pl.read_csv(output_file)
assert "word" in loaded_df.columns
assert "count" in loaded_df.columns
counts_dict = dict(zip(loaded_df["word"], loaded_df["count"]))
assert counts_dict["the"] == 3
assert counts_dict["fox"] == 2
assert counts_dict["quick"] == 2Run the suite:
uv run pytest -v tests/test_integration.pyCommit the test:
git add tests/test_integration.py
git commit -m "test: add integration test for file processing boundary"Scientific Analysis: Descriptive Zipf Fit
Now we use TDD to extend our research project with scientific analysis: computing a descriptive log-log linear fit for Zipf’s law.
Zipf’s law states that the frequency of a word is inversely proportional to its rank:
\[\text{frequency} \propto \frac{1}{\text{rank}^\alpha} \implies \log(\text{frequency}) \approx \beta_0 - \alpha \log(\text{rank})\]
Performing ordinary least squares (OLS) regression on log-transformed data provides a useful descriptive summary of decay rate (slope \(\approx -1.0\)) and fit quality (\(R^2\)).
However, OLS on log-transformed binned power-law data can produce biased estimates and does not mathematically prove that empirical data follows a true power law distribution (which requires maximum likelihood estimation). In this course, we treat this regression strictly as a descriptive fit.
Adding the Zipf Tests First (TDD)
Create tests/test_zipf.py:
"""Unit tests for Zipf's law fitting functions."""
from pathlib import Path
import polars as pl
import pytest
from bookstats.zipf import compute_zipf_fit, fit_all_books
def test_compute_zipf_fit_linear_decay():
# Construct synthetic data with an exact 1/rank Zipf decay
ranks = list(range(1, 11))
counts = [int(1000 / r) for r in ranks]
df = pl.DataFrame({"word": [f"w{i}" for i in ranks], "count": counts})
fit = compute_zipf_fit(df)
# Slope should be approximately -1.0 with high R^2
assert pytest.approx(fit.slope, rel=0.1) == -1.0
assert fit.r_squared > 0.95
assert "rank" in fit.data.columns
assert "log_rank" in fit.data.columns
assert "log_count" in fit.data.columns
assert "fitted_log_count" in fit.data.columns
assert "fitted_count" in fit.data.columns
def test_compute_zipf_fit_empty():
df = pl.DataFrame(
{"word": [], "count": []}, schema={"word": pl.String, "count": pl.UInt32}
)
fit = compute_zipf_fit(df)
assert fit.slope == 0.0
assert fit.intercept == 0.0
assert fit.r_squared == 0.0
assert len(fit.data) == 0Add scipy as a dependency for regression:
uv add scipyImplementing src/bookstats/zipf.py
Create src/bookstats/zipf.py with dataclass outputs and regression computation:
"""Descriptive Zipf's law analysis and log-log linear fitting."""
from __future__ import annotations
from dataclasses import dataclass
from pathlib import Path
import numpy as np
import polars as pl
from scipy.stats import linregress
@dataclass(frozen=True)
class ZipfFitResult:
"""Results from a descriptive log-log linear fit.
Attributes
----------
slope : float
Slope of the descriptive log-log linear fit.
intercept : float
Intercept of the descriptive log-log linear fit.
r_squared : float
Coefficient of determination (R^2).
data : polars.DataFrame
DataFrame with rank, count, log_rank, log_count, fitted values.
"""
slope: float
intercept: float
r_squared: float
data: pl.DataFrame
def compute_zipf_fit(counts_df: pl.DataFrame) -> ZipfFitResult:
"""Compute ranks, log transforms, and descriptive log-log linear fit.
Parameters
----------
counts_df : polars.DataFrame
DataFrame containing at least 'count' (and optionally 'word').
Returns
-------
ZipfFitResult
Fit metrics and transformed DataFrame.
"""
if len(counts_df) == 0:
empty_df = pl.DataFrame(
schema={
"rank": pl.UInt32,
"count": pl.UInt32,
"log_rank": pl.Float64,
"log_count": pl.Float64,
"fitted_log_count": pl.Float64,
"fitted_count": pl.Float64,
}
)
return ZipfFitResult(slope=0.0, intercept=0.0, r_squared=0.0, data=empty_df)
sorted_df = counts_df.sort("count", descending=True)
ranks = np.arange(1, len(sorted_df) + 1, dtype=np.float64)
counts = sorted_df["count"].to_numpy().astype(np.float64)
log_ranks = np.log(ranks)
log_counts = np.log(counts)
if len(ranks) > 1:
res = linregress(log_ranks, log_counts)
slope = float(res.slope)
intercept = float(res.intercept)
r_squared = float(res.rvalue**2)
else:
slope = 0.0
intercept = float(log_counts[0]) if len(log_counts) > 0 else 0.0
r_squared = 1.0
fitted_log_counts = slope * log_ranks + intercept
fitted_counts = np.exp(fitted_log_counts)
result_df = sorted_df.with_columns(
[
pl.Series("rank", ranks.astype(np.uint32)),
pl.Series("log_rank", log_ranks),
pl.Series("log_count", log_counts),
pl.Series("fitted_log_count", fitted_log_counts),
pl.Series("fitted_count", fitted_counts),
]
)
return ZipfFitResult(
slope=slope, intercept=intercept, r_squared=r_squared, data=result_df
)Run pytest to verify that all tests pass:
uv run pytest -vCommit the implementation:
git add src/bookstats/zipf.py tests/test_zipf.py pyproject.toml uv.lock
git commit -m "feat: implement descriptive log-log linear Zipf fit with unit tests"Interactive Visualization: Overlaying the Fit in Marimo
Now open notebooks/visualize.py in Marimo:
uv run marimo edit notebooks/visualize.pyImport compute_zipf_fit from bookstats.zipf. When the reader selects a book (whether Frankenstein, Dracula, or another title from Project Gutenberg), Marimo will: 1. Filter the dataset to the selected book. 2. Run compute_zipf_fit(book_df). 3. Display fit statistics (Slope \(\alpha\), Intercept, and \(R^2\)). 4. Plot the empirical word counts as scatter points alongside the regression fit line.
%%{init: {"flowchart": {"curve": "basis"}, "theme": "base", "themeVariables": {"fontFamily": "Arial"}}}%%
flowchart LR
Dropdown["Dropdown Selection ('frankenstein')"] --> Filter["Filter Data"]
Filter --> Fit["compute_zipf_fit()"]
Fit --> Stats["Display: Slope = -1.04, R² = 0.98"]
Fit --> Chart["Altair Layered Chart: Points + Fit Line"]
classDef ui fill:#fef3c7,stroke:#b45309,color:#000
classDef calc fill:#dbeafe,stroke:#2563eb,color:#000
class Dropdown,Stats,Chart ui
class Filter,Fit calc
In Marimo, layered charts in Altair are combined with the + operator:
# Empirical observations (points)
points = alt.Chart(chart_data).mark_circle(opacity=0.5, color="#2563eb").encode(
x=alt.X("rank:Q", scale=alt.Scale(type="log"), title="Rank (log scale)"),
y=alt.Y("count:Q", scale=alt.Scale(type="log"), title="Frequency (log scale)"),
tooltip=["word", "rank", "count"]
)
# Descriptive linear fit (line)
line = alt.Chart(chart_data).mark_line(color="#dc2626", strokeWidth=2).encode(
x="rank:Q",
y="fitted_count:Q"
)
chart = points + lineCommit your updated notebook and merge the feature branch:
git add notebooks/visualize.py
git commit -m "feat: overlay descriptive Zipf regression line in marimo notebook"
git push -u origin add-testsOpen a Pull Request on GitHub, review the changes, and merge into main.
OIST Transfer Task
Take 10–15 minutes to reflect on your own research project at OIST:
- Identify one critical failure mode: What is one silent error that could corrupt your analysis (e.g., empty input files, unexpected negative values, NaNs in experimental measurements)?
- Draft an automated check: Write one unit test, input validation assertion, or scientific plausibility check for that failure.
- Commit the check: Commit the test or assertion to your OIST repository branch.
Checklist for Completion
Before proceeding to Session 4, confirm that your project meets these milestones: