%%{init: {"flowchart": {"curve": "basis"}, "theme": "base", "themeVariables": {"fontFamily": "Arial"}}}%%
flowchart LR
Raw["data/raw/<br>(Raw Books)"] --> Inter["data/intermediate/<br>(Per-Book Counts)"]
Inter --> Proc["data/processed/<br>(Combined Counts)"]
Proc --> Fits["output/tables/<br>(Zipf Fits CSV)"]
Proc --> Figs["output/figures/<br>(Static & HTML Plots)"]
Proc --> Site["_site/<br>(Published Marimo App)"]
classDef data fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000
classDef out fill:#dcfce7,stroke:#16a34a,stroke-width:2px,color:#000
class Raw,Inter,Proc data
class Fits,Figs,Site out
Automation and Publication
“A computer lets you make more mistakes faster than any invention in human history—with the possible exceptions of handguns and tequila.”
— Mitch Ratcliffe
In research, an analysis is rarely run just once. Datasets are updated, bug fixes are made to preprocessing scripts, colleagues request new figures, and reviewers ask for revisions months after submission.
If reproducing your figures requires remembering a 12-step sequence of manual terminal commands, subtle discrepancies will inevitably creep in. In this final session, we transform our research project into a fully reproducible, self-documenting pipeline:
- Represent the project as a Directed Acyclic Graph (DAG) of files and dependencies.
- Encode the DAG in GNU Make to automate incremental rebuilds without redundant computation.
- Achieve one-command clean reproduction (
make clean && make check && make all). - Publish an interactive Marimo analysis to the web using GitHub Actions and GitHub Pages.
The Analysis Pipeline as a Directed Acyclic Graph (DAG)
In the previous sessions, we developed code to download books, count words, fit Zipf’s law, and visualize results. Let’s look at the flow of data through our project.
Every research workflow can be expressed as a Directed Acyclic Graph (DAG) where: - Nodes are files (data inputs, intermediate tables, code files, and final figures). - Edges (arrows) represent transformations (which files are needed to produce which outputs). - Acyclic means the workflow flows strictly forward; an output cannot be its own input.
Conceptual Workflow View
At a high level, raw texts flow through counting, aggregation, and curve fitting to produce final figures and an interactive report:
Data and Result Directory Semantics
Notice how our directories reflect clear data life cycles:
| Directory | Purpose | Version Controlled? | Example File |
|---|---|---|---|
data/raw/ |
Original immutable inputs (Gutenberg text files). | Yes | data/raw/frankenstein.txt |
data/intermediate/ |
Generated per-book count tables. | No (Git ignored) | data/intermediate/frankenstein.csv |
data/processed/ |
Combined, analysis-ready multi-book dataset. | No (Git ignored) | data/processed/book-counts.csv |
output/tables/ |
Reader-facing tabular summaries. | No (Git ignored) | output/tables/zipf-fits.csv |
output/figures/ |
Reader-facing plots (SVG and standalone HTML). | No (Git ignored) | output/figures/zipf-law.svg |
_site/ |
Complete published interactive Marimo web application. | No (Git ignored) | _site/index.html |
A table belongs under data/ when another analysis script consumes it. A table belongs under output/ when its primary audience is a human reader or reviewer. File format alone (.csv) does not determine whether something is data or output.
File-Level DAG for Dracula and Frankenstein
Here is the exact file-level DAG that connects our Python modules to our input and output files:
%%{init: {"flowchart": {"curve": "basis"}, "theme": "base", "themeVariables": {"fontFamily": "Arial"}}}%%
flowchart LR
D_txt["data/raw/dracula.txt"]:::raw --> D_csv["data/intermediate/dracula.csv"]:::inter
F_txt["data/raw/frankenstein.txt"]:::raw --> F_csv["data/intermediate/frankenstein.csv"]:::inter
Code_counts["src/bookstats/counts.py"]:::code -.-> D_csv
Code_counts -.-> F_csv
D_csv --> Combined["data/processed/book-counts.csv"]:::proc
F_csv --> Combined
Code_counts -.-> Combined
Combined --> Zipf_csv["output/tables/zipf-fits.csv"]:::out
Combined --> Fig_svg["output/figures/zipf-law.svg"]:::out
Combined --> Fig_html["output/figures/zipf-law.html"]:::out
Code_zipf["src/bookstats/zipf.py"]:::code -.-> Zipf_csv
Code_zipf -.-> Fig_svg
Code_zipf -.-> Fig_html
Combined --> Site["_site/index.html"]:::out
Code_viz["notebooks/visualize.py"]:::code -.-> Site
classDef raw fill:#e0e7ff,stroke:#4338ca,stroke-width:1.5px,color:#000
classDef inter fill:#f1f5f9,stroke:#64748b,stroke-width:1.5px,color:#000
classDef proc fill:#dbeafe,stroke:#2563eb,stroke-width:1.5px,color:#000
classDef out fill:#dcfce7,stroke:#16a34a,stroke-width:1.5px,color:#000
classDef code fill:#fef3c7,stroke:#b45309,stroke-dasharray: 5 5,color:#000
Notice that Python scripts (src/bookstats/counts.py, src/bookstats/zipf.py) are prerequisites just like data files! If you alter the word-tokenizing algorithm in counts.py, all downstream counts and figures must be updated.
Encoding the DAG with GNU Make
GNU Make is a build automation tool that has been a standard in computing for decades. Make uses a configuration file called a Makefile to define targets, their prerequisites, and the commands needed to build them.
Check that Make is installed on your machine:
make --versionThe Anatomy of a Make Rule
Every rule in a Makefile has three components:
target: prerequisite1 prerequisite2
recipe_command- Target: The file you want Make to build (or an action name).
- Prerequisites: The files required to produce the target.
- Recipe: The shell command that creates the target.
Every recipe line must start with a literal TAB character, not spaces. In VS Code, you can press Cmd/Ctrl+Shift+P, type Convert Indentation to Tabs, and select it.
Let’s inspect how Make decides whether to execute a recipe: 1. Does the target exist? If not, run the recipe. 2. Are any prerequisites newer than the target (by checking file modification timestamps)? If yes, run the recipe. 3. If the target exists and is newer than all prerequisites, Make skips the recipe because the target is up to date!
Building the Project Makefile
Create a feature branch for automation:
git switch -c add-automationCreate a file named Makefile in the root of your bookstats repository.
Public Targets and Conventions
We establish clear public targets for our project:
.POSIX:
.PHONY: all check clean help site
# Default target: builds all analysis tables and figures
all: data/processed/book-counts.csv output/tables/zipf-fits.csv output/figures/zipf-law.svg output/figures/zipf-law.html.PHONY mean?
A .PHONY declaration tells Make that a target is an action (like clean or check), rather than a physical file on disk. This prevents collisions if a file with that name ever exists.
Code Quality Check Target
Add a target to run all our linting, formatting, and test checks in one step:
# Run linter and automated test suite
check:
uv run ruff check .
uv run ruff format --check .
uv run pytest -vTry running it:
make checkPattern Rules for Intermediate Per-Book Counts
Instead of writing a separate rule for every single book in data/raw/, Make allows us to define a pattern rule using the % wildcard (stem matching):
RAW_BOOKS = $(wildcard data/raw/*.txt)
INTERMEDIATE_COUNTS = $(patsubst data/raw/%.txt,data/intermediate/%.csv,$(RAW_BOOKS))
# Pattern rule: Raw book -> Intermediate per-book word count CSV
data/intermediate/%.csv: data/raw/%.txt src/bookstats/counts.py
@mkdir -p data/intermediate
uv run python -m bookstats.counts $< $@Here: - $< is an automatic variable representing the first prerequisite (e.g. data/raw/dracula.txt). - $@ is an automatic variable representing the target (e.g. data/intermediate/dracula.csv). - $(wildcard ...) and $(patsubst ...) dynamically discover any books added by the user—whether Frankenstein, Dracula, or any other novel from Project Gutenberg!
Aggregate Counts and Result Outputs
Add the rules to combine counts and produce the reader-facing Zipf fit table and figures:
PROCESSED_DATA = data/processed/book-counts.csv
FIT_TABLE = output/tables/zipf-fits.csv
FIGURE_SVG = output/figures/zipf-law.svg
FIGURE_HTML = output/figures/zipf-law.html
# Combined processed dataset
$(PROCESSED_DATA): $(INTERMEDIATE_COUNTS) src/bookstats/counts.py
@mkdir -p data/processed
uv run python -m bookstats.counts --combine $(INTERMEDIATE_COUNTS) -o $@
# Reader-facing Zipf fit summary table
$(FIT_TABLE): $(PROCESSED_DATA) src/bookstats/zipf.py
@mkdir -p output/tables
uv run python -m bookstats.zipf $< --table $@
# Reader-facing figures: SVG and standalone interactive HTML
$(FIGURE_SVG) $(FIGURE_HTML): $(PROCESSED_DATA) src/bookstats/zipf.py
@mkdir -p output/figures
uv run python -m bookstats.zipf $< --svg $(FIGURE_SVG) --html $(FIGURE_HTML)Publishing the Interactive Marimo Website
Add the target that exports our interactive Marimo notebook into a standalone HTML publication artifact:
SITE_INDEX = _site/index.html
# Interactive Marimo application publication artifact
$(SITE_INDEX): $(PROCESSED_DATA) notebooks/visualize.py
@mkdir -p _site
uv run marimo export html notebooks/visualize.py -o $@
site: $(SITE_INDEX)Clean and Help Targets
Finally, add targets to clean generated artifacts and display help:
# Clean all generated, non-committed data, output artifacts, and website build
clean:
rm -rf data/intermediate data/processed output _site .pytest_cache .ruff_cache
# Display available Make targets
help:
@echo "Available make targets:"
@echo " all - Build processed data, fit tables, and figures (default)"
@echo " check - Run Ruff linting/formatting checks and pytest test suite"
@echo " clean - Remove all generated data, outputs, and build artifacts"
@echo " site - Export interactive Marimo website to _site/index.html"
@echo " help - Show this help message"Observing Incremental Rebuilds
The magic of build automation is incremental computation.
Run
make all. Make runs the counting scripts, combines tables, and produces figures.Run
make alla second time immediately. Make prints:make: Nothing to be done for `all'.Because no files were modified, Make executes zero commands.
Now touch a single book:
touch data/raw/dracula.txt make allWatch Make’s output carefully:
- Make rebuilds
data/intermediate/dracula.csv. - Make does not rebuild
data/intermediate/frankenstein.csv! - Because intermediate Dracula counts changed, Make rebuilds
data/processed/book-counts.csv,output/tables/zipf-fits.csv, and the figures.
- Make rebuilds
Now modify
notebooks/visualize.py:touch notebooks/visualize.py make siteMake rebuilds
_site/index.html, but does not re-parse books or re-fit tables.
The Ultimate Test: Clean One-Command Reproduction
To prove that your project is fully reproducible and has no hidden manual dependencies, run:
make clean
make check
make all
make siteIn under 10 seconds, Make cleans all intermediate state, validates code quality and tests, regenerates the full analysis from raw texts, and compiles the interactive website.
Commit the Makefile:
git add Makefile
git commit -m "feat: encode full analysis DAG in Makefile"Continuous Integration and Publication: GitHub Actions & Pages
Now that our analysis pipeline is automated locally, we can deploy a robot in the cloud to run this exact pipeline whenever changes are pushed to GitHub.
Create a GitHub Actions workflow file at .github/workflows/publish.yml:
name: Check and Publish
on:
push:
branches:
- main
workflow_dispatch:
permissions:
contents: read
pages: write
id-token: write
concurrency:
group: pages
cancel-in-progress: false
jobs:
build-and-deploy:
environment:
name: github-pages
url: ${{ steps.deployment.outputs.page_url }}
runs-on: ubuntu-latest
steps:
- name: Checkout source code
uses: actions/checkout@v4
- name: Install uv
uses: astral-sh/setup-uv@v5
with:
enable-cache: true
version: "latest"
- name: Set up Python 3.12
run: uv python install 3.12
- name: Install dependencies
run: uv sync --frozen
- name: Run code quality checks and tests
run: make check
- name: Run analysis pipeline
run: make all
- name: Build publication website
run: make site
- name: Setup Pages
uses: actions/configure-pages@v5
- name: Upload Pages artifact
uses: actions/upload-pages-artifact@v3
with:
path: _site
- name: Deploy to GitHub Pages
id: deployment
uses: actions/deploy-pages@v4Enabling GitHub Pages
Commit and push the workflow:
git add .github/workflows/publish.yml git commit -m "ci: automate testing, DAG pipeline, and GitHub Pages deployment" git push -u origin add-automationMerge the Pull Request on GitHub into
main.In your repository on GitHub, navigate to Settings → Pages.
Under Build and deployment → Source, select GitHub Actions.
%%{init: {"theme": "base", "themeVariables": {"fontFamily": "Arial"}}}%%
flowchart LR
Push["git push origin main"] --> Action["GitHub Actions Runner"]
Action --> Sync["uv sync --frozen"]
Sync --> Check["make check (Ruff & Pytest)"]
Check --> Build["make all & make site"]
Build --> Pages["Deploy to GitHub Pages URL"]
classDef git fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000
classDef pub fill:#dcfce7,stroke:#16a34a,stroke-width:2px,color:#000
class Push,Action,Sync,Check,Build git
class Pages pub
Within 1–2 minutes, the GitHub Action will finish running make check, make all, and make site, and will publish your interactive Marimo analysis to https://<your-username>.github.io/bookstats/!
OIST Transfer Task & Reflection
Reflect on your scientific computing workflows at OIST:
- Map your DAG: On paper or in a markdown file, sketch the Directed Acyclic Graph of one of your active research projects. What are the raw instruments/sequencer files? What intermediate files are generated? What figures end up in manuscripts?
- Identify Bottlenecks: Are there steps you currently run by copying and pasting terminal commands? Could a single
Makefilerule eliminate the risk of forgetting a step? - Continuous Verification: How could automated tests and linting in GitHub Actions protect your group’s shared code repository from accidental regressions?
Congratulations!
Across four sessions, you have progressed from an empty folder to a published, reproducible scientific research project: - Session 1: Managed version control, branches, issues, and peer code reviews on GitHub. - Session 2: Built an isolated, reproducible environment with uv, modularized scripts into an installable Python package, and created an interactive Marimo exploratory app. - Session 3: Enforced formatting with Ruff, guarded commits with pre-commit hooks, applied Test-Driven Development with pytest, and derived descriptive Zipf fits. - Session 4: Encoded the end-to-end dependency DAG in GNU Make, verified one-command reproduction, and published an interactive analysis to GitHub Pages with GitHub Actions.
These principles form the bedrock of transparent, reproducible, and good-enough scientific research.