OpenWiki: Repo Docs, Admin Browser, Nightly Cron, and a PaperTrail Feed
⚠️ Cookbook example — not live code. (KEEP THIS CALLOUT.) Every code block below is an example snippet, not part of the llamapress.ai codebase, and not running on this server. This is a reference recipe for a Leo instance (an AI coding agent) to implement in its own app — read it to understand the pattern, then recreate it there.
OpenWiki is a command-line agent that reads a repository and writes a wiki about it
into an openwiki/ folder. You run it once with --init to build the wiki, then on a
schedule with --update to keep it current. It writes plain Markdown files with YAML
frontmatter, so the wiki is just files in your repo — you can grep it, diff it, and
serve it.
This guide has four parts:
- Install OpenWiki and generate the wiki.
- Serve it inside your own app at
/admin/open-wiki(sidebar, Markdown rendering, working internal links). - Schedule it with a nightly cron that fails loudly instead of silently.
- Feed it more than code. OpenWiki only reads files. A wiki built from code alone documents what the software can do, never what it did. Part 4 exports a PaperTrail digest — which models real users changed, how often, and which columns — into a Markdown file in the repo, so the next OpenWiki run reads it and writes it into the wiki.
When to use: a codebase big enough that new engineers (or AI agents) can't hold it in their head, where you want documentation that refreshes itself. When not to: a small app where a hand-written README stays accurate. OpenWiki costs LLM tokens on every run.
The 80/20 in one breath
npm install --global openwiki, then authenticate a model provider.- Run
openwiki code --initin the repo root. It writesopenwiki/*.md. - Write
openwiki/INSTRUCTIONS.md— the brief that steers every future run. This is the file you edit to change the wiki, not the pages themselves. - Bind-mount
./openwikiread-only into the Rails container, add two routes, and add a read-only controller that parses frontmatter server-side and renders the body with marked.js. - Add a cron entry running a wrapper script (
flock,timeout, a log, and a loud failure marker) that callsopenwiki code --update --printnightly. - Optional but high value: a second, earlier cron writes a PaperTrail digest to
docs/audit/data-change-digest.md, andINSTRUCTIONS.mdtells OpenWiki to synthesize it.
Layer 1 — Install and generate
OpenWiki is a Node CLI. Install it globally on the machine that will run the cron.
# Node 22+. If you use nvm, note WHICH node — the cron section below depends on it.
npm install --global openwiki
cd ~/YourRepo
openwiki auth openai-chatgpt # or set OPENWIKI_PROVIDER + an API key
openwiki code --init --print # plans the wiki; see the warning below
⚠️
--initdoes not build the wiki. It plans one. On any repository big enough to be worth documenting, that single agent turn is spent writingopenwiki/_skeleton.md, anindex.md, and a message offering to continue — then it exits 0. Nothing tells you the "wiki" is a plan:--updatepasses, each one turn, until the page count stops climbing:for i in $(seq 1 8); do openwiki code --update --print "Continue building this wiki from openwiki/_skeleton.md. Write the next few pages COMPLETELY, in the skeleton's priority order. Do not reply with a plan and do not ask whether to proceed — write the pages now. Delete openwiki/_skeleton.md when every planned page exists with real content." # stop when the page count stops rising, or _skeleton.md disappears doneBudget for this: ~25 pages over three passes on a mid-size Rails app.
Two modes exist and they are easy to confuse:
| Command | What it does | Writes to |
|---|---|---|
openwiki code |
Documents the current repository | <repo>/openwiki/ |
openwiki personal |
A local personal brain over configured connectors | ~/.openwiki/wiki |
Always pass code explicitly in scripts. Useful flags:
openwiki code --update --print # one non-interactive run, prints the summary
openwiki code --update "focus on the jobs directory" # steer a single run with a message
openwiki code --modelId <model-id> --update
After a run, OpenWiki records what it did in openwiki/.last-update.json:
{
"updatedAt": "2026-08-07T08:03:11.568Z",
"command": "update",
"gitHead": "636f08d71b3992cee07df3f5b2e7279e44289092",
"model": "gpt-5.6-terra"
}
That gitHead is the mechanism of the whole system. An --update run diffs the repo
from that commit to HEAD and only rewrites the pages the changed files affect. It is
incremental, not a full rebuild. Two consequences follow, and both bite people:
- A run right after a big merge is slow and expensive. A run with no new commits is nearly free.
- Work that never lands in a commit is invisible to the update. Part 4 is built around this fact.
Layer 2 — Steer the wiki with INSTRUCTIONS.md
openwiki/INSTRUCTIONS.md is the standing brief. Every run reads it. Editing a generated
page is pointless — the next run overwrites it. Editing the brief changes the wiki
permanently.
A brief that works has five sections:
This is the internal wiki for <system>. Its readers are <who> and the AI agents that
operate this system.
## Scope — synthesize ALL of these sources
- `app/`, `lib/`, `bin/` — the code.
- `docs/dev/`, `docs/incidents/` — designs and postmortems.
- `docs/audit/` — the PaperTrail data-change digest (see Layer 5).
- `.claude/skills/*/SKILL.md` — operational playbooks with exact commands.
## HARD RULE — secrets
Never copy credentials, API keys, tokens, `.env` values, or SSH key material into wiki
pages. Naming the env var is fine; the value never is.
## Required coverage (build pages for these; keep them current)
1. <Topic> — <the source files it must synthesize>
2. …
## Audience convention (maintain on every run)
Every page's `tags` list carries exactly ONE audience tag: `audience-engineering`,
`audience-business`, or `audience-all`. Preserve the tag on every existing page.
## Frontmatter and recall
Agents find pages by grepping `description:` fields. Write descriptions
grep-optimized: include exact command names, class names, error strings, and domain
terms a searcher would type.
The "Required coverage" list is the highest-leverage part. Without it the agent writes whatever the diff suggested, and important-but-stable subsystems slowly go undocumented because nobody edits them.
Layer 3 — The admin browser UI
The wiki is Markdown on disk. Serving it needs three pieces: a mount, a read-only controller, and a client-side renderer.
3a. Mount the folder read-only
The Rails container cannot see the repository root. Mount just the wiki folder:
# docker-compose.yml
services:
llamapress:
volumes:
- ./openwiki:/rails/openwiki:ro # generated docs, browsed at /admin/open-wiki
This is a directory mount, so regenerated pages appear with no restart. (Single-file mounts do not behave this way — see Gotchas.)
3b. Routes
# config/routes.rb
namespace :admin do
# `format: false` keeps a `.md` suffix inside the wildcard instead of Rails
# parsing it as a response format.
get '/open-wiki', to: 'open_wiki#index', as: :open_wiki
get '/open-wiki/*path', to: 'open_wiki#show', as: :open_wiki_page, format: false
end
3c. The controller — read-only, path-guarded
# app/controllers/admin/open_wiki_controller.rb
#
# Read-only browser for the generated openwiki/ folder. This controller NEVER
# writes: the pages are generated, and hand-edits get clobbered on regeneration.
class Admin::OpenWikiController < ApplicationController
before_action :authenticate_user!
before_action :ensure_admin
WIKI_DIR = Rails.root.join("openwiki").freeze
# Each path segment maps straight to a filename: leading alphanumeric, then
# word chars / dashes / dots. Blocks dotfiles, `..`, and absolute paths.
SEGMENT_RE = /\A[A-Za-z0-9][A-Za-z0-9\-_.]*\z/
def index = render_page("index.md")
def show
rel = resolve_page(params[:path].to_s)
return not_found unless rel
render_page(rel)
end
private
def render_page(rel)
unless Dir.exist?(WIKI_DIR)
redirect_to admin_path, alert: "The openwiki folder is not mounted into the container." and return
end
path = WIKI_DIR.join(rel)
return not_found unless File.file?(path)
raw = File.read(path)
render plain: raw, content_type: "text/markdown" and return if params[:raw].present?
@current = rel
@meta = frontmatter(raw)
@body = strip_frontmatter(raw)
@updated = File.mtime(path)
@tree = page_tree
@last_update = wiki_last_update
render "admin/open_wiki/show"
end
# "" → index.md · "architecture" → architecture/index.md · "quickstart" → quickstart.md
def resolve_page(raw_path)
segments = raw_path.split("/").reject(&:blank?)
return "index.md" if segments.empty?
return nil unless segments.all? { |s| s.match?(SEGMENT_RE) }
rel = segments.join("/")
candidates = []
candidates << rel if rel.end_with?(".md")
candidates << "#{rel}/index.md"
candidates << "#{rel}.md"
found = candidates.find { |c| File.file?(WIKI_DIR.join(c)) }
return nil unless found
# Belt-and-braces: the resolved absolute path must stay inside WIKI_DIR.
abs = File.expand_path(WIKI_DIR.join(found))
abs.start_with?("#{File.expand_path(WIKI_DIR)}/") ? found : nil
end
# All pages grouped by directory ("" = root), for the sidebar. index.md sorts
# first within a group, then by title.
def page_tree
Dir.glob(WIKI_DIR.join("**/*.md")).map { |file|
rel = Pathname.new(file).relative_path_from(WIKI_DIR).to_s
dir = File.dirname(rel)
{ rel: rel,
dir: dir == "." ? "" : dir,
index: File.basename(rel) == "index.md",
title: frontmatter(File.read(file))["title"].presence ||
File.basename(rel, ".md").tr("-", " ").capitalize }
}.group_by { |p| p[:dir] }
.sort_by { |dir, _| dir }.to_h
.transform_values { |ps| ps.sort_by { |p| [p[:index] ? 0 : 1, p[:title].downcase] } }
end
# The generator's own metadata: when the wiki was last rebuilt, and by which model.
def wiki_last_update
file = WIKI_DIR.join(".last-update.json")
File.file?(file) ? JSON.parse(File.read(file)) : nil
rescue JSON::ParserError
nil
end
def not_found = redirect_to(admin_open_wiki_path, alert: "That wiki page doesn't exist.")
# A malformed page must not 500 the whole wiki.
def frontmatter(raw)
m = raw.match(/\A---\s*\n(.*?\n)---\s*\n/m)
return {} unless m
YAML.safe_load(m[1], permitted_classes: [], aliases: false) || {}
rescue Psych::Exception
{}
end
def strip_frontmatter(raw) = raw.sub(/\A---\s*\n.*?\n---\s*\n/m, "")
end
3d. The view — render Markdown in the browser, and fix the links
Two things make this view non-trivial. First, there is no server-side Markdown gem in
most pinned images, so rendering happens client-side with marked.js. Second, OpenWiki
writes standard relative Markdown links (quickstart.md, ../workflows/foo.md,
architecture/). Those 404 under an admin route unless you rewrite them.
<%# app/views/admin/open_wiki/show.html.erb (body of the page) %>
<%# Raw page body — embedded inertly (Rails-escaped), rendered client-side. %>
<script type="text/plain" id="openwiki-md-src"><%= @body %></script>
<script src="https://cdn.jsdelivr.net/npm/marked@12.0.0/marked.min.js"></script>
<script>
(function () {
// The blob above was HTML-escaped by <%%= %>; decode back to true markdown.
function decodeEntities(s) {
var ta = document.createElement('textarea'); ta.innerHTML = s; return ta.value;
}
var raw = decodeEntities(document.getElementById('openwiki-md-src').textContent || '');
var bodyEl = document.getElementById('openwiki-body');
// Relative links resolve against the CURRENT page's directory.
var current = <%= @current.to_json.html_safe %>; // e.g. "architecture/overview.md"
var baseDir = current.split('/').slice(0, -1).join('/');
function rewriteHref(href) {
// External, absolute, protocol-relative, and pure-anchor links pass through.
if (!href || /^([a-z][a-z0-9+.-]*:|\/\/|\/|#)/i.test(href)) return null;
var m = href.match(/^([^#?]*)([#?].*)?$/);
var p = m[1], suffix = m[2] || '';
if (p === '') return null;
var parts = (baseDir ? baseDir + '/' + p : p).split('/'), out = [];
for (var i = 0; i < parts.length; i++) {
var s = parts[i];
if (s === '' || s === '.') continue;
if (s === '..') { out.pop(); continue; }
out.push(s);
}
var joined = out.join('/');
if (/\.md$/i.test(joined)) joined = joined.slice(0, -3);
return '/admin/open-wiki/' + joined + suffix;
}
marked.setOptions({ gfm: true, breaks: false });
bodyEl.innerHTML = marked.parse(raw);
bodyEl.querySelectorAll('a[href]').forEach(function (a) {
var to = rewriteHref(a.getAttribute('href'));
if (to) a.setAttribute('href', to);
});
})();
</script>
Render the sidebar from @tree, and show @last_update['updatedAt'] and
@last_update['model'] in the page header. The freshness line matters more than it
looks: a wiki that quietly stopped updating three weeks ago reads exactly like a wiki
that is current.
Add a Raw .md link (?raw=1) on every page. Agents fetch that endpoint directly.
Layer 4 — The nightly cron
openwiki code --update --print is the whole job. Everything around it exists because
this job runs unattended and can fail in ways that look like success.
#!/usr/bin/env bash
# openwiki-nightly.sh — nightly OpenWiki update for the internal wiki.
#
# SCHEDULE: crontab, 0 8 * * * UTC.
# LOG: /home/ubuntu/openwiki-nightly.log (truncated at 1MB).
# STEERING: edit openwiki/INSTRUCTIONS.md — NOT this script.
#
# HARD-WON NOTES:
# - openwiki is installed under nvm Node 22, NOT system Node 18 — the PATH export
# below is load-bearing; cron does not source ~/.bashrc.
# - `openwiki --help` prints "provider: OpenAI" before loading its env file; that
# banner is not evidence of misconfiguration.
# - AUTH EXPIRES SILENTLY. An OAuth refresh token can be revoked (e.g. the account
# is re-signed-in elsewhere). The run then dies instantly. Cron still fires — the
# fingerprint is a MISSING "=== done" line, not a missing start line. That is why
# the trap below writes a greppable "*** FAILED" marker.
set -euo pipefail
export PATH="$HOME/.nvm/versions/node/v22.23.1/bin:$PATH"
LOG="$HOME/openwiki-nightly.log"
LOCK="/tmp/openwiki-nightly.lock"
REPO="$HOME/YourRepo"
# Truncate log if over 1MB
if [ -f "$LOG" ] && [ "$(stat -c%s "$LOG")" -gt 1048576 ]; then
tail -c 262144 "$LOG" > "$LOG.tmp" && mv "$LOG.tmp" "$LOG"
fi
{
echo "=== openwiki-nightly $(date -u '+%Y-%m-%d %H:%M:%S UTC') ==="
# flock: skip this run entirely if the previous one is still going.
flock -n 9 || { echo "SKIP: previous run still holds lock"; exit 0; }
cd "$REPO"
trap 'rc=$?; [ "$rc" -ne 0 ] && echo "*** FAILED $(date -u "+%F %T UTC"), exit $rc — check auth ***"' EXIT
timeout 3600 openwiki code --update --print
echo "=== done $(date -u '+%Y-%m-%d %H:%M:%S UTC'), exit $? ==="
} 9>"$LOCK" >> "$LOG" 2>&1
Install it:
chmod +x bin/local/openwiki-nightly.sh
crontab -e
# 0 8 * * * /home/ubuntu/YourRepo/bin/local/openwiki-nightly.sh
Check health in one command:
grep -c '=== done' ~/openwiki-nightly.log # successful runs
grep -n '\*\*\* FAILED' ~/openwiki-nightly.log # dead runs, with the exit code
Why nightly and not a git hook. A per-commit hook looks tempting because the update is diff-driven. Do not do it. Each run takes minutes and spends model quota, and the run produces wiki changes that themselves want a commit — so the hook recurses. One batch run per night collapses a day of commits into a single diff.
The GitHub Actions alternative
If you would rather run it in CI than on a box, the workflow is small:
# .github/workflows/openwiki-update.yml
name: OpenWiki Update
on:
workflow_dispatch:
schedule:
- cron: "0 8 * * *"
permissions:
contents: write
pull-requests: write
jobs:
update:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: "22" }
- run: npm install --global openwiki
- run: openwiki code --update --print
env:
OPENWIKI_PROVIDER: openrouter
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
OPENWIKI_MODEL_ID: <model-id>
- uses: peter-evans/create-pull-request@v7
with:
add-paths: |
openwiki
AGENTS.md
CLAUDE.md
.github/workflows/openwiki-update.yml
branch: openwiki/update
commit-message: "docs: update OpenWiki"
title: "docs: update OpenWiki"
Pick one path and mean it. The Actions path opens a pull request. If your team does not merge those pull requests, the wiki silently stops advancing while the workflow keeps reporting green — so the local cron becomes the real path and the workflow is noise. OpenWiki also regenerates this workflow file on every run, so never hand-edit it.
Layer 5 — Feed it more than code, with PaperTrail
Everything above documents the code. That leaves a real gap. The wiki can tell you
that a ProductionProjectItem has a welder column and which controller writes it. It
cannot tell you that the column is edited 200 times a week by two people, that imports
destroy and recreate every row, or that a table nobody mentions in standup carries the
most churn in the system.
PaperTrail already records that. It writes one row to a versions table for every
create, update, and destroy on a model you opt in. The plan is a small pipeline:
versions table → nightly rollup query → docs/audit/data-change-digest.md
→ git commit → OpenWiki --update reads it → a wiki page
The key insight is the arrow in the middle: OpenWiki reads files in a repository, so anything you want it to know must first become a file in the repository. That makes it extensible far past PaperTrail — the same pattern works for error telemetry, job runtimes, or support-ticket themes.
5a. Turn PaperTrail on
# Gemfile
gem "paper_trail", "~> 15.2"
# db/migrate/XXXXXXXX_create_versions.rb
class CreateVersions < ActiveRecord::Migration[7.2]
def change
create_table :versions do |t|
t.string :item_type, null: false
t.bigint :item_id, null: false
t.string :event, null: false # create | update | destroy
t.string :whodunnit
t.jsonb :object # the record BEFORE the change
t.jsonb :object_changes # {"col": [from, to]}
t.datetime :created_at
end
add_index :versions, %i[item_type item_id]
add_index :versions, :created_at
end
end
# config/initializers/paper_trail.rb
if defined?(PaperTrail)
PaperTrail.config.enabled = true
PaperTrail.config.track_associations = false
end
# app/controllers/application_controller.rb
# Without this, EVERY version has a null whodunnit and the digest cannot name actors.
before_action :set_paper_trail_whodunnit
# app/models/production_project_item.rb
class ProductionProjectItem < ApplicationRecord
has_paper_trail # bare = all columns tracked
# has_paper_trail only: [:welder, :weld_date] # or narrow it
end
Opt in deliberately. Bare has_paper_trail on a hot table can double its write
volume and grow versions past the table it audits. Start with the 5–10 models whose
history someone would actually ask about.
5b. The digest script — roll up, never dump
Put the script in lib/ (or rails/lib), not the app root — on Leo boxes the Rails root
is not bind-mounted, only its subdirectories are, so a script at the root will not exist
inside the container.
# lib/papertrail_digest.rb — run with: bin/rails runner /rails/lib/papertrail_digest.rb
#
# Prints a Markdown digest of the last 7 days of PaperTrail activity to stdout,
# between two markers so the caller can slice it out of Rails' log noise.
#
# HARD RULE: emit COUNTS and COLUMN NAMES only. Never the values in `object` or
# `object_changes` — those are your users' real data, and this file gets committed
# and read by a model provider.
ActiveRecord::Base.logger = nil
DAYS = 7
conn = ActiveRecord::Base.connection
def rows(conn, sql) = conn.select_all(sql).to_a
puts "--- BEGIN DIGEST ---"
puts <<~HEAD
---
type: Data Change Digest
title: "Data change digest (last #{DAYS} days)"
description: "PaperTrail rollup: which models real users create, update, and destroy, which columns churn, and who the actors are. Generated nightly; counts and column names only, never values."
tags: [audit, paper_trail, audience-engineering]
---
# Data change digest — last #{DAYS} days
Generated #{Time.current.utc.iso8601} from the `versions` table. Counts only.
HEAD
# 1) Volume by model and event.
puts "\n## Change volume by model\n\n| Model | Creates | Updates | Destroys |\n|---|---:|---:|---:|"
rows(conn, <<~SQL).each { |r| puts "| #{r['item_type']} | #{r['creates']} | #{r['updates']} | #{r['destroys']} |" }
SELECT item_type,
count(*) FILTER (WHERE event = 'create') AS creates,
count(*) FILTER (WHERE event = 'update') AS updates,
count(*) FILTER (WHERE event = 'destroy') AS destroys
FROM versions
WHERE created_at >= now() - interval '#{DAYS} days'
GROUP BY 1 ORDER BY count(*) DESC LIMIT 25
SQL
# 2) Which COLUMNS actually churn. jsonb_each over object_changes gives the keys;
# we take the key and throw the value away.
puts "\n## Hottest columns (updates only)\n\n| Model | Column | Edits |\n|---|---|---:|"
rows(conn, <<~SQL).each { |r| puts "| #{r['item_type']} | `#{r['column_name']}` | #{r['edits']} |" }
SELECT v.item_type, c.key AS column_name, count(*) AS edits
FROM versions v, jsonb_each(v.object_changes) c
WHERE v.event = 'update'
AND v.created_at >= now() - interval '#{DAYS} days'
AND c.key NOT IN ('updated_at', 'created_at')
GROUP BY 1, 2 ORDER BY edits DESC LIMIT 30
SQL
# 3) Who. whodunnit is a users.id AS A STRING — PaperTrail stores no name. Join it.
puts "\n## Actors\n\n| User | Changes |\n|---|---:|"
rows(conn, <<~SQL).each { |r| puts "| #{r['actor']} | #{r['n']} |" }
SELECT coalesce(u.email, 'unattributed (' || coalesce(v.whodunnit, 'null') || ')') AS actor,
count(*) AS n
FROM versions v
LEFT JOIN users u ON u.id = nullif(v.whodunnit, '')::bigint
WHERE v.created_at >= now() - interval '#{DAYS} days'
GROUP BY 1 ORDER BY n DESC LIMIT 15
SQL
# 4) Coverage — the honest denominator. A model with no has_paper_trail is
# UNMEASURED, not quiet, and the wiki must say so.
Rails.application.eager_load!
tracked = ApplicationRecord.descendants.select { |m| m.respond_to?(:paper_trail_options) }.map(&:name).sort
untracked = ApplicationRecord.descendants.map(&:name).sort - tracked
puts "\n## Coverage\n"
puts "Versioned (#{tracked.size}): #{tracked.join(', ')}"
puts "\nNOT versioned (#{untracked.size}) — absence of history here means UNMEASURED, not unchanged:"
puts untracked.join(', ')
puts "--- END DIGEST ---"
5c. The wrapper cron
#!/usr/bin/env bash
# papertrail-digest-nightly.sh — write the data-change digest, then commit it so the
# OpenWiki run (which is git-diff driven) actually sees it.
#
# RUNS BEFORE openwiki-nightly.sh. 15 minutes of headroom is plenty.
set -euo pipefail
REPO="$HOME/YourRepo"
OUT="$REPO/docs/audit/data-change-digest.md"
cd "$REPO"
mkdir -p "$(dirname "$OUT")"
# `rails runner` interleaves boot output with your stdout. Slice on the markers
# instead of trusting a clean stdout.
docker compose exec -T llamapress bin/rails runner /rails/lib/papertrail_digest.rb \
| sed -n '/--- BEGIN DIGEST ---/,/--- END DIGEST ---/p' \
| sed '1d;$d' > "$OUT.tmp"
# Never publish an empty digest over a good one.
if [ "$(wc -l < "$OUT.tmp")" -lt 10 ]; then
echo "*** FAILED: digest too short ($(wc -l < "$OUT.tmp") lines), keeping previous"; exit 1
fi
mv "$OUT.tmp" "$OUT"
# The OpenWiki update diffs from the gitHead in .last-update.json. An uncommitted
# file is not in that diff, so the run will not notice the digest changed.
if ! git diff --quiet -- "$OUT"; then
git add "$OUT"
git commit -m "chore: nightly data-change digest"
fi
45 7 * * * /home/ubuntu/YourRepo/bin/local/papertrail-digest-nightly.sh >> /home/ubuntu/papertrail-digest.log 2>&1
0 8 * * * /home/ubuntu/YourRepo/bin/local/openwiki-nightly.sh
5d. Tell OpenWiki the file exists
Two edits to openwiki/INSTRUCTIONS.md. Without them, OpenWiki may treat the digest as
just another data file and skip it.
## Scope — synthesize ALL of these sources
- `docs/audit/data-change-digest.md` — a nightly PaperTrail rollup of what real users
changed. Treat it as evidence of ACTUAL system usage, distinct from what the code
makes possible.
## Required coverage
11. **Data change patterns** — from `docs/audit/data-change-digest.md`: which models
carry real write volume, which columns churn, who the actors are, and which models
are NOT versioned (absence of history there means unmeasured, not unchanged). Cross-
reference the hot models against the code pages so a reader can jump from "this
table changes constantly" to the controller that writes it.
Now the wiki says things a code-only wiki never could: "tender_line_item is the
highest-churn model in the system; 94% of edits touch four columns; Welder is not
versioned, so employee-record edits are unknowable." That is the sentence that changes
what an engineer does next.
Gotchas (the hard-won stuff)
OpenWiki itself
--initplans; it does not build. See the warning in Layer 1. The tell is a_skeleton.mdin the output folder and a page count in the low single digits.- The first pass writes GENERIC pages, and generic is worse than missing. Asked to synthesize across dozens of documents in one turn, the agent reaches for plausible themes instead of reading the evidence: a cross-client insight page came back saying clients want "clear reporting that ties work to outcomes" — true of every agency on earth, and citing nothing. A page that could have been written without the data looks like insight and is not. Fix it with a second, narrower pass per deliverable that demands citations: "For every theme you claim, cite at least two specific conversations by name and date and quote the source. Drop any theme you cannot cite." The rewrite came back with six themes, each carrying three or four dated citations and direct quotes.
- It reads the FILESYSTEM, not git — so
.gitignoredoes not hide.envfrom it. Add an.openwikiignore(same syntax) covering.env,*.pem,*.key, database dumps,backups/andlogs/, or the agent reads your live credentials and may paraphrase them into a page. This is the single most important file in the setup. - The update is git-diff driven from
.last-update.json'sgitHead. Uncommitted work is largely invisible. If you generate an input file for the wiki, commit it (Layer 5c does) — and commit the wiki too, or the first checkpoint rollback on an agent-run box deletes it. - Its model registry goes stale. v0.3.1 warns that a current model "is not a known Anthropic model (it belongs to GitHub Copilot)". The warning is cosmetic — the call is still made — so do not chase it when the real error is underneath.
- Never hand-edit
openwiki/index.mdor the GitHub Actions workflow file. OpenWiki deterministically overwrites both on every run. Same for any generated page: your edit survives until the next run and then vanishes, which is worse than never making it. openwiki --helpprints the provider banner before it loads its env file. Seeing "provider: OpenAI" when you configured something else is not a misconfiguration.- OAuth auth expires silently and cron keeps firing. A revoked refresh token kills
the run in seconds. The log then shows a start line and nothing else — identical to
"still running". Three nights were lost to exactly this. The fix is structural: the
trapwriting*** FAILED, and monitoring for a missing=== donerather than a missing start line. - Under nvm, cron cannot find the binary. Cron does not source
~/.bashrc, soopenwikiinstalled under Node 22 is not on cron'sPATH. The explicitexport PATHis load-bearing, and it hardcodes a Node version — a Node upgrade breaks the cron silently. Re-check it after any nvm change. - Use
flockandtimeout. A slow run overlapping the next night's run produces two agents writing the same files.timeout 3600caps a hung run. - The secrets rule belongs in
INSTRUCTIONS.md, not in your head. The agent reads your whole repo and writes summaries. Say explicitly that env var names are fine and values never are. - The wiki folder is generated output, but it is not free to publish. Decide whether
openwiki/ships to downstream forks or customer boxes. If it must stay internal, keep it off your deploy allowlist and setvisibility: adminon the browser.
The admin browser
format: falseon the wildcard route. Without it, a request for/admin/open-wiki/architecture/overview.mdmakes Rails parse.mdas a response format and the route misses.- Validate every path segment against a regex and re-check the expanded path. A
wildcard route that reads files is a directory-traversal hole by default. The
SEGMENT_REcheck plus theabs.start_with?check are both needed — the first blocks..in the request, the second catches anything that slips past. - A malformed page must not 500 the wiki. Frontmatter is model-generated, so it will
occasionally be invalid YAML. Rescue
Psych::Exceptionand return{}. - Use
YAML.safe_loadwithpermitted_classes: []. PlainYAML.loadon a generated file is remote code execution waiting for a bad run. - A directory bind-mount hot-reloads; a single-file mount does not.
./openwikias a directory means new pages appear with no restart. If you instead mount an individual file, an atomic-write editor swaps the host inode and detaches it from the mount — the host file changes and the container keeps reading the old one, silently. - Rewrite relative links or the wiki is unnavigable. OpenWiki writes portable
Markdown links. Under
/admin/open-wiki/...every one of them 404s until you resolve it against the current page's directory and strip the.md. - Show the last-updated date and model in the header. A stale wiki looks exactly like a fresh one. This is the cheapest possible staleness alarm.
- Escape the body, then decode it in JS. Interpolating raw Markdown into a
<script>tag unescaped lets a generated page break out of it.
PaperTrail
- Confirm the gem is actually in your image before designing around it. On this
system, the same initializer and migration ship everywhere, but the mothership image
does not bundle
paper_trail— sodefined?(PaperTrail)isfalse, the initializer no-ops, andversionssits at 0 rows while looking perfectly configured. The fleet image does bundle it (paper_trail ~> 15.2). A table existing is not proof the gem is loaded. Checkdefined?(PaperTrail), not the schema. - The gem being present is still not tracking. Nothing is versioned until a model
declares
has_paper_trail. A fresh box has the gem, the table, and zero coverage. - Bulk inserts skip callbacks, so they produce NO version.
insert_all,upsert_all, and most Excel/CSV importers write rows PaperTrail never sees. One real project had 1,549 versions, every one anupdateand not a singlecreate— the values arrived with the insert. Never read a missing create version as "nobody set it." Say this in the digest, or the wiki will confidently state the opposite. - Re-imports that destroy and recreate rows break history continuity. History follows
the row id, not the thing the row represents. On that same project,
destroywas the single largest event type (34,685 of 46,396). A record's story ends at each re-import. whodunnitis a user id stored as a string — no name, no email. Join it yourself, and expect nulls from anything that runs outside a request (jobs, console, rake).- There is an install-date horizon. Nothing before the day you added
has_paper_trailis knowable. Put that date in the digest so nobody mistakes the horizon for a quiet period. - Never let the digest emit
objectorobject_changesvalues. Those columns hold your users' real data verbatim — names, addresses, amounts. Column names and counts answer every question the wiki needs and leak nothing. This matters twice over because the digest gets committed and fed to a model provider. - Slice
rails runneroutput on markers. Boot logs and warnings interleave with your stdout; redirecting it straight into a Markdown file gets you a Markdown file with a Docker warning at the top. - Guard against writing an empty digest. A failed query producing a 3-line file that overwrites a good one turns a monitoring system into a source of false calm.
Files this pattern touches
openwiki/INSTRUCTIONS.md # the standing brief — edit THIS
openwiki/**/*.md # generated; never hand-edit
openwiki/.last-update.json # gitHead + model of the last run
docker-compose.yml # ./openwiki:/rails/openwiki:ro
config/routes.rb # 2 routes, format: false
app/controllers/admin/open_wiki_controller.rb # read-only browser
app/views/admin/open_wiki/show.html.erb # sidebar + marked.js + link rewriting
bin/local/openwiki-nightly.sh # cron wrapper: flock, timeout, markers
bin/local/papertrail-digest-nightly.sh # cron wrapper: digest + commit
lib/papertrail_digest.rb # the rollup queries
docs/audit/data-change-digest.md # generated input to the wiki
Gemfile # gem "paper_trail"
config/initializers/paper_trail.rb
db/migrate/XXXXXXXX_create_versions.rb
app/controllers/application_controller.rb # set_paper_trail_whodunnit
app/models/*.rb # has_paper_trail on chosen models
.github/workflows/openwiki-update.yml # optional; regenerated by OpenWiki
How to adapt to your stack
- Swap the model provider. Set
OPENWIKI_PROVIDERplus that provider's key. A subscription-billed provider (ChatGPT/Codex-backed) avoids per-token metering on a nightly job; an API key is simpler to automate. Pick before you schedule it — the auth-expiry failure mode in Gotchas is specific to OAuth providers. - Not on Docker? Drop the
docker compose execfrom the digest wrapper and runbin/rails runner lib/papertrail_digest.rbdirectly. The controller does not change; only the mount goes away. - Not Postgres? The rollup queries use
FILTER (WHERE …)andjsonb_each. On MySQL, useSUM(event = 'create')andJSON_KEYS(object_changes). On SQLite, storeobject_changesas text and roll up in Ruby instead. - Public wiki instead of admin-only? Drop
ensure_admin, move the routes out of theadminnamespace, and re-read the secrets rule inINSTRUCTIONS.mdfirst — an agent summarizing your repo into a public page is a disclosure path. - Feed it something other than PaperTrail. The pattern is: roll up a data source
into a committed Markdown file, then name that file in
INSTRUCTIONS.md. Good candidates are exception counts by class, slowest jobs by runtime, and support-ticket themes. Keep each digest to one file with stable headings so the diff stays readable and the run stays cheap. - Safe to drop: the GitHub Actions workflow (pick cron or CI, not both), the
?raw=1endpoint if no agents read the wiki over HTTP, and the coverage section of the digest once every model you care about is versioned.