Things I build.

Tools for working with coding agents, and a few older experiments. All outside my day job, all on GitHub.

kstrl

A software factory for AI coding agents. Hand it a spec and it plans, builds, and measures the result with checks the agent did not write: tests, types, lint and diff scope first, then a reviewer from a different model family on every acceptance criterion, then a security reviewer. Where the measurement disagrees with the agent, the gap goes back as the next instruction, parsed to file and line with a fix hint. Bounds on iterations, time, tokens, cost and what a merge may touch stop it running away. You stay on the loop, not in it: undecomposable specs, approvals and exhausted budgets come to you, everything else flows and is recorded. On PyPI as kstrl.

Python · CLI and TUI · MIT · active
[ fig.01 ] kstrl
AGENT12345671 IMPLEMENT · seconds to minutes · openthe agent; nothing that reads the code measures it2 ACCEPT · minutes · closeda retry with context; verify, review, security3 INTEGRATE · tens of minutes · closedschedule and merge; contract tests4 INTAKE · hours · closedqueue admission; queue state, spend, inbox5 TRUST · days · closed, two inputs unwiredan autonomy level change; run outcomes6 LEARN · weeks · open, sensor builtplaybook and prompt edits; attribution, calibration7 OPERATE · not builtthe release driver; runtime error ratesINNERMOST FASTEST. EACH BAND MEASURES THE ONE INSIDE IT WITH SOMETHING THE AGENT DID NOT WRITE.
seven loops around the agent, innermost fastestFIG

telltale

A flight recorder for coding agents. It reads the traces Claude Code and Codex already emit and keeps every tool call, file touched, command, test result, token spent, compaction and commit, and never the prompt, the reply or the contents of an edit. Then it asks what that record can forecast, and tests every answer against four rules of thumb with the bar written down before the run. What an agent is about to spend is partly knowable; what happens to the code afterwards is not, so far, beyond a one-line rule of thumb. The open question needs both halves of a change, and only a running recorder can supply the second.

Python · daemon · 3,300 sessions read, and counting
[ fig.02 ] telltale
WHILE THE AGENT WORKStool calls · tests run · tokens · compactions, recorded as they happenWHAT GIT KEEPSthe commitTHE SESSION HALF EXISTS ONLY IF SOMETHING WAS RECORDING
the session half is only ever recorded liveFIG

systemap

Your coding agent writes faster than you read, and the picture you had of the system stops matching it. systemap has the agent draw the map from the code, following a written procedure, and a checker refuses a map that leaves a module unplaced, routes a journey through a part it does not touch, or is older than the tree. Every pull request then says what it did to the shape of the system, not just which lines changed. Python only. Ships as a Claude Code plugin and a plain skill.

Python · skill and checker · maps itself
[ fig.03 ] systemap
THE MAP, DRAWN BY THE AGENT, HELD TO ELEVEN RULESGATEWAYEMITTERCONTRACTSno cardsolid edge: an import backs it · dashed: nothing doessystemap delta --base mainmoved: pkg.old -> pkg.newGateway still names pkg.old: rename itadded: pkg.thing, claimed by no cardgive it a card, or say why notevidence lost: Emitter -> Contractsnothing in the code backs that edge nowEXIT 1: SOMETHING NEEDS A DECISIONposted on the pull request, before the diffgit says which lines changedSYSTEMAP SAYS WHAT CHANGED IN THE SYSTEM
what the pull request did to the systemFIG

detailed-code-review

A review skill for Codex and Claude Code whose output is a merge decision, not a long comment. It fixes the comparison range first, maps the change before judging lines, and accepts a finding only with a trigger, an observable result and evidence in the repository. Review lanes run as independent specialists, coverage is recorded, and the verdict is one of five. An evaluator measures recall, precision and noise on real pull requests; the README says exactly how far each measurement falls short.

Skill · evaluator · paired evaluation on real pull requests
[ fig.04 ] detailed-code-review
A FINDING, BEFORE IT COUNTSTRIGGEROBSERVABLE RESULTEVIDENCE IN THE REPOONE MERGE VERDICT, ONE OF FIVE, FROM BLOCK TO APPROVEANY SLOT STILL EMPTY: THE FINDING IS NOISE, HOWEVER LONGUSEFUL DEFECTS FOUND, NOISE REMOVED, UNCERTAINTY VISIBLE
no failure path, no findingFIG

explain-pr

Teach a pull request instead of reviewing it. The skill hands the diff to a reader who did not make the change, because an author cannot see the reasoning that never reached the diff, and builds one arc: where this sits, why each decision went that way, what you need for the next change, and what would make it wrong. It ships with Human-Outward, a writing style for anything meant to be understood: the idea before its name, the answer you rejected, and a mark on how each claim is known.

Skill · output style · four checkers, limits stated
[ fig.05 ] explain-pr
THE DIFFread bya stranger1 WHERE THIS SITS2 WHY EACH DECISION WENT THAT WAY3 WHAT YOU NEED FOR THE NEXT CHANGE4 WHAT WOULD MAKE IT WRONGTHE AUTHOR CANNOT SEE THE REASONING THAT NEVER REACHED THE DIFF
a lesson from a reader who did not make the changeFIG

voice-writing

You can tell when a machine wrote something even when you cannot say what gave it away. This skill names the shapes and strips them, one paragraph at a time, in four passes that each carry one rule and repeat until a sweep changes nothing. A rewrite keeps every fact, hedge and attribution, checked against a ledger built before a word is touched. The profile that ships is mine; it will build one from your own writing.

Skill · voice profile · rewrite and author modes
[ fig.06 ] voice-writing
ONE PARAGRAPH AT A TIMEFOUR PASSES, ONE RULE EACHJOINSREADERASIDESDECORATIONAGAIN, UNTIL A SWEEP CHANGES NOTHINGEVERY FACT, HEDGE AND ATTRIBUTION SURVIVES
four passes, paragraph by paragraphFIG

praxis

An AI usage coach built as a loop. Commit to one behaviour for the week, get it cued into every Claude Code and Codex session through hooks, answer a thirty-second check-in afterwards, and review what actually changed. The scoring underneath, six prompting dimensions and a learning-or-atrophying trajectory read from your own sessions, exists so the coaching is honest about whether you did the thing.

Python · Claude Code, Codex and Copilot hooks · weekly loop
[ fig.07 ] praxis
COMMITCUEREFLECTONE BEHAVIOUR A WEEK · IN-SESSION
commit, cue, reflect, reviewFIG

rust2py

It migrates whole Rust systems to Python on its own, file by file, with full type safety and behavioural equivalence testing, learning from its own failures as it goes. Tested on real third-party crates: json-rust (4,710 lines, circular deps, unsafe blocks) translated end to end with 796 tests passing under mypy --strict.

Python · tree-sitter · tested on real crates
[ fig.08 ] rust2py
.RSAGENT.PY796 TESTS GREEN · mypy --strict
agentic code migrationFIG

RLM-RS

A reference implementation of Recursive Language Models (arXiv:2512.24601): a loop that runs model-written Python over corpora far larger than any context window, sandboxed in Lambda, with span citations you can check and budgets you can set.

Python · AWS Lambda · paper implementation
[ fig.09 ] rlm-rs
CORPUS · 3.2M TOKENSWINDOW · 128KSPAN · CITEDRECURSE UNTIL IT FITS · arXiv:2512.24601
recursive language modelsFIG

pptx-layout

A pure-Python engine that lays out PowerPoint slides: flexbox and CSS Grid semantics for decks, resolved down to EMU coordinates. The same input always produces the same slide, which matters once something automated is generating them.

Python · layout engine · deterministic
[ fig.10 ] pptx-layout
FLEXBOX FOR SLIDES · NO DRIFT
deterministic slide layoutFIG

Claude Skills

Reusable skill frameworks for AI consulting and GenAI training delivery: creative ideation, consulting playbooks and a five-tier training curriculum, written as modules an agent loads directly.

Markdown · skills authoring · on GitHub
[ fig.11 ] claude-skills
CONSULTING.SKILLTRAINING.SKILLEVALS.SKILLAGENTREUSABLE DELIVERY FRAMEWORKS
skills as modulesFIG

CringeIn

A Chrome extension that spots cringe LinkedIn posts and blurs them as you scroll, with caching and an adjustable sensitivity dial. I built it because my feed had become genuinely hard to read.

JavaScript · Chrome extension · side project
[ fig.12 ] cringedin
MINMAX
sensitivity dial - blur onFIG

More experiments on GitHub ↗