The repo already knows how to check itself. The agent has no compact way to ask.
The same four lines, forever
“Run the tests, check coverage, audit dependencies, check migration status.” Watch what each side actually does with that.
............................................
plugins: cov-5.0.0, anyio-4.4.0
Name Stmts Miss Cover Missing
app/api.py 312 41 87% 88-93, 210
Found 1 known vulnerability in 1 package
INFO [alembic.runtime.migration] Context impl
… ~21.5 KB of log to read
Same work, same wording. The agent spends its thinking on the one high finding instead of on remembering flags.
Round 1, measured
Top bar: baseline. Bottom bar: with the command surface.
The gap widens as the task repeats
Run the same check twice a day for a month. The saving is not a one-off discount — it compounds.
Cumulative tool output the agent has to read. Run 1 uses measured round-1 figures; later runs use the rounds 2–3 average, where the surface is already familiar. A projection, not a benchmark.
What you actually run
One markdown Skill file. You point it at a repo; it does the design work, you keep the diff.
Let the agent decide what. Let the repo own how.