▮▮▮SYS.LOG // 04 agents deployed // unattended since day one
Agents that run unattended
Four systems that read a repo, draft the mail, pull the filings and publish the
report. No prompt box, no operator, no demo mode — a schedule, a toolbelt,
and a ledger that records what every run cost.
▮ agent_ticket_system / live▮ agent_auto_system / 4 crews scheduled▮ finance_data / daily pipeline▮ investskill / open source▮ every run written to sqlite
▮ agent_ticket_system / live▮ agent_auto_system / 4 crews scheduled▮ finance_data / daily pipeline▮ investskill / open source▮ every run written to sqlite
[ 001 ] The loop
Five stages, no hand-offs
Every system on this page walks the same five stages. Only the source and the toolbelt
change. Pick one.
Stage 01 / 05
Signal
Touches
signal / sources▮
[ 002 ] The systems
Four things I built and kept running
012025
Agent Ticket System
Give it a local path or a GitHub URL. It clones, indexes and reads the codebase, then writes tickets an engineer can actually act on — title, description, acceptance criteria, priority. Served over REST. No database anywhere in it.
llmfastapiuvicornpythonpytest
StackFastAPI / PythonStorenone — 0-DBTestspytest in CIStatus▮ liveView repo ↗
022025
Agent Auto System
Multi-agent office automation on CrewAI and GPT-4o. Four crews draft email, summarise documents and run US equity analysis on a schedule. Every run is costed into SQLite and surfaced in a Streamlit dashboard, so the bill is never a surprise.
Collects SEC filings — 10-K, 10-Q, 13-F, 6-K — for 30+ US equities, structures them into fundamental, insider and technical reports with Claude, then rebuilds and publishes the whole site daily from GitHub Actions. Nobody presses build.
21 structured prompt frameworks that turn any LLM into an equity analyst — Piotroski scoring, DCF, 13-F tracking, earnings-call sentiment. Slash commands inside Claude Code, plain text everywhere else. Nothing to install, nothing to run.
agent_auto_systemOne entry point, one schedule, four crews that share nothing, and one ledger they all write to. Take a box to read it.
[ 005 ] KV cache
Read the prefix once, pay for it once
A job that fires every weekday morning sends the same system prompt, the same toolbelt
and the same output schema every single time. Attention already did that arithmetic
yesterday. Keeping the prefix warm is the difference between paying for the preamble
250 times a year and paying for it once.
Prefill tape — 48k prompt, 1k a blockwarm prefix
cached — read at 0.1× 36krecomputed this run 12k
Time to first token775 ms
Prefix recomputed12k
Input cost, one run$0.047
A model, not a benchmark. 48k of prefix at $3 / MTok in and $0.30 / MTok on a cache
read — the published ten-to-one cache discount. Latency assumes prefill scales
with the uncached tail over a 300 ms floor. Bars are read against the cold run.
[ 006 ] Retrieval
Nearest is not the same as right
finance_data answers out of filings, not out of the model. Embedding search gets you
close — close is a synonym, a stale press release, a paragraph that mentions the
word. The reranker is where the top-k stops being trivia and starts being evidence.
01Queryquestion → sub-questions
02Embed1536-d, cached by hash
03Searchtop-40 by cosine
04Rerankcross-encode → top-4
05Groundanswer + span cited
Candidates — “how did gross margin move, and why”
10-K · mgmt-disc §70.84kept
earnings call · Q&A 020.81kept
10-Q · liquidity §30.79kept
press release · 04-180.77kept
10-K · cost of revenue0.74dropped
10-K · risk factors 1A0.72dropped
blog · 2024 recap0.70dropped
8-K · item 2.020.66dropped
Same forty candidates, two orderings. Cosine keeps the press release and the blog recap;
the cross-encoder throws both out and pulls the cost-of-revenue note up to third —
the paragraph that actually carries the number. Scores are illustrative, the swap is not:
two of the four the vector search was proudest of do not survive a reader.
[ 007 ] Agentic loop
A loop that knows how to stop
An agent is four moves on repeat, and the fourth is the one people skip. Anything can
observe, plan and act; a system you can leave alone overnight is one that checks its own
work and then decides the run is over.
observe / inputs▮
Halt conditions — the loop exits on the first one that fires
Goal met
The verifier passes and the artefact exists on disk. Nothing else counts as done.
Budget spent
Tokens and dollars are decremented per step. At zero the run stops and says so.
Tool refuses twice
One retry with backoff. A second failure is a fact about the world, not a prompt problem.
Step ceiling
Twelve. A loop with no ceiling is not an agent, it is a bill.
[ 008 ] Harness engineering
The model is the easy part. The harness is the product
Swapping one frontier model for another is an afternoon. The year goes on everything
around it: what is allowed into the window, what gets thrown away, what happens on the
second failure, and what is written down before anybody reads the answer.
Context window — 200k, one step of one run
tools
retrieved
history
scratch
free
Tools + system18k
Retrieved96k
History62k
Scratch14k
Headroom10k
190kIn the window
10kLeft for the answer
22%Of it actually read
A model, not a benchmark. Compaction is a summariser over the history. The harness adds
three things on top: only the tools this step can use, retrieval that returns spans
instead of whole documents, and a pinned block of facts that never gets summarised away.
Budget, not vibes
Every step decrements a token and dollar allowance. The ledger is written before the output is read.
Idempotent tools
A retried send must not send twice. Side effects carry a key, not a hope.
Fail closed
An unparseable answer is a failed run, not a shipped one. Schema first, then the write.
Replayable
Inputs, prompt hash and output land in runs.db, so yesterday’s eight o’clock can be re-run today.
[ 009 ] Open to build
Got something that has to run without anyone watching?
That is the part I enjoy — the agent being right when nobody is checking.
Every repo above is open. Mail me if one of them is useful to you.