Personal tooling · in use

Agent infrastructure

Personal tooling to run coding agents across three machines, search everything they've done, and keep an eye on machine health and provider limits.

The problem

I run coding agents across several tools and machines. That creates work on top of the task itself: deciding where an agent should run, checking whether it can keep going, and finding out what it already tried.

I want my laptop free for my own work, so builds, indexing and browser sessions run elsewhere. The catch is that remote work is harder to follow. A useful decision ends up buried in an old conversation on another machine. I built this setup to keep that work reachable and its history searchable.

The context

I use several harnesses and models. Codex, Claude Code and OpenCode each keep their own history, in their own format. The app I coordinate them from can hold a copy of the same conversation. Personal and work profiles add more stores.

When I remember a decision, I rarely remember the exact words or which agent made it. I need to search across all of those stores, read the exchange behind a result, and tell a proposed change from a finished one. I also want one view of machine resources and account limits, without logging into each machine.

Architecture and decisions

Three machines on a private network:

Architecture · three-machine private network

hrec: search across every agent conversation

hrec (now also called recall) searches past agent conversations across tools and machines. It never modifies the original histories. Exporters read each tool’s native store and extract sessions and messages with their original IDs and locations. Unchanged sessions are fingerprinted, so re-indexing skips them.

The index, the embedding model and its runtime live on pop-agent, with a replica on om-agent. Keeping them off the laptop means the Mac doesn’t carry the database, the model environment or a resident embedding process. On the Mac there’s only a small Go client that queries the desktops over SSH.

Search combines two methods. SQLite full-text search with BM25 ranking handles exact terms, with stemming and typo tolerance. A local multilingual E5 model embeds passages for semantic search. The two rankings are merged per conversation. Embeddings run on my own machines, so no conversation is sent to an external API, and if the model is down, keyword search still works.

Results start compact. From a hit I can expand into the surrounding messages or the whole session, with a character budget so it doesn’t flood an agent’s context. Every result says which machine, tool and session it came from, so I can check it before relying on it.

The example below is illustrative and uses synthetic content.

Illustrative · fictional search output

hrec

Searches every past agent conversation across tools and machines, and shows where each result came from.

$ hrec "headless browser"
Example excerpt: “Run the browser on the agent machine.”source: illustrative tool / session
machine: pop-agent

If pop-agent is unreachable, the Mac client falls back to om-agent. If both are down, search on the laptop is unavailable. I left that dependency visible on purpose rather than quietly keeping another index on the laptop.

Agent Fleet: what’s running and what’s left

Agent Fleet is a macOS menu-bar app showing machine health, recent usage and provider quota windows. The shell is Objective-C hosting web views; Python collectors report remote resource usage. It shows when quotas reset, caches its probes and marks stale data. It polls less often when hidden or in Low Power Mode.

How I delegate

Each agent gets a bounded brief: the question, the sources it may use, the files it owns, and the evidence it has to return. Research and reviews run read-only. A separate reviewer checks the result before I integrate it, and if quota limits force a different reviewer, that’s recorded.

What it does today

On 10 October 2026 the index held 452,006 messages from 5,057 sessions across four tools: Codex, Claude Code, OpenCode and T3 Code. Some conversations are stored in more than one place; after removing those duplicates, about 437,000 remain.

Most of that is the agents working: about 196,000 tool calls and 152,000 tool results. The conversation itself is about 19,000 prompts I wrote and 67,000 agent replies. All of it is searchable, because the decision I’m looking for is often in a tool result rather than in a reply.

Agent Fleet currently tracks four provider families. I haven’t measured a productivity or battery gain from either tool yet.

Limits and what’s next

Mirrored conversations are indexed more than once, so search results can repeat. Attachments aren’t transcribed, some tools’ histories can’t be read, and semantic passages can miss details buried deep in long answers. A healthy status doesn’t guarantee fresh history either; new laptop conversations normally sync only when it’s plugged in.

Next I want a public demo of the whole loop on a synthetic task: an agent does the work, hrec finds the decision, a reviewer checks it, and the fallback is visible. I also want proper measurements of retrieval quality, latency and laptop energy use.