How I Build Agents
I’ve been an early user of coding agents like Claude Code, Codex and many others, but a general purpose agent that I could make my own and use for non-coding related tasks, was more interesting. Very “early open-source software feels”.
I’ve been testing and building agents since OpenClaw (back when it was called ClawdBot in Dec 2025 or so). It was exciting - a shiny new tool to play around with, with the potential to do all the work I don’t like doing? sign me right up!
OpenClaw quickly became too cumbersome to build, maintain and use. It was annoying to work with to do even the most basic work with. I tried Goose, Zo Computer and a bunch of others too.
I then tried Pi and that really opened up my mind on the potential of agents. Pi is designed for coding agentic work, but it is compact enough to be repurposed for other activities too.
With agentic workflows, I was most worried about token usage and therefore cost. Agents are notorious token guzzlers, with no real guardrails. Pi was the first one where I knew it was specific in its task and tool management but moreover, it’s system was not built with the bloat that comes with Claude Code or Codex. The agent harness was light weight - 1,000 or so tokens consumed to respond to “hello”, vs. 15-20K tokens consumed by Claude for the same input prompt.
I repurposed Pi to do agentic stock trading for me on the Indian markets. Because Pi was built such that the agent had access to its own docs, it could self-heal and improve any aspects of its own code. That felt powerful. Pi was able to work autonomous all through the day call tools, search the internet, learn from trading mistakes, plan for the next day and so on. Pi made learning its system fun.
In the process I learned how to “code” and agents’ “soul”, give it arms and legs (tools), help it to think, give it memory, help it with memory and task management, frontend GUI development, virtual private server (VPS) deployment, pseudo PuTTY control and deployment and much more. Would never have learned any of this stuff even last year (2025).
I then thrashed all of the code and moved to Hermes - thanks to Nikhil Jois. Hermes felt like a steep learning curve, so I never gave it a shot despite the immense social media buzz around it. But once I started using it, it felt very natural. Moreover, things just worked with Hermes, as opposed to my experience building and using openclaw.
Just as things were settling down on personal agents, Instinct and then Muse were launched. Their agentic harnesses are just too good. Grok Bot is just too expensive for me to use at the moment. But I’m sure Claude, OpenAI, deepseek and every other major labs are prepping their personal agent products, so hopefully, we get to enjoy low priced/free personal agents until then.
I currently have three agents live right now: stardust trades NSE equities with real capital, vita tracks my family’s health, medications and appointments, vera manages all my investing, emailing, research related activities. Different domains, different stakes, same skeleton.
All of this context, to say, I’ve learned a lot and wanted to share some of my learnings, tips and tricks. I’m also hoping this attracts like-minded people who reach out and we learn and improve together. So here goes.
The stack
All my agents are built and managed on Hermes. Instinct is the other personal agent I’m dailying, at least until Muse launches. I use OpenRouter for all LLM calls and reasoning. I use Deepseek within OpenRouter for the same. It costs me barely $25-30/month in API usage.
I try and do everything in markdown - all pdfs, images, etc, everything is converted and stored in markdown for agentic use. It’s token efficient, which is why I can do everything really affordably. PDFs and image files have a lot of unnecessary bloat tokens which makes agentic token consumption high. Further, agents don’t even need to see styling, so I don’t use HTML. Markdown is pure and efficient.
One profile, one job
Every agent is its own Hermes gateway profile — its own systemd service, its own config root, its own bot identity, its own process. Not a shared assistant with a trading mode and a health mode bolted on. stardust can crash, get rate-limited, or get corrupted state without vita ever knowing it happened, because they don’t share a process, a context window, or a directory.
The model reasons. The code still decides.
Every agent splits into the same three layers:
- A skill — a markdown playbook, loaded fresh into context per job. It teaches the model how to think: the mandate, the invariants, the accumulated pitfalls. It contains no code the model can run from memory.
- A tool — the only thing the model is allowed to act with. Every unit of real work (a quote, a medication log, an email triage) is a named tool, never prose telling the model to run a script.
- An engine — the already-tested CLI or code path the tool shells into.
stardust’s tools all funnel throughstardust_cli.py.vita’s all funnel through a SQLite-backed CLI. Same shape both times.
The model never touches the broker, the database, or the filesystem directly, and never runs Python it wasn’t handed. It reasons in language, picks a named tool, and the tool runs the same hardened path a human would run by hand. This is also why none of these use an MCP server in front of an API that’s already just functions — wrapping an already-thin layer a second time just means two credential surfaces to keep in sync instead of one.
Tools in a toolshed
I learned this process/workflow from Stripe Dev’s engg post on Minions - Stripe’s internal, one shot end-to-end coding agent. Their Minion system is built on top of Goose, by Block - Jack Dorsey’s financial services company.
I found that explaining every action in the SKILLS.md file doesn’t lead to consistency of outcomes. The agent knows the skill, because it loads the skill in its session context. But if you tell it to summarise a pdf because you defined it in the SKILLS.md file, it can fumble hard and not know for sure whether to convert the file to markdown, then summarise, or upload the file to an API service, read the text and then summarise or what. It wouldn’t know how to handle images in the PDF file etc. It’s small stuff like this you’d get frustrated with if you gave the job to a person and they fumbled. This ruins the quality and consistency of the outcome.
Building a tool your agent knows of and can use, addresses exactly this. An agentic tool is basically set of deterministic instructions and processes with a defined outcome. IMO, without tools, the agent is basically a headless chicken. Tools are like giving your agents arms and legs to actually do the work, while the SKILLS.md file is the brain, telling your arms and legs what to do.
As your tool-set increases, you need to build a toolshed, an index of sorts, which your agent can refer to know which tool makes the most sense to use for the given task at hand. My vera investment agent has some 70-80 tools across 6-7 categories. stardust stock trading agent has 35-40 tools, across 4-5 categories.
Any unit of work that can be defined, that’s repetitive and that I expect to have a specific outcome gets made into a tool, added to the toolshed and ready for use by the agent.
If the toolshed is small enough, I’ve experimented with having a simple json file listing out all tools, and for larger toolsheds, I’ve experimented with having an sqlite database file, where again, the agent has the full schema/structure of.
Ground truth never lives in the model
Depends from use case to use case, but the source of truth is important. stardust’s own ledger can disagree with the broker’s numbers. When it does, the broker’s numbers become the source of truth, and stardust’s ledgers get corrected to match, never defended. vita works the same way: the SQLite DB is the record, the model’s conversational memory of “I think we logged that already” is not. For vera, the source of truth is the md file vault in github.
Source of truth is crucial for grounding to ensure there’s zero room for hallucinations.
The pattern generalises: whatever external system already has to be right for other reasons (a broker, a database, an inbox) is the one source of truth. The model’s own memory is a cache that can be stale or wrong, not a second opinion that gets to argue with it.
Gate expensive reasoning behind cheap logic
Not every decision needs an LLM call. stardust’s cadence controller is a five-line deterministic state machine that runs every five minutes and decides whether the trading agent should even wake up. It only rewrites the cron schedule when the tier actually changes, never drives on weekends, and costs zero tokens.
vita does the same thing structurally rather than behaviorally: every mutating tool call and every outbound reply passes through Jev, a small, fast confidence-gating model first — is this value actually stated by the user, does it look like a duplicate, is this drifting into clinical advice — before the action is allowed through. The expensive model only gets invoked once the free/cheap gate has already decided the call is worth reasoning about at all.
Across all three agents, the rule is the same: script what’s deterministic, reserve reasoning for what actually requires judgment, and put a cheap check between the trigger and the expensive call wherever one is possible.
Bound the blast radius structurally, not behaviorally
None of the hard limits in any of these agents exist because the prompt asks nicely. They exist because the tool layer makes the wrong action physically unavailable:
stardustcan only act on a security already present in its own tracked state — anything else on the broker account has to be deliberately adopted first.- Every live entry immediately places a real exchange-side stop order, so protection exists between wake-ups, not just when the agent happens to notice a breach.
vitastructurally blocks any reply that gives clinical judgment rather than logging or summarizing — enforced by the gate, not by an instruction the model could choose to ignore under pressure.verahas multiple entry points (email, buzz, telegram etc.) but it’s all bound by approved channels only. It knows which channel it has been invoked in and who invoked it and therefore knows how to react. This helps to avoid prompt injection issues. Vera also has an org chart tool, so it can even tag other agents or humans when needed for specific tasks.
If a mistake would be expensive, I don’t write a rule against it — I make the tool layer incapable of executing it.
Foreman - the task manager behind my agents
For my day-to-day, like anyone else, I juggle a lot of tasks - portfolio support, deal evaluation, sourcing, investor updates, personal work and much more.
Unfortunately, I realised too late in life that it’s best to complete a task immediately instead of procrastinating. I hate to admit it but sometimes I procrastinate on tasks and more often than not I end up dropping the ball.
Us humans have a finite context length - only so many things we can keep in mind. Sleep helps us start new context lengths. Agents work the same way too, they have limited context length 250K to 1M tokens. However, with agents, a new conversation is like a factory reset on the agent. That’s just not acceptable when you want to get work done.
Moreover, sometimes there are tasks that need to be completed over days or weeks. Plain reminders don’t cut it. The agent needs to load the context of that task, the last known state and the steps needed to complete it.
So I (and my coding agents) built Foreman, an open source Task Manager for Agents. With Foreman my agents don’t drop the ball on tasks that need to be completed. It gives agents a durable place to keep the work itself — objectives, plans, decisions, state, progress, blockers, and artifact references — independent of the agent runtime doing the work.
Agent runtimes change quickly. A project might start in Hermes today, continue in OpenClaw tomorrow, and eventually be picked up by another agent entirely.
Without a durable work layer, the useful state of the project is trapped inside the conversation or inside one runtime.
Foreman separates those concerns:
- Agent runtime — executes the work, uses tools, reasons, and communicates with the user. Foreman — owns durable representation of the work and makes it resumable.
- Artifacts — hold substantial research, documents, code, and other outputs; Foreman stores references to them rather than stuffing their contents into every context window.
- The goal is not to replace an agent’s native task system. A runtime such as Hermes can remain the execution backend.
Foreman provides a portable work representation that can survive a runtime change.
The prompt is a living incident log, not a spec
There’s no clever system prompt anywhere in this. The base identity (SOUL.md) is specific to the domain but deliberately high level. All of the actual judgment and decision making lives one level up, in a skill file that gets rewritten every time something breaks in production.
A representative entry from stardust’s pitfall log: positions() and holdings() are different broker endpoints, and a delivery position from a prior day won’t show up in the first one. That’s not a comment — it’s a standing instruction, re-read every single session, written generally enough to prevent the whole class of bug rather than just that one instance.
A session restart should never mean re-learning the same lesson twice. The skill is the institutional memory; the model is replaceable underneath it.
The shape, restated
One agent profile per job, isolated down to the process. A skill that reasons, a task manager to not drop the ball, a tool that’s the only way to act, an engine that’s already been tested. One ground truth outside the model, deferred to completely. Cheap logic gating expensive reasoning wherever a cheap check is possible.
Limits enforced by what the tools can do, not by what the prompt asks for. Every production mistake folded back into the thing that gets re-read next time. Scope that widens only after the narrower version has held up. That’s the system — the domain is just whatever gets plugged into it next.
Onward.