Give each agent a clear job
I use Claude Code, Codex, and Cursor through Conductor. Separate workspaces keep investigations, features, and fixes organized. Each task carries its context, scope, open questions, and expected result.
I developed Robust Agents to structure my work with AI, from investigation and planning to implementation, independent review, and testing. I stay responsible for the decisions and what ships.
Before implementation, I establish the user problem, inspect the code, and separate what we know from what we still need to decide. I turn that into a short spec, milestones, and checks that define a working result.
MCP integrations connect agents to project tools: product context from Notion and Slack, designs from Figma or Paper, and behavior from Sentry, Mixpanel, and BigQuery. GitHub, Context7, Expo, and Vercel support code research, documentation, testing, and delivery. I use the source that can answer the question.
I use Claude Code, Codex, and Cursor through Conductor. Separate workspaces keep investigations, features, and fixes organized. Each task carries its context, scope, open questions, and expected result.
Small changes stay simple. For larger or uncertain work, my Robust Agents workflow separates implementation from independent review. Reviewers check the actual change against the original problem, and I verify findings before accepting a fix.
Specs, source links, decisions, and test evidence stay with the task. That lets me return to an investigation, hand work to another agent, or understand why a particular approach was chosen.
Agents help investigate options and carry out defined work. I decide the scope and tradeoffs, align product questions with the team, and keep merges and releases under explicit control.
This replays how a review runs on a small example: a login change in a demo app. The app, code, and findings are invented for this page. The steps match my workflow: classify the risk, review independently, then verify each finding in the code.
Pull requestAdd “Remember me” to the login formnotes-demo, an example app
18const ONE_HOUR = 60 * 60;
19const THIRTY_DAYS = 30 * 24 * ONE_HOUR;
20
21res.cookie("session", token, {
22 httpOnly: true,
23 sameSite: "lax",
24 maxAge: rememberMe ? THIRTY_DAYS : ONE_HOUR,
25});
Express reads maxAge in milliseconds. A test login without “Remember me” expired after 3.6 seconds.
3export const loginBody = z.object({
4 email: z.string().email(),
5 password: z.string().min(1),
6 rememberMe: z.boolean().optional(),
7});
The schema accepts only a boolean. A request with "false" as a string returns 400 before the handler runs.
One confirmed finding. Multiply maxAge by 1,000 and add a test for both cookie lifetimes.
A successful build is one check. I also test the flow a person will use. For native work, cloud workspaces connect to iOS simulators and Android emulators on my Mac through my remote testing setup.
Argent helps drive the app and capture screenshots or recordings. Separate simulator lanes keep parallel tasks tied to their own workspace. I check the installed app and the services it connects to before trusting the result.
I can manage cloud work and inspect progress from my phone, with the Mac providing the native environment.
I keep reusable instructions for investigations, code review, device QA, writing, and session handoffs in a versioned personal repository. Shared context carries the relevant project decisions across agents and machines.
My Conductor routines cover production scans, release monitoring, inbox triage, and learning. A routine points to its instructions in the repository, so improvements have one maintained source.
Monitoring combines errors, product data, and user feedback. Findings need evidence and a clear impact before they become work. Temporary monitors have a stopping condition so they don’t keep producing reports after the question is resolved.
I keep active work in Conductor and priorities, ideas, and parked tasks in Notion. Short updates, early demos, and explicit next actions help me coordinate several initiatives without losing the thread.
Read the report, inspect the affected flow, and connect it to code and production evidence. Record what is still uncertain.
Give the agent the scope and acceptance checks. Keep unrelated cleanup outside the task.
Check review findings against the code. Reproduce the flow on the relevant platform and attach evidence.
Track the rollout and relevant signals. Close the task when the evidence supports it, and capture useful lessons for the next one.
I keep up with new models and study emerging AI patterns. I try them on actual tasks, compare the results, and refine my skills and instructions around what proves useful.
I form a hypothesis, make a small prototype, and share it early. User feedback, analytics, and design conversations help me decide whether to continue, change direction, or stop. Product taste and scientific thinking both have a place in that decision.
Moving background and subtle signal interference.