Robust Agents / The workflow I developed

A system for
building with AI.

I developed Robust Agents to structure my work with AI, from investigation and planning to implementation, independent review, and testing. I stay responsible for the decisions and what ships.

Make the unknowns
visible.

Before implementation, I establish the user problem, inspect the code, and separate what we know from what we still need to decide. I turn that into a short spec, milestones, and checks that define a working result.

MCP integrations connect agents to project tools: product context from Notion and Slack, designs from Figma or Paper, and behavior from Sentry, Mixpanel, and BigQuery. GitHub, Context7, Expo, and Vercel support code research, documentation, testing, and delivery. I use the source that can answer the question.

  1. 01Understand the problem
  2. 02Define the result
  3. 03Build and challenge
  4. 04Test and follow through

Focused work.
Independent review.

Give each agent a clear job

I use Claude Code, Codex, and Cursor through Conductor. Separate workspaces keep investigations, features, and fixes organized. Each task carries its context, scope, open questions, and expected result.

Match the process to the risk

Small changes stay simple. For larger or uncertain work, my Robust Agents workflow separates implementation from independent review. Reviewers check the actual change against the original problem, and I verify findings before accepting a fix.

Keep a reviewable trail

Specs, source links, decisions, and test evidence stay with the task. That lets me return to an investigation, hand work to another agent, or understand why a particular approach was chosen.

Keep human decisions explicit

Agents help investigate options and carry out defined work. I decide the scope and tradeoffs, align product questions with the team, and keep merges and releases under explicit control.

Two reviewers.
One verdict.

This replays how a review runs on a small example: a login change in a demo app. The app, code, and findings are invented for this page. The steps match my workflow: classify the risk, review independently, then verify each finding in the code.

Pull requestAdd “Remember me” to the login formnotes-demo, an example app

  1. Risk classifierSorts the diff by riskReported
  2. Reviewer AClaudeReported
  3. Reviewer BCodexReported
  4. VerifierChecks each finding in the codeReported
  1. Risk classifierThe diff changes how long a login session lasts.
  2. Risk classifierRisk: Nontrivial. Route: two independent reviewers, then a verifier.
  3. Reviewer AReading the diff against the original issue.
  4. Reviewer BReading the diff and the login tests.
  5. Reviewer AmaxAge is set in seconds, but Express expects milliseconds.
  6. Reviewer BrememberMe could arrive as the string “false”, which is truthy.
  7. Reviewer AI disagree. The request schema should reject a string before the handler sees it.
  8. Reviewer BAgreed on the cookie units. That one blocks the merge.
  9. VerifierRunning a test login without “Remember me”.
  10. VerifierConfirmed: the session expired after 3.6 seconds.
  11. VerifierRejected: posting "false" as a string returns 400.

Findings

The session cookie expires after seconds, not hoursConfirmed
src/auth/login.ts
const ONE_HOUR = 60 * 60;
const THIRTY_DAYS = 30 * 24 * ONE_HOUR;

res.cookie("session", token, {
  httpOnly: true,
  sameSite: "lax",
  maxAge: rememberMe ? THIRTY_DAYS : ONE_HOUR,
});

Express reads maxAge in milliseconds. A test login without “Remember me” expired after 3.6 seconds.

“false” as a string would keep every user signed inRejected
src/auth/schema.ts
export const loginBody = z.object({
  email: z.string().email(),
  password: z.string().min(1),
  rememberMe: z.boolean().optional(),
});

The schema accepts only a boolean. A request with "false" as a string returns 400 before the handler runs.

Changes requested

One confirmed finding. Multiply maxAge by 1,000 and add a test for both cookie lifetimes.

See it work.
On the device.

A successful build is one check. I also test the flow a person will use. For native work, cloud workspaces connect to iOS simulators and Android emulators on my Mac through my remote testing setup.

Argent helps drive the app and capture screenshots or recordings. Separate simulator lanes keep parallel tasks tied to their own workspace. I check the installed app and the services it connects to before trusting the result.

I can manage cloud work and inspect progress from my phone, with the Mac providing the native environment.

The useful parts
become repeatable.

Skills hold the method

I keep reusable instructions for investigations, code review, device QA, writing, and session handoffs in a versioned personal repository. Shared context carries the relevant project decisions across agents and machines.

Routines handle recurring work

My Conductor routines cover production scans, release monitoring, inbox triage, and learning. A routine points to its instructions in the repository, so improvements have one maintained source.

Signals become next actions

Monitoring combines errors, product data, and user feedback. Findings need evidence and a clear impact before they become work. Temporary monitors have a stopping condition so they don’t keep producing reports after the question is resolved.

Organization keeps work moving

I keep active work in Conductor and priorities, ideas, and parked tasks in Notion. Short updates, early demos, and explicit next actions help me coordinate several initiatives without losing the thread.

From a report
to a checked fix.

  1. 01 / Investigate

    Find the user’s actual failure.

    Read the report, inspect the affected flow, and connect it to code and production evidence. Record what is still uncertain.

  2. 02 / Build

    Make the smallest complete change.

    Give the agent the scope and acceptance checks. Keep unrelated cleanup outside the task.

  3. 03 / Verify

    Challenge the result.

    Check review findings against the code. Reproduce the flow on the relevant platform and attach evidence.

  4. 04 / Follow through

    Check what happens after release.

    Track the rollout and relevant signals. Close the task when the evidence supports it, and capture useful lessons for the next one.

Keep learning.
Keep questioning.

Try new models in real work

I keep up with new models and study emerging AI patterns. I try them on actual tasks, compare the results, and refine my skills and instructions around what proves useful.

Bring product judgment to the process

I form a hypothesis, make a small prototype, and share it early. User feedback, analytics, and design conversations help me decide whether to continue, change direction, or stop. Product taste and scientific thinking both have a place in that decision.

See my work at Rosebud ↗Work with me ↗
Atmosphere
Atmosphere

Moving background and subtle signal interference.