There is a thread in r/gtmengineering called “My Claude Code Setup to Automate Cold Email” and another in r/AiAutomations titled “How I automated my Cold Email Outreach with Claude Code”. Both are worth reading, and both converge on the same lesson: the agent is brilliant at the boring middle of outbound, and dangerous at the edges.

I run outbound systems for clients at KomsGro, and I have been using agent CLIs in that work since they got good at reading files, calling APIs, and running in loops. This is the practical version of what I delegate, what I refuse to delegate, and the setup that makes it repeatable.

The split: what Claude Code should and should not do

Think of cold email as a factory with a research department, a copywriter, a QA desk, and a dispatch office. Here is the honest assignment.

Delegate to the agent:

  1. Research briefs per account. Given a domain, pull the context that makes a first line credible: what they sell, who they hire for, recent announcements, the stack they run. This is the single highest-leverage use, because it is the task humans skip when volume rises.
  2. Drafting the first line. Not the whole sequence, the first line. One specific, verifiable sentence per prospect, drawn from the research brief.
  3. List QA. Finding the rows that break your rules: missing fields, free-email addresses, obvious competitors, duplicates against the CRM.
  4. Reply triage. Classifying incoming replies into interested, not now, wrong person, and unsubscribe, then routing each into the right folder with a suggested next action.
  5. Sequence drafting and rewriting. Turning a rough outline plus three real examples into a four-touch draft that a human edits.

Never delegate: 6. Sending. The agent does not talk to the sender. Volume, rotation and reputation belong to a purpose-built tool and a human watching it. 7. Deliverability configuration. Domain choices, mailbox counts, authentication, and warmup are decisions, not tasks. Automating them is how you burn domains. (The mailbox math is in how many mailboxes per domain.) 8. The final approval. Every message that leaves is approved by a human. This is not a technical limitation, it is a judgment call about accountability. An agent that can send unsupervised is one bad prompt away from emailing your whole list twice.

The setup that makes it repeatable

The difference between a fun experiment and a system is that the system has a repo. Mine looks like this:

gtm-agent/
  context/          # who we are, what we sell, ICP definition, tone rules
  leads/            # one CSV or JSON per campaign, the only source of truth
  research/         # agent output: one brief per account
  drafts/           # agent output: first lines and sequences, pending review
  logs/             # every run: what was processed, what was skipped, why
  skills/
    research.md     # the research brief prompt, with the output schema
    firstline.md    # the copy rules, with examples of good and bad
    qa.md           # the list rules, expressed as checks
    triage.md       # reply classification categories and routing

Four practices make this work:

Schemas, not vibes. Each skill file defines the exact output shape. A research brief is a fixed set of fields; a first line is one sentence under a character limit with the source fact cited. When the agent deviates, the run fails loudly instead of quietly producing slop.

Small batches with review. Process 25 accounts, review the output, then scale. The temptation is to point it at 5,000 rows on day one. The teams that do that get 5,000 mediocre first lines and no idea which part broke.

Everything logged. Every run writes what it did and why it skipped a row. When reply rates move, the log tells you whether the research got worse or the list did.

Human review as a node, not a hope. The review step is a folder, not a promise. Drafts sit in drafts/ until a person moves them to approved/. Anything that bypasses the folder is a bug.

A worked example

Take one account. The workflow I actually run:

  1. Input: a domain and a contact name.
  2. Research: the agent reads the site, the careers page, and a news search, then writes a brief with five fields: what they do, who they sell to, a recent signal, a probable pain, and one verifiable fact.
  3. First line: a second skill reads the brief and writes one sentence that uses the verifiable fact and connects it to our offer. It must cite which field it used.
  4. Review: I read twenty of these in five minutes, mark the bad ones, and refine the skill file. The refit is usually a tighter example set, not a longer prompt.
  5. Sequence: a third skill expands the approved first line into a four-touch draft following our structure, then it waits.
  6. Send: the approved drafts go into the sender manually or via a queue I control.

The whole loop costs cents per account and removes the two tasks that make outbound feel like a treadmill: research and the blank page.

Claude Code vs Cursor vs Codex for this job

You will see this comparison in every GTM community right now, and the honest answer is that the differences matter less than the setup around them. From our own use and the public threads:

  • Claude Code is the strongest at long, multi-step, file-and-terminal workflows, which is exactly what a research-to-draft pipeline is. It holds context across a chain of steps well, and its file handling suits the repo structure above.
  • Cursor shines when a human and the model are editing the same file together. For iterating on one prompt file or one script, its in-editor loop is faster.
  • Codex is a capable alternative that shows up in the community comparisons, and many teams run two agents side by side, one for research chains and one for code.

My recommendation: pick one agent and invest in the repo, the schemas, and the review loop. Switching agents changes maybe ten percent of the outcome; having no schemas changes all of it.

The guardrails that keep this from going wrong

  • No PII in prompts you do not need. Send the minimum.
  • No sending from the agent. Ever.
  • Rate limits on every external call. The agent will happily loop.
  • A kill switch. One file that, when present, stops all runs.
  • A weekly audit. Read ten outputs at random. If quality slipped, refit the examples, not the prompt language.

The bottom line

Claude Code does not replace your outbound strategy. It removes the two hours per campaign that used to sit between “we have a list” and “we have something worth sending”, and it does it while leaving every judgment call with a human.

If you want the surrounding system rather than the experiment, that is what we build at KomsGro’s outbound marketing service: the infrastructure, the qualification, and the sequences, with AI doing the assembly work and people making the calls. Start with one skill file and twenty accounts. The repo is the product.

Common questions about using Claude Code for outbound

Does the agent write the emails that get sent? It drafts them. A human approves them. The workflow puts a literal folder between the agent and the sender: outputs land in a pending directory and only a person moves them into the approved one. That single design choice is the difference between leverage and a liability.

How much does this cost per campaign? Pennies per account for the research and drafting, plus your existing tools. The larger cost is the setup time for the repo and the skill files, which is a few hours once, not per campaign.

Can I use Cursor or Codex instead? Yes. The pipeline in this guide is agent-agnostic: context files, schemas, small batches, review gates. Claude Code is the strongest at long multi-step chains, which is what a research-to-draft pipeline is; Cursor is faster for editing one file at a time; Codex is a capable alternative. The repo matters more than the choice.

How do I stop it from hallucinating facts about a prospect? Require a citation field. Every research brief must name the source field each claim came from, and every first line must cite which brief field it used. Output that cannot cite is rejected automatically.

What about privacy? Send the minimum data needed for the task, keep prospect data in your own repo rather than pasting it into prompts unnecessarily, and never connect the agent to the sender. Those three rules cover most of the risk surface.

The one-file start

If the repo structure above sounds like a project, start smaller: create one folder and one skill file. Call it research.md. Put the ICP definition at the top, then five required output fields, then two examples of a good brief and one of a bad brief. Run it on twenty-five accounts.

That single file is the whole model in miniature: context, schema, examples, small batch, human review. Everything else in this guide is that loop applied to the next task.