H
Book a consultation
Claude + n8n GTM pipeline

Ten times the outreach, at the same reply rate

Vessl AI's GTM team spent 30 minutes researching every cold email. With H.ai they rebuilt the motion as one pipeline: n8n moves the data, Claude makes the judgement calls, a human approves.

Client
Vessl AI
Industry
B2B sales
Solution
Claude + n8n GTM pipeline
Status
Deployed
Measured
8 weeks

With Claude, Vessl AI:

  • Produces around 200 personalised drafts a day, up from 20
  • Holds a 5% reply rate, unchanged at ten times the volume
  • Sends 95% of drafts without edits, up from 73% in week one
  • Cut human time per email from 30 minutes to around 9 seconds
The challenge

30 minutes of research per email.

Vessl AI sells to researchers in the US, and the motion runs on conferences: before each major event, the team writes to the people presenting there. Doing it well meant knowing who was attending, what they worked on, and what they had just published.

Doing it at all meant a day of tab-switching. Check the conference calendar. Identify the research domain. Pull matching contacts from the CRM. Google each person's recent work. Draft a hook. Then do it again, for every contact.

Thirty minutes per email, twenty emails a day. The team spent the whole day on mechanical effort before a single real conversation happened. The process was well-defined. It was fixed. It was repetitive.

The solution

One pipeline. n8n moves the data, Claude makes the judgement calls.

H.ai rebuilt the workflow as a single pipeline. n8n orchestrates the schedule and the data path across HubSpot, Google Sheets, Gmail and Google Search. Every step that requires judgement runs on the Anthropic API. Each morning, before anyone opens a laptop, the pipeline checks the calendar, classifies the research domain, matches contacts against it, researches each person, and drafts a personalised email for review.

Choosing Claude 3 Sonnet for its balance of cost, performance and tone

Any model can produce 200 emails a day. The constraint was that each one had to survive a researcher reading it for two seconds and deciding it was not machine-written. Claude has a natural tone that made it the right candidate for this task. It also had to fail safely: when the research came back thin, the draft needed to say less rather than invent a paper the recipient never wrote.

Two jobs, two prompts

The pipeline never asks one prompt to do two things. Classification runs first and separately: Claude assigns each conference a research domain, then decides which CRM contacts belong to it. Only what survives that gate reaches the drafting prompt. Keeping them apart made the system debuggable: when a draft came back wrong, the team could tell whether the model had misread the person or misread the match.

The rest of the work was edge cases. Deduplicating contacts who attend several conferences. Suppressing anyone silent for 60 days. Distributing 200 daily sends across aliased accounts to protect deliverability. Giving every dynamic hook an audit trail back to its source, so a reviewer can check a claim in seconds.

Results, measured

A full day of outreach, approved before the first coffee.

Baseline
The team's own pre-deployment workflow, timed step by step
Measurement window
8 weeks
Volume observed
Around 200 drafts per working day
Metric Before After Change
Drafts produced per day 20 ~200 10×
Reply rate on the sequence 5% 5% Unchanged
Human time per email 30 min ~9 sec −99%
Daily human time on outreach 8 hours 30 min −94%
Meetings booked baseline 5× baseline 5×
Drafts sent untouched n/a 95% 73% → 95%
Data sources touched by hand 4 0 one pipeline

The team's edits were the eval set. Every rewrite in week one was a labelled failure: wrong hook, wrong tone, wrong person. H.ai collected them and fixed the prompt with few-shot examples. By week five there was little left to learn from, which is what 95% untouched actually means.

Share of drafts sent without human edits
Week 1
73%
Week 2
82%
Week 3
90%
Weeks 4 to 8
95%

Ten times the volume at an unchanged reply rate is the figure that matters. Raising output alone is trivial and usually just moves the problem into a spam folder. The reply rate holding flat at 5% is what says the drafts were still worth reading. The human approval gate never came out, and was never meant to.

Want something similar for your GTM motion?

H.ai will walk through your stack, your edge cases, and what a human-in-the-loop Claude pipeline could look like. A practical conversation, no pitch deck.