With Claude, Vessl AI:
- Produces around 200 personalised drafts a day, up from 20
- Holds a 5% reply rate, unchanged at ten times the volume
- Sends 95% of drafts without edits, up from 73% in week one
- Cut human time per email from 30 minutes to around 9 seconds
30 minutes of research per email.
Vessl AI sells to researchers in the US, and the motion runs on conferences: before each major event, the team writes to the people presenting there. Doing it well meant knowing who was attending, what they worked on, and what they had just published.
Doing it at all meant a day of tab-switching. Check the conference calendar. Identify the research domain. Pull matching contacts from the CRM. Google each person's recent work. Draft a hook. Then do it again, for every contact.
Thirty minutes per email, twenty emails a day. The team spent the whole day on mechanical effort before a single real conversation happened. The process was well-defined. It was fixed. It was repetitive.
One pipeline. n8n moves the data, Claude makes the judgement calls.
H.ai rebuilt the workflow as a single pipeline. n8n orchestrates the schedule and the data path across HubSpot, Google Sheets, Gmail and Google Search. Every step that requires judgement runs on the Anthropic API. Each morning, before anyone opens a laptop, the pipeline checks the calendar, classifies the research domain, matches contacts against it, researches each person, and drafts a personalised email for review.
Choosing Claude 3 Sonnet for its balance of cost, performance and tone
Any model can produce 200 emails a day. The constraint was that each one had to survive a researcher reading it for two seconds and deciding it was not machine-written. Claude has a natural tone that made it the right candidate for this task. It also had to fail safely: when the research came back thin, the draft needed to say less rather than invent a paper the recipient never wrote.
Two jobs, two prompts
The pipeline never asks one prompt to do two things. Classification runs first and separately: Claude assigns each conference a research domain, then decides which CRM contacts belong to it. Only what survives that gate reaches the drafting prompt. Keeping them apart made the system debuggable: when a draft came back wrong, the team could tell whether the model had misread the person or misread the match.
The rest of the work was edge cases. Deduplicating contacts who attend several conferences. Suppressing anyone silent for 60 days. Distributing 200 daily sends across aliased accounts to protect deliverability. Giving every dynamic hook an audit trail back to its source, so a reviewer can check a claim in seconds.
A full day of outreach, approved before the first coffee.
- Baseline
- The team's own pre-deployment workflow, timed step by step
- Measurement window
- 8 weeks
- Volume observed
- Around 200 drafts per working day
| Metric | Before | After | Change |
|---|---|---|---|
| Drafts produced per day | 20 | ~200 | 10× |
| Reply rate on the sequence | 5% | 5% | Unchanged |
| Human time per email | 30 min | ~9 sec | −99% |
| Daily human time on outreach | 8 hours | 30 min | −94% |
| Meetings booked | baseline | 5× baseline | 5× |
| Drafts sent untouched | n/a | 95% | 73% → 95% |
| Data sources touched by hand | 4 | 0 | one pipeline |
The team's edits were the eval set. Every rewrite in week one was a labelled failure: wrong hook, wrong tone, wrong person. H.ai collected them and fixed the prompt with few-shot examples. By week five there was little left to learn from, which is what 95% untouched actually means.
- Week 1
- 73%
- Week 2
- 82%
- Week 3
- 90%
- Weeks 4 to 8
- 95%
Ten times the volume at an unchanged reply rate is the figure that matters. Raising output alone is trivial and usually just moves the problem into a spam folder. The reply rate holding flat at 5% is what says the drafts were still worth reading. The human approval gate never came out, and was never meant to.