Outbound Catalyst

Playbook

How we actually run outbound: the full playbook

Six phases, three checkpoints that block bad work from reaching an inbox, and the operational rules that took years of failed campaigns to earn.

18 min read

Key takeaways

  • Outbound is not a channel, it is a validation engine. You choose exactly who tests your message, and you find out whether your positioning, ICP, and offer hold up before you scale something broken across ads, content, and a sales team.
  • If you cannot describe the client's business to a seven year old, keep digging. Clever copy on top of a fuzzy understanding of the buyer produces confident, well-formatted emails that say nothing true.
  • Market size sets the entire strategy. Under 1,000 accounts and high-volume plays will exhaust your list before it regenerates. The tier, not the tactic, decides contact frequency and campaign mix.
  • Infrastructure math is fixed, not creative. Roughly 1.75 warmup emails for every cold email, three mailboxes per domain maximum, and sender IP geography matched to target geography before anything else.
  • Sales starts with no. A hard not-interested is worth a short follow-up asking what would need to be true for the answer to change. That answer is some of the best positioning feedback a company will ever get.
  • A campaign that stops working is almost never a copy problem first. Check the infrastructure and the targeting before you touch a word of the message.

Most outbound fails for the same reason most diets fail. Not because the mechanics are secret, but because nobody runs the full system for long enough to find out if it works. A founder buys a tool, buys a list, writes three emails on a Sunday night, sends them to five thousand people who were never qualified in the first place, and calls the result we tried outbound, it does not work for us.

It is not that outbound does not work. It is that what most companies run is not outbound. It is activity dressed up as a strategy.

We think about outbound differently, and we run it as a system with six phases, three checkpoints that block bad work from reaching a prospect's inbox, and a set of operational rules that took years of failed campaigns to earn. This is that system, in full. We are publishing it because the parts that matter most, the diagnosis, the targeting logic, the infrastructure math, the reply discipline, are not secret. They are just tedious enough that most agencies skip them and most in-house teams never get around to writing them down.

If you read one thing before anything else, read this: outbound is not a channel, it is a validation engine. You choose exactly who tests your message. You get real signal from real buyers, fast. You control the variables. And you find out whether your positioning, your ICP and your offer actually hold up, before you spend real budget scaling something broken across ads, content and a sales team. Everything below is in service of that one idea.

Phase 0 and 1: diagnose before you touch a single list

The single most common reason a campaign underperforms has nothing to do with copy. It is that nobody did the diagnostic work before building it.

Before we write anything, we run every new engagement through an onboarding process built to answer one question honestly: do we actually understand this business well enough to represent it to a stranger. The internal rule we hold ourselves to is blunt: if you cannot describe the client's business to a seven year old, keep digging. Clever copy on top of a fuzzy understanding of the buyer just produces confident, well-formatted emails that say nothing true.

The diagnostic covers ground most agencies skip entirely.

  • What the business actually does, not the pitch deck version. What problem it solves best, and just as important, who it is a bad fit for. We ask for that explicitly, because knowing who to exclude is half of targeting.
  • What a great result actually looks like, in numbers. Not more pipeline. A specific number of meetings per month, a specific pipeline figure, a specific definition of success that both sides agreed to before a single email went out.
  • The real ICP, described two ways: the account (industry, size, geography, must-have traits, technology used or avoided, hiring signals) and the persona (who the champion is versus who the actual decision maker is, because in a lot of B2B motions those are two different people who need two different messages). We ask for real profiles of both, five or six examples of each, because a persona description in the abstract is nearly useless and a persona anchored to real humans is not.
  • The pain, in the buyer's own words. Not the five pains the founder assumes. The pains prospects have actually expressed, in sales calls, in support tickets, in lost-deal notes.
  • What has already been tried. Whether outbound has run before, what worked, what did not, and why. Most companies that say outbound does not work for us ran one bad campaign and generalized from it. That history is diagnostic gold if you actually read it instead of starting from zero.
  • The CRM, not the sales team's memory of the CRM. We pull the last twelve months of closed-won deals and run a real win and loss pattern analysis: industry, company size, persona, common objections. We pull closed-lost and the stated reasons. This is where assumptions about who we sell to get corrected by what actually happened.

Out of this diagnostic we produce three things before a single list gets built: a written summary of the business and its ICP that both sides sign off on, a market mapping plan that tells the data team exactly which sources to pull from and which fields are non-negotiable, and a first-pass campaign roadmap, not the finished copy, but which campaigns run in what order and why.

Nobody proceeds past this checkpoint until the campaign hypothesis is genuinely understood and approved. Skipping this step to move faster is the single most expensive shortcut in outbound, because everything downstream inherits whatever mistake it contains, at volume.

Phase 2: market mapping, or sizing the market honestly

Once the diagnosis is done, the next question is simple to ask and genuinely easy to get wrong: how big is this market, actually?

Get this number wrong and everything downstream breaks. Treat a market of eight hundred accounts like it has eighty thousand, and you will burn through your entire addressable list in six weeks running high-volume plays that needed a small, patient, relationship-led approach instead. Treat a market of eighty thousand like it is eight hundred, and you will leave most of your pipeline on the table running one careful campaign at a time.

We estimate total addressable market two ways and reconcile them, because disagreement between the two methods is signal, not failure.

The top-down method starts with the total number of companies in the target vertical and geography, then narrows using a realistic size distribution. Most markets are heavily weighted toward small companies, so a size band like 50 to 500 employees is usually a much smaller slice of the total than intuition suggests.

The bottom-up method reasons from a countable proxy, almost always a filtered search on a professional network using industry, headcount and geography filters. It is less theoretical and it is checkable in minutes by anyone on the team, which matters more than precision.

From the account number we derive a contact number: qualifying accounts, multiplied by the average number of matching personas inside each account, multiplied by a reachability discount for the people you can actually find and message. That discount is real. Somewhere between half and four fifths of your theoretical contacts are the ones you will actually be able to locate and reach, and pretending otherwise just inflates a number nobody can act on.

The output that matters most is not the number. It is the tier, because the tier determines the entire strategy that follows.

Market sizeAccountsWhat it changes
SmallUnder 1,000Contact every 6 to 12 months. Evergreen and signal-based plays only. High-volume campaigns will exhaust the list before it can regenerate. Account-based selling is viable if deal size is high enough to justify it.
Medium1,000 to 10,000Contact every 3 to 6 months. A balanced mix of everything below works.
Large10,000 and upContact every 3 months. The full playbook is available, and whatever wins first gets scaled hardest.

This single classification is the reason two companies who sell superficially similar products need completely different outbound strategies. A vendor selling into a niche of six hundred accounts and a vendor selling into a market of sixty thousand accounts cannot run the same playbook, even if their product, price point and sales cycle look identical on paper.

Once the tier is set, we source accounts from every relevant data source, apply an elimination pass to strip anything obviously irrelevant before spending enrichment budget on it, then qualify what is left against the ICP criteria from the diagnostic. Qualification is never a single filter. It is a funnel: cheap signals first, expensive ones only on what survives. Scrape the website, check basic keywords, then only run the expensive AI classification and deep enrichment on companies that already look like plausible fits. We only pull people data after the account has qualified, never before, because enriching contacts at companies that were never going to be a fit is the fastest way to burn a data budget on nothing.

Nothing moves to campaign building until the data sources are locked and the segmentation approach is decided. That is the second checkpoint.

Phase 3: campaign building, the playbook that is not one playbook

Here is the part most agencies get backwards. They pick one channel, one message, one audience, and run it until it stops working. We run a portfolio, deliberately mixed across three types of campaign that trade off effort, predictability and ceiling.

The three campaign types

  • Evergreen campaigns are the universal, reliable plays that work for nearly every B2B company, refreshed every three to six months: first-degree network connections, warm intros through customers, event-based outreach, re-engaging closed-lost deals, thought-leadership engagers, lookalike audiences built from your best customers, and straightforward pain-based messaging. Low effort, high predictability. This is where you start, because it validates your ICP and messaging fast, with real, if modest, results.
  • Always-on campaigns are set-and-forget systems that run continuously once configured: engaging your own page followers, monitoring competitor page followers, reaching people who viewed your profile or your pricing page, tracking website visitors, following up on content downloads, and watching for customer champions who change jobs. That last one only works once you have enough happy customers for the signal to be meaningful. Low maintenance after setup, and it becomes the steady baseline flow of opportunity underneath everything else.
  • Signal-led campaigns are the experimental, high-ceiling, high-effort plays: reaching out on hiring spikes for relevant roles, new leadership appointments, funding announcements, product launches, technology stack changes, partnership announcements, and other real-world triggers that indicate timing. These are unpredictable and take real build time, and we deliberately do not commit to fully custom, one-off signal experiments in the first month of a new engagement. You earn the right to experiment once the fundamentals are proven.

The relevance ladder

The deeper idea underneath the taxonomy is what we call the relevance ladder, and it is the closest thing we have to a unifying theory of why some outreach lands and most does not.

The first, warmest layer is an existing connection: you are already linked on a professional network, they are a customer, they have visited your site. The second layer is a pain signal: they follow a competitor, they used to be a customer of someone you replaced, they are hiring for a role that implies the problem you solve, they follow the right thought leaders. The leads that do not fit either layer are where genuinely custom signal work comes in.

The insight that actually moves the needle is combining layers. A lead who is connected to your team, follows your page, and happens to be evaluating a specific tool category is not just in the ICP, they are showing you exact timing. That combination does not always need to change what you say to them. It should change how you reach them: manual outreach, a phone call, something that matches the urgency the signal is telling you about, instead of dropping them into the same automated sequence as everyone else.

How we decide what to build first

Every candidate campaign gets scored on four criteria: how fast it can be built to a quality bar, how well it fits this specific offer and audience, how differentiated it is from what the prospect sees from everyone else, and how many leads it can produce and for how long. Three variables shape the actual roadmap on top of that score.

  • Market size, from Phase 2, sets the volume ceiling as described above.
  • Deal size and sales cycle shape which plays make sense at all. A high-value, long-cycle sale calls for relationship-driven plays: first-degree connections, thought leadership, account-based approaches, patient timing. Volume-oriented tactics like hiring triggers or broad competitor-follower plays are usually a mismatch here, they read as impersonal on a deal that needs trust. A high-value, short-cycle sale rewards warm channels and precise timing signals. A lower-value deal, whatever the cycle length, needs volume to work at all: broader signal plays, pain-based messaging, vertical campaigns, and it is usually the wrong economics for account-based or fully custom signal work.
  • The size and quality of the founder's or sales team's own network determines how much weight the first-degree play should carry. A founder with a large, relevant network should treat first-degree outreach as the primary engine, not a nice-to-have add-on, because a warm introduction converts at a completely different rate than a cold one and the leverage is sitting right there unused in most companies.

The default starting sequence, absent a specific reason to deviate: first-degree connections launch first, because they reliably produce early meetings and build trust in the process while everything else ramps. One always-on campaign gets configured in parallel to start the steady baseline flow. One clear quick win, usually a closed-lost re-engagement or a sharp pain-based campaign, runs alongside it to build momentum. Signal-led campaigns get planned based on the market size and the patterns visible in the CRM, and they launch once the fundamentals are proven, not before.

Nothing launches until three things are true: the accurate market size is documented, the qualification logic is tested and finalized, and the copy is approved internally and by the client. That is the third checkpoint, and it exists because the cost of a bad launch compounds. Every send under a damaged reputation makes the next campaign harder, which is why the next section matters as much as anything above it.

Phase 4: execution, the unglamorous machinery that determines outcomes

This is the part that does not photograph well for a case study slide, and it is the part that decides whether any of the above ever reaches an inbox at all.

Infrastructure: the math is fixed, do not argue with it

Sender reputation is a one-way door. Once an inbox or a domain gets flagged, recovery takes four to six weeks if it happens at all, and often it just does not. That single fact is why infrastructure discipline is not optional, it is structural, and treating it as a creative decision instead of an operational one is where most in-house outbound programs quietly die.

We run a deliberately diversified stack rather than betting on one provider: our own infrastructure for Microsoft inboxes, paired with established providers for Google Workspace and for custom SMTP setups. The mix matters because of a deliverability rule that has held up for years: same-provider delivery tends to perform better, so sending from a mix of Microsoft and Google inboxes, matched against a mixed list of Microsoft and Google recipients, consistently outperforms putting everything on one provider.

The factor that gets missed most often, and the one worth paying the most attention to, is where the sending IP is physically based, relative to where you are sending. A European IP sending into European inboxes performs well. A US IP sending into US inboxes performs well. A European IP sending heavily into the US gets harder, and the same holds in reverse. An IP based somewhere with a poor sending reputation as a region, regardless of how clean your setup otherwise is, performs badly no matter who you are sending to. Match your IP geography to your target geography first, before you touch anything else about the setup.

For every one cold email a mailbox sends, it needs roughly 1.75 warmup emails running alongside it. Drop that ratio and reputation drops with it. This is the baseline invariant everything else is derived from.

A mailbox at full ramp sends around 15 cold emails a day, alongside 25 to 26 warmup emails, roughly 40 total. Each domain carries a maximum of three mailboxes, never more, because reputation lives at the domain level, and spreading risk across domains is the only durable way to scale. A domain never carries the company's real, transactional email; sending domains are separate, built specifically to absorb the risk that cold outreach carries.

Domain age matters more than warmup duration. Buy your domains and let them sit for seven to fourteen days before you do anything with them. That single step does more for deliverability than an extra two or three weeks of formal warmup on a fresh domain. Once a domain has aged, do not wait for a clean multi-week warmup cycle before sending anything real. Start sending very low volumes immediately, something like one to five emails a day on both the Microsoft and Gmail side, alongside the warmup traffic. Low-volume real sending from day one simulates genuine human behavior better than warmup traffic alone, and it gives you an early, honest read on whether the setup is working before you have committed real list volume to it. Ramp from there based on what you actually observe, not a fixed calendar.

Reply volume speeds up warmup more than almost anything else. A domain that is generating thirty to fifty replies a day reads as a domain with real, wanted activity, and it warms up noticeably faster as a result, especially where a meaningful share of those replies are positive and produce genuine back-and-forth. This cuts the other way too: replying to negative replies, rather than letting them sit unanswered, is itself a useful engagement signal and worth doing even when there is no deal to save.

Scaling volume is just multiplication once you know the ratio: fifty sends a day needs roughly four mailboxes across two domains; a hundred a day needs about seven mailboxes across three domains; two hundred a day, which is close to the practical ceiling for a single well-run setup, needs around fourteen or fifteen mailboxes spread across five domains, with something like 350 warmup emails a day running underneath it.

Deliverability: ten rules, learned the expensive way

Every one of these came out of a real campaign that broke, not a best-practices article. Treat them as fixed constraints, not suggestions.

  • Rotate sender domains rather than hammering one indefinitely.
  • Send plain text only on the first touch: no HTML signature, no images, no tracking pixels. Save any formatting for the third touch onward.
  • Suppress every account on your do-not-contact list, and suppress anyone your qualification process has already flagged as not callable.
  • Enforce a rolling cooldown window, typically thirty days, across every active campaign for a given account. A single prospect should never be inside two campaigns at once.
  • Cap every sequence at three touches, full stop. A fourth or fifth touch is disproportionately where stop contacting me replies come from, and every extra touch degrades the reputation your next campaign depends on.
  • Never send a broken personalization token. Either fall back to a generic-but-warm greeting or skip the row entirely.
  • Retire subject lines the moment they start reading as templated rather than personal. Prospects pattern-match on this faster than most teams realize.
  • Run a deliverability audit before every launch and check inbox placement weekly once live, with a target of at least 80 percent landing in the primary inbox on the major providers.
  • Watch spam complaint rate obsessively: keep it under a tenth of a percent, and pause immediately if it crosses half that.

Copy: rules that came from data, not taste

No bullet points in the first touch. Bullets read as a template the moment a human sees them; if you use them at all, save them for the third touch. Keep every touch under roughly eighty words, it is the single strongest predictor of reply quality we have measured. Sign off with a first name, not a title; authority does not come from a signature block. Keep subject lines under forty characters.

And always end the first touch with a question: in our own data, a question at the end of the first email produces close to a 49 percent positive reply rate against roughly 31 percent for emails that do not end on one. Use the exact language your buyer's sub-segment actually uses to describe themselves and their work, never the generic category label your own team defaults to internally.

Reply handling: where most of the actual revenue lives

A campaign that gets replies and then handles them badly is worse than a campaign that gets no replies at all, because it burns the goodwill of everyone who took the time to respond. We treat this stage as seriously as the campaign build itself, arguably more so.

The internal target is a response within three minutes during working hours, with ten minutes as a hard ceiling. If that target cannot be hit, the rule is to flag it immediately to the team rather than let a conversation go cold, because conversations genuinely get lost to slow replies.

Every reply gets one of three tags. Positive, where the goal is always to lock in a specific date and time directly in the conversation rather than just sending a booking link, because a directly confirmed meeting gives you far more control than one someone books unsupervised on their own schedule. Grey zone, anything that needs a clarifying follow-up, which gets a task created and a human follow-up rather than being left to auto-resolve. Negative, which splits three ways: genuinely not interested, in which case the contact goes back into a future nurture list rather than being discarded; explicit do-not-contact, which goes onto a permanent suppression list immediately; and wrong person, which is treated as an opportunity rather than a dead end, the reply is used to ask for the right contact and that new person enters a referral flow.

The mindset we train into this is the one thing worth repeating: sales starts with no. A no is a starting point for understanding why, not an ending. Even a hard not interested is worth a short, respectful follow-up asking what would need to be true for the answer to be different, because that answer is some of the best product and positioning feedback a company will ever get, and it is sitting for free in a reply nobody read carefully.

We also run a phone layer on top of replies wherever a phone number is available, because a proactive call to someone who already engaged converts at a materially higher rate than waiting for them to reply to a second email. The single best predictor of a strong month is not cleverer copy, it is how fast and how completely the team works every reply that already came in.

Cold calling: a five-part structure, not a script to memorize

Where a calling layer exists, whether cold or as a follow-up on replies, we run a consistent five-part structure rather than a memorized script, because the goal is a real conversation, not a recitation.

An opener that confirms you have the right person and their actual role, and gracefully redirects if you do not. A short framing of the three frustrations people in that role commonly report, delivered as an observation rather than a pitch, ending on a question that invites the prospect to either agree or push back. Genuine curiosity about whichever pain they engaged with: how long it has been a problem, what they have already tried to fix it, what it is actually costing them, and how it makes them feel, because staying in that emotional register rather than jumping straight to features is what surfaces the real motivation to act. And a close that is honest about uncertainty rather than falsely confident: an offer to explore whether you can help, not a guarantee that you can.

Every call has three legitimately good outcomes, and only one of them is booking a meeting: you book a meeting, you learn something that will help you book one later, or you disqualify quickly and save everyone real time. Treating the second and third outcomes as failures is why a lot of cold calling programs burn out the people running them for no good reason.

Phase 5: reporting and iteration

A campaign that launches and then runs unmonitored is not a system, it is a bet. Every active campaign gets a weekly review covering what is live, how much runway is left in the underlying list before it needs a refresh, reply and positive reply rate broken out by channel, meetings booked, and how quickly the client side is following up on positive replies, because a slow response on a warm lead wastes the hardest part of the work.

We separate two kinds of pipeline credit, because conflating them produces a distorted picture of what is actually working. Attributed pipeline is a direct reply that converted through the channel it came in on. Influenced pipeline is a prospect we contacted who later came in through a different channel entirely. A channel like this does not get credit for that meeting unless you are tracking it deliberately. Reporting both separately, rather than only the more flattering one, is what makes the reporting trustworthy enough to actually act on.

The decision rules that follow from the weekly numbers are simple by design: if reply rate on a campaign drops below its usual range, pause it and review copy and targeting before scaling anything further, do not just push more volume through a leak. If positive reply rate is strong, that is the signal to increase volume or expand the audience, not a coincidence to admire quietly. If a market is running out of fresh addressable accounts, that is the trigger to open a new segment or source new lists rather than re-contacting the same people too soon. And when a campaign proves itself over time, it graduates from an experiment into a standing part of the evergreen rotation, run again on a fixed few-month cycle rather than reinvented from scratch each time.

What actually breaks campaigns: a composite case

Here is a pattern we have seen enough times that it is worth describing plainly, built from more than one real engagement so no single client is identifiable in it.

A campaign had been running well over a year on the same sending domains, against a fairly narrow, repeatedly re-contacted list. Reply volume had collapsed by more than 80 percent over that period while the genuine positive rate among the replies that did arrive stayed roughly flat. Nothing about the offer had gotten worse. Five things had happened at once: the sending domains had absorbed eighteen months of continuous cold volume with no rotation and reputation had quietly eroded; the core template had been reused so many times it now read as boilerplate to a list that had mostly seen it before; a seasonal hook that made sense for a few months had kept running year-round long after it stopped being relevant; and the campaign had drifted from targeted, signal-anchored outreach into a broad, signal-less list that violated the exact targeting discipline described above.

The rebuild was not a copy refresh. It was structural: fresh sending domains warmed properly from zero, a hard three-touch cap enforced without exception, a thirty-day cooldown applied across every active campaign rather than within just one, and the list rebuilt entirely around specific, current signals instead of a static audience contacted repeatedly out of habit. Reply quality recovered well before volume did, because quality follows targeting discipline and volume follows domain reputation, and reputation takes longer to earn back than it took to lose.

The lesson generalizes past this one case: a campaign that stops working is almost never a copy problem first. Check the infrastructure and the targeting before you touch a word of the message.

The system, in one paragraph

Diagnose the business and the buyer honestly before building anything. Size the market with two independent methods and let the tier it produces set the entire strategy, because contact frequency and campaign mix both flow directly from that one number. Build a genuine portfolio across evergreen, always-on and signal-led plays rather than betting everything on one channel, starting with quick wins and earning the right to run expensive experiments. Treat sending infrastructure and deliverability as fixed engineering constraints, not creative choices, because reputation only moves in one direction once it is damaged. Answer every reply like the person on the other end deserves a fast, honest, human response, because that is where a disproportionate share of the actual revenue lives. And review the numbers every week, honestly enough to pause what is not working and scale what is, before assuming the market is the problem.

None of this is complicated in the sense of being hard to understand. It is complicated in the sense of requiring real discipline, held consistently, for longer than most companies are willing to hold it. That is the actual edge. Not a secret tactic, just a system run properly for long enough to compound.

Want this running for your team?

Book a 20-min call