Most pilots are designed so that they cannot fail. Which is the same thing as designing them so that they cannot inform.
Three hundred emails over four weeks, and then a meeting about whether it worked. At a two percent reply rate that pilot produces six replies. At three percent it produces nine. Nobody in that meeting can tell the difference between a good campaign and a bad one, because three replies is not a difference, it is noise. So the decision gets made on tone of voice and how the calls felt, which is exactly the decision you were trying to avoid making.
TL;DR
A lead generation pilot is a measurement exercise, and most are sized too small to measure anything. Using the standard sample size formula published by NIST, telling a two percent reply rate apart from a four percent one takes roughly 483 sends. Telling two percent from three percent takes about 1,747. Telling two percent from two and a half takes about 6,587. A 300-email pilot can only detect enormous differences, so it will almost always come back inconclusive and get read as a failure. Length has the same problem in reverse: Google's own sender guidance tells you to "Start with a low sending volume to engaged users, and slowly increase the volume over time", so the first weeks of any new sending setup measure your warm-up rather than your offer. Size the pilot to the difference you actually need to detect, run it long enough to clear the ramp, and agree the decision rule in writing before the first email goes out.
The pilot that cannot fail is the pilot that cannot inform
A pilot is a measurement, not a trial run. Everyone treats it as a low-commitment way to see whether an agency or a channel is any good. It is really an attempt to estimate one number, your reply or meeting rate, precisely enough to make a spending decision worth five or six figures.
Which means it has a required size, and the size is not negotiable by wanting it smaller. If you want to know whether your reply rate is two percent or four percent, there is an amount of sending below which you cannot know. Running less than that does not give you a weaker answer. It gives you a number that is indistinguishable from chance.
The failure mode is not a bad result. It is an ambiguous one. An ambiguous pilot gets interpreted, and interpretation is where the person who wanted to buy sees promise and the person who did not sees waste. Both are reading the same six replies.
So the real design question is: what difference do I need to detect? Not "how many emails can I afford". Work backwards from the decision. If you would sign a retainer at a three percent reply rate and walk away at two, then three versus two is the difference you must be able to see, and that sets everything else.
Size it against the difference you need to see
The formula is public and it is not complicated. NIST's Engineering Statistics Handbook publishes the sample size needed to test whether a proportion differs from an assumed value by a detectable amount. For a two-sided test at the conventional 95 percent confidence and 80 percent power, against a baseline of two percent, it produces the following.
To detect two percent versus six percent, about 141 sends. A tripling is easy to see. If your pilot only needs to answer "is this catastrophically broken", a small pilot is fine.
To detect two percent versus four percent, about 483 sends. A doubling. This is the most common real threshold, and it is already well above what most pilots run.
To detect two percent versus three percent, about 1,747 sends. A fifty percent improvement. This is the difference that usually decides whether a channel is worth funding, and it needs roughly six times the volume of a typical pilot.
To detect two percent versus two and a half, about 6,587 sends. At this point you are no longer running a pilot. You are running the campaign.
Read the ladder rather than the individual numbers. Halving the difference you want to detect roughly quadruples the sending you need. That relationship is why "let's start small and see" quietly fails: small pilots can only find effects so large that you would have noticed them without measuring.
And if 1,747 sends is not affordable, that is useful information too. It means you cannot answer the question you asked, so you should either change the question to one a smaller pilot can answer, or accept that the first phase is an operational test rather than a performance test, and say so out loud. Our piece on cold email benchmarks covers what rates people actually report and why you should not plan against them.
The first weeks measure your warm-up, not your offer
Google tells you to ramp, in its own sender guidelines. The published advice is to "Start with a low sending volume to engaged users, and slowly increase the volume over time", and to "Send email at a consistent rate. Avoid sending email in bursts." A new domain and a new mailbox going straight to full volume is the pattern those guidelines are written against.
So a four week pilot on new infrastructure spends most of itself climbing. Whatever reply rate you see in weeks one and two is a measurement of your deliverability ramp, not of your message. If you then average the whole period, you have deliberately contaminated the number you are paying to learn.
You also cannot see the deliverability data at that volume. Google's own Postmaster Tools documentation states that "Data might be missing if the total number of messages for a given day is too low. This is to protect users' privacy." A small pilot therefore sits in a gap where you can neither read your reputation signals nor trust your reply rate. Our cold email deliverability guide covers the setup that has to be right before any of this is worth measuring.
Which is why the pilot needs a stated ramp and a stated measurement window. Weeks one and two are the ramp. Week three is where list quality shows up, in bounces and in silence from addresses that should have replied. Weeks four onward are the measurement period, and they are the only weeks whose numbers go into the decision. Write that split down before you start, because after the fact everyone will want to include or exclude the early weeks depending on which way the numbers went.
The volume thresholds also mean the pilot is not the same compliance surface as the campaign. Google's bulk sender requirements apply to senders of "more than 5,000 messages per day to Gmail accounts", so a pilot below that line is not exercising the requirements your scaled programme will have to meet. Plan for them in the pilot even though they do not yet bite.
The unsubscribe and identification duties apply from the first email, though. The UK regulator's electronic mail marketing guidance states that "You can send unsolicited electronic mail marketing to corporate subscribers without consent or a soft opt-in", and in the same guidance that "You must not disguise or hide your identity in messages to either type of subscriber. You must provide a valid contact address for recipients to opt out or unsubscribe." A pilot is not a period during which those stop applying. This is not legal advice, and the position differs by jurisdiction.
![[SCREENSHOT NEEDED: Google Postmaster Tools showing a sparse or empty chart, illustrating the low-volume data gap]](https://framerusercontent.com/images/xWPcY6nhFv3FTXyOY90ti5dXd50.webp)
Agree the decision rule before the first email
Write down the number that means yes. Not "we will see how it goes". A specific rate, on a specific metric, measured over a specific window. If you cannot name it before you start, you will name it afterwards, and you will name it to match whatever happened.
Write down the number that means no, separately. These are not the same threshold with a different sign. There is usually a middle band where the honest answer is "extend and re-measure", and deciding in advance that the band exists prevents an inconclusive result from being argued into a conclusive one.
Name who owns the list. A pilot that produces a targeting list, verified contacts and a working sequence has produced an asset. Whether that asset stays with you if you do not continue is a contract question, not a goodwill question, and it is much easier to ask now.
Name what happens to the domains and mailboxes. If the pilot ran on infrastructure the agency owns, then the warmed sending reputation you paid several weeks to build does not come with you. If it ran on yours, you keep it and you also keep the risk.
Agree the reporting cadence and the raw data. Weekly is right. Ask for sends, deliveries, opens if they are tracked at all, replies split into positive, neutral and negative, and meetings booked, as counts rather than percentages. Percentages on small denominators are the mechanism by which a pilot gets oversold.
Agree who writes the copy and who approves it. The most common reason a pilot underperforms is three rounds of internal approval sanding the message down to something nobody could object to, which is also something nobody replies to.
Fix the offer for the duration. Changing the message halfway through a pilot means you have run two underpowered tests instead of one adequately powered one, which is strictly worse than either.
The list matters more than the copy at this size
At pilot volumes, targeting error dominates everything else. A thousand well-chosen contacts and a mediocre email will beat five thousand loosely chosen ones and a good email, because the reply rate you are measuring is a property of the pairing, not of the message alone.
So build the list before you write anything. If the pilot is a test of whether this segment responds, the segment has to be defined tightly enough that the result generalises to the segment you would scale into. A pilot list assembled from whoever was easiest to find tells you about ease of finding, and nothing else. Our ICP framework is the version of this we use.
And verify it, because bounces come out of your denominator and your reputation at the same time. A ten percent bounce rate on a 500-send pilot removes fifty sends from an already thin sample and damages the sending reputation you will need for the real campaign.
One segment, not four. Splitting a small pilot across multiple industries or personas produces four samples that are each too small to read, and then a comparison between them that is entirely noise. If you want to compare segments, that is a separate and much larger exercise.
What a pilot genuinely cannot tell you
Whether the channel works for your business over a full sales cycle. A pilot measures the top of the funnel. If your cycle is six months, the pilot cannot tell you about close rates, and any projection to revenue is arithmetic performed on an assumption.
Whether this agency is good. It tells you whether this agency, on this list, with this offer, in this window, produced this many replies. Operational quality shows up over months, in how they handle the weeks that go badly. Our piece on the first 90 days with an outbound agency covers what that period actually looks like.
Whether the offer is right. A pilot tests one offer. A poor result is at least as likely to mean the offer was wrong as the channel was, and those two conclusions lead to opposite decisions.
What it will cost at scale. Per-meeting cost in a pilot is calculated on a handful of meetings and is close to meaningless as a unit economic. Treat it as a range, widely.
What we do not publish here
Benchmark reply rates, meeting rates or cost per meeting. Ours come from a specific set of clients, offers and markets, and publishing them as though they were general would be exactly the kind of number this article argues against planning around.
A recommended pilot budget. It follows from the volume you need, which follows from the difference you need to detect, which is specific to your decision.
A recommended pilot length in weeks. It depends on how quickly your infrastructure can ramp and how many sends per day the list supports without burning it.
Any claim about how inbox providers decide placement. Neither Google nor Microsoft publishes its filtering mechanics, and everything in circulation on the subject is inference.
Legal advice. The sender requirements referenced here are technical, not legal, and the rules on contacting people without prior consent differ by jurisdiction. Take advice on your own markets.
FAQ
How many emails should a lead generation pilot send?
Enough to detect the difference that changes your decision. Against a two percent baseline, distinguishing two from four percent takes roughly 483 sends, two from three percent roughly 1,747, and two from two and a half roughly 6,587, using the standard sample size formula for a proportion at 95 percent confidence and 80 percent power. Decide which of those questions you are asking first.
How long should a lead generation pilot run?
Long enough that the measurement window sits after the sending ramp, not across it. Google's guidance is to start at low volume and increase slowly, so on new infrastructure the first two weeks are warm-up and should be excluded from the decision by agreement, not by argument afterwards.
What is a good reply rate for a pilot?
We do not publish one, and you should be careful with anyone who does without saying which market, offer and list it came from. The more useful question is what rate would make you sign, because that is the number the pilot has to be sized to detect.
Should a pilot test more than one segment?
No, not at pilot volumes. Splitting a small sample across segments produces several samples too small to read individually and a comparison between them that is noise. Pick the segment you would scale into and test that one properly.
Who should own the domains a pilot runs on?
Decide it in writing before the pilot starts. Sending reputation takes weeks to build and does not transfer, so if the pilot runs on infrastructure you do not own, you start again from zero if you continue with someone else.
What should you do if the pilot result is inconclusive?
Extend and re-measure rather than interpret, which is why the middle band belongs in the agreement from the start. An inconclusive result is a sample size problem, and the only thing that fixes a sample size problem is more sample.
Bottom line
A pilot is not a smaller version of the campaign. It is a measurement exercise with a required size, and if you run it below that size you have bought an anecdote at campaign prices. Work out which difference in reply rate would change your decision, look up how much sending that difference needs, and either commit to that volume or change the question. Then put the ramp period, the measurement window, the yes number, the no number, the middle band, the list ownership and the domain ownership in writing before the first email leaves. None of that costs anything, and all of it prevents the meeting where six replies get argued about for an hour.
Want a pilot sized to answer the question rather than to look affordable? Book a call with GROU. We run lead generation and outbound inside B2B revenue engines across verticals.
We are GROU, a B2B pipeline agency that runs lead generation, outbound, and LinkedIn content for clients across manufacturing, fintech, iGaming, software, and professional services. The sample size figures in this article are computed by us from the published NIST formula and shown with their inputs so you can check them. Nothing here is legal advice.
Most pilots are designed so that they cannot fail. Which is the same thing as designing them so that they cannot inform.
Three hundred emails over four weeks, and then a meeting about whether it worked. At a two percent reply rate that pilot produces six replies. At three percent it produces nine. Nobody in that meeting can tell the difference between a good campaign and a bad one, because three replies is not a difference, it is noise. So the decision gets made on tone of voice and how the calls felt, which is exactly the decision you were trying to avoid making.
TL;DR
A lead generation pilot is a measurement exercise, and most are sized too small to measure anything. Using the standard sample size formula published by NIST, telling a two percent reply rate apart from a four percent one takes roughly 483 sends. Telling two percent from three percent takes about 1,747. Telling two percent from two and a half takes about 6,587. A 300-email pilot can only detect enormous differences, so it will almost always come back inconclusive and get read as a failure. Length has the same problem in reverse: Google's own sender guidance tells you to "Start with a low sending volume to engaged users, and slowly increase the volume over time", so the first weeks of any new sending setup measure your warm-up rather than your offer. Size the pilot to the difference you actually need to detect, run it long enough to clear the ramp, and agree the decision rule in writing before the first email goes out.
The pilot that cannot fail is the pilot that cannot inform
A pilot is a measurement, not a trial run. Everyone treats it as a low-commitment way to see whether an agency or a channel is any good. It is really an attempt to estimate one number, your reply or meeting rate, precisely enough to make a spending decision worth five or six figures.
Which means it has a required size, and the size is not negotiable by wanting it smaller. If you want to know whether your reply rate is two percent or four percent, there is an amount of sending below which you cannot know. Running less than that does not give you a weaker answer. It gives you a number that is indistinguishable from chance.
The failure mode is not a bad result. It is an ambiguous one. An ambiguous pilot gets interpreted, and interpretation is where the person who wanted to buy sees promise and the person who did not sees waste. Both are reading the same six replies.
So the real design question is: what difference do I need to detect? Not "how many emails can I afford". Work backwards from the decision. If you would sign a retainer at a three percent reply rate and walk away at two, then three versus two is the difference you must be able to see, and that sets everything else.
Size it against the difference you need to see
The formula is public and it is not complicated. NIST's Engineering Statistics Handbook publishes the sample size needed to test whether a proportion differs from an assumed value by a detectable amount. For a two-sided test at the conventional 95 percent confidence and 80 percent power, against a baseline of two percent, it produces the following.
To detect two percent versus six percent, about 141 sends. A tripling is easy to see. If your pilot only needs to answer "is this catastrophically broken", a small pilot is fine.
To detect two percent versus four percent, about 483 sends. A doubling. This is the most common real threshold, and it is already well above what most pilots run.
To detect two percent versus three percent, about 1,747 sends. A fifty percent improvement. This is the difference that usually decides whether a channel is worth funding, and it needs roughly six times the volume of a typical pilot.
To detect two percent versus two and a half, about 6,587 sends. At this point you are no longer running a pilot. You are running the campaign.
Read the ladder rather than the individual numbers. Halving the difference you want to detect roughly quadruples the sending you need. That relationship is why "let's start small and see" quietly fails: small pilots can only find effects so large that you would have noticed them without measuring.
And if 1,747 sends is not affordable, that is useful information too. It means you cannot answer the question you asked, so you should either change the question to one a smaller pilot can answer, or accept that the first phase is an operational test rather than a performance test, and say so out loud. Our piece on cold email benchmarks covers what rates people actually report and why you should not plan against them.
The first weeks measure your warm-up, not your offer
Google tells you to ramp, in its own sender guidelines. The published advice is to "Start with a low sending volume to engaged users, and slowly increase the volume over time", and to "Send email at a consistent rate. Avoid sending email in bursts." A new domain and a new mailbox going straight to full volume is the pattern those guidelines are written against.
So a four week pilot on new infrastructure spends most of itself climbing. Whatever reply rate you see in weeks one and two is a measurement of your deliverability ramp, not of your message. If you then average the whole period, you have deliberately contaminated the number you are paying to learn.
You also cannot see the deliverability data at that volume. Google's own Postmaster Tools documentation states that "Data might be missing if the total number of messages for a given day is too low. This is to protect users' privacy." A small pilot therefore sits in a gap where you can neither read your reputation signals nor trust your reply rate. Our cold email deliverability guide covers the setup that has to be right before any of this is worth measuring.
Which is why the pilot needs a stated ramp and a stated measurement window. Weeks one and two are the ramp. Week three is where list quality shows up, in bounces and in silence from addresses that should have replied. Weeks four onward are the measurement period, and they are the only weeks whose numbers go into the decision. Write that split down before you start, because after the fact everyone will want to include or exclude the early weeks depending on which way the numbers went.
The volume thresholds also mean the pilot is not the same compliance surface as the campaign. Google's bulk sender requirements apply to senders of "more than 5,000 messages per day to Gmail accounts", so a pilot below that line is not exercising the requirements your scaled programme will have to meet. Plan for them in the pilot even though they do not yet bite.
The unsubscribe and identification duties apply from the first email, though. The UK regulator's electronic mail marketing guidance states that "You can send unsolicited electronic mail marketing to corporate subscribers without consent or a soft opt-in", and in the same guidance that "You must not disguise or hide your identity in messages to either type of subscriber. You must provide a valid contact address for recipients to opt out or unsubscribe." A pilot is not a period during which those stop applying. This is not legal advice, and the position differs by jurisdiction.
![[SCREENSHOT NEEDED: Google Postmaster Tools showing a sparse or empty chart, illustrating the low-volume data gap]](https://framerusercontent.com/images/xWPcY6nhFv3FTXyOY90ti5dXd50.webp)
Agree the decision rule before the first email
Write down the number that means yes. Not "we will see how it goes". A specific rate, on a specific metric, measured over a specific window. If you cannot name it before you start, you will name it afterwards, and you will name it to match whatever happened.
Write down the number that means no, separately. These are not the same threshold with a different sign. There is usually a middle band where the honest answer is "extend and re-measure", and deciding in advance that the band exists prevents an inconclusive result from being argued into a conclusive one.
Name who owns the list. A pilot that produces a targeting list, verified contacts and a working sequence has produced an asset. Whether that asset stays with you if you do not continue is a contract question, not a goodwill question, and it is much easier to ask now.
Name what happens to the domains and mailboxes. If the pilot ran on infrastructure the agency owns, then the warmed sending reputation you paid several weeks to build does not come with you. If it ran on yours, you keep it and you also keep the risk.
Agree the reporting cadence and the raw data. Weekly is right. Ask for sends, deliveries, opens if they are tracked at all, replies split into positive, neutral and negative, and meetings booked, as counts rather than percentages. Percentages on small denominators are the mechanism by which a pilot gets oversold.
Agree who writes the copy and who approves it. The most common reason a pilot underperforms is three rounds of internal approval sanding the message down to something nobody could object to, which is also something nobody replies to.
Fix the offer for the duration. Changing the message halfway through a pilot means you have run two underpowered tests instead of one adequately powered one, which is strictly worse than either.
The list matters more than the copy at this size
At pilot volumes, targeting error dominates everything else. A thousand well-chosen contacts and a mediocre email will beat five thousand loosely chosen ones and a good email, because the reply rate you are measuring is a property of the pairing, not of the message alone.
So build the list before you write anything. If the pilot is a test of whether this segment responds, the segment has to be defined tightly enough that the result generalises to the segment you would scale into. A pilot list assembled from whoever was easiest to find tells you about ease of finding, and nothing else. Our ICP framework is the version of this we use.
And verify it, because bounces come out of your denominator and your reputation at the same time. A ten percent bounce rate on a 500-send pilot removes fifty sends from an already thin sample and damages the sending reputation you will need for the real campaign.
One segment, not four. Splitting a small pilot across multiple industries or personas produces four samples that are each too small to read, and then a comparison between them that is entirely noise. If you want to compare segments, that is a separate and much larger exercise.
What a pilot genuinely cannot tell you
Whether the channel works for your business over a full sales cycle. A pilot measures the top of the funnel. If your cycle is six months, the pilot cannot tell you about close rates, and any projection to revenue is arithmetic performed on an assumption.
Whether this agency is good. It tells you whether this agency, on this list, with this offer, in this window, produced this many replies. Operational quality shows up over months, in how they handle the weeks that go badly. Our piece on the first 90 days with an outbound agency covers what that period actually looks like.
Whether the offer is right. A pilot tests one offer. A poor result is at least as likely to mean the offer was wrong as the channel was, and those two conclusions lead to opposite decisions.
What it will cost at scale. Per-meeting cost in a pilot is calculated on a handful of meetings and is close to meaningless as a unit economic. Treat it as a range, widely.
What we do not publish here
Benchmark reply rates, meeting rates or cost per meeting. Ours come from a specific set of clients, offers and markets, and publishing them as though they were general would be exactly the kind of number this article argues against planning around.
A recommended pilot budget. It follows from the volume you need, which follows from the difference you need to detect, which is specific to your decision.
A recommended pilot length in weeks. It depends on how quickly your infrastructure can ramp and how many sends per day the list supports without burning it.
Any claim about how inbox providers decide placement. Neither Google nor Microsoft publishes its filtering mechanics, and everything in circulation on the subject is inference.
Legal advice. The sender requirements referenced here are technical, not legal, and the rules on contacting people without prior consent differ by jurisdiction. Take advice on your own markets.
FAQ
How many emails should a lead generation pilot send?
Enough to detect the difference that changes your decision. Against a two percent baseline, distinguishing two from four percent takes roughly 483 sends, two from three percent roughly 1,747, and two from two and a half roughly 6,587, using the standard sample size formula for a proportion at 95 percent confidence and 80 percent power. Decide which of those questions you are asking first.
How long should a lead generation pilot run?
Long enough that the measurement window sits after the sending ramp, not across it. Google's guidance is to start at low volume and increase slowly, so on new infrastructure the first two weeks are warm-up and should be excluded from the decision by agreement, not by argument afterwards.
What is a good reply rate for a pilot?
We do not publish one, and you should be careful with anyone who does without saying which market, offer and list it came from. The more useful question is what rate would make you sign, because that is the number the pilot has to be sized to detect.
Should a pilot test more than one segment?
No, not at pilot volumes. Splitting a small sample across segments produces several samples too small to read individually and a comparison between them that is noise. Pick the segment you would scale into and test that one properly.
Who should own the domains a pilot runs on?
Decide it in writing before the pilot starts. Sending reputation takes weeks to build and does not transfer, so if the pilot runs on infrastructure you do not own, you start again from zero if you continue with someone else.
What should you do if the pilot result is inconclusive?
Extend and re-measure rather than interpret, which is why the middle band belongs in the agreement from the start. An inconclusive result is a sample size problem, and the only thing that fixes a sample size problem is more sample.
Bottom line
A pilot is not a smaller version of the campaign. It is a measurement exercise with a required size, and if you run it below that size you have bought an anecdote at campaign prices. Work out which difference in reply rate would change your decision, look up how much sending that difference needs, and either commit to that volume or change the question. Then put the ramp period, the measurement window, the yes number, the no number, the middle band, the list ownership and the domain ownership in writing before the first email leaves. None of that costs anything, and all of it prevents the meeting where six replies get argued about for an hour.
Want a pilot sized to answer the question rather than to look affordable? Book a call with GROU. We run lead generation and outbound inside B2B revenue engines across verticals.
We are GROU, a B2B pipeline agency that runs lead generation, outbound, and LinkedIn content for clients across manufacturing, fintech, iGaming, software, and professional services. The sample size figures in this article are computed by us from the published NIST formula and shown with their inputs so you can check them. Nothing here is legal advice.
Most pilots are designed so that they cannot fail. Which is the same thing as designing them so that they cannot inform.
Three hundred emails over four weeks, and then a meeting about whether it worked. At a two percent reply rate that pilot produces six replies. At three percent it produces nine. Nobody in that meeting can tell the difference between a good campaign and a bad one, because three replies is not a difference, it is noise. So the decision gets made on tone of voice and how the calls felt, which is exactly the decision you were trying to avoid making.
TL;DR
A lead generation pilot is a measurement exercise, and most are sized too small to measure anything. Using the standard sample size formula published by NIST, telling a two percent reply rate apart from a four percent one takes roughly 483 sends. Telling two percent from three percent takes about 1,747. Telling two percent from two and a half takes about 6,587. A 300-email pilot can only detect enormous differences, so it will almost always come back inconclusive and get read as a failure. Length has the same problem in reverse: Google's own sender guidance tells you to "Start with a low sending volume to engaged users, and slowly increase the volume over time", so the first weeks of any new sending setup measure your warm-up rather than your offer. Size the pilot to the difference you actually need to detect, run it long enough to clear the ramp, and agree the decision rule in writing before the first email goes out.
The pilot that cannot fail is the pilot that cannot inform
A pilot is a measurement, not a trial run. Everyone treats it as a low-commitment way to see whether an agency or a channel is any good. It is really an attempt to estimate one number, your reply or meeting rate, precisely enough to make a spending decision worth five or six figures.
Which means it has a required size, and the size is not negotiable by wanting it smaller. If you want to know whether your reply rate is two percent or four percent, there is an amount of sending below which you cannot know. Running less than that does not give you a weaker answer. It gives you a number that is indistinguishable from chance.
The failure mode is not a bad result. It is an ambiguous one. An ambiguous pilot gets interpreted, and interpretation is where the person who wanted to buy sees promise and the person who did not sees waste. Both are reading the same six replies.
So the real design question is: what difference do I need to detect? Not "how many emails can I afford". Work backwards from the decision. If you would sign a retainer at a three percent reply rate and walk away at two, then three versus two is the difference you must be able to see, and that sets everything else.
Size it against the difference you need to see
The formula is public and it is not complicated. NIST's Engineering Statistics Handbook publishes the sample size needed to test whether a proportion differs from an assumed value by a detectable amount. For a two-sided test at the conventional 95 percent confidence and 80 percent power, against a baseline of two percent, it produces the following.
To detect two percent versus six percent, about 141 sends. A tripling is easy to see. If your pilot only needs to answer "is this catastrophically broken", a small pilot is fine.
To detect two percent versus four percent, about 483 sends. A doubling. This is the most common real threshold, and it is already well above what most pilots run.
To detect two percent versus three percent, about 1,747 sends. A fifty percent improvement. This is the difference that usually decides whether a channel is worth funding, and it needs roughly six times the volume of a typical pilot.
To detect two percent versus two and a half, about 6,587 sends. At this point you are no longer running a pilot. You are running the campaign.
Read the ladder rather than the individual numbers. Halving the difference you want to detect roughly quadruples the sending you need. That relationship is why "let's start small and see" quietly fails: small pilots can only find effects so large that you would have noticed them without measuring.
And if 1,747 sends is not affordable, that is useful information too. It means you cannot answer the question you asked, so you should either change the question to one a smaller pilot can answer, or accept that the first phase is an operational test rather than a performance test, and say so out loud. Our piece on cold email benchmarks covers what rates people actually report and why you should not plan against them.
The first weeks measure your warm-up, not your offer
Google tells you to ramp, in its own sender guidelines. The published advice is to "Start with a low sending volume to engaged users, and slowly increase the volume over time", and to "Send email at a consistent rate. Avoid sending email in bursts." A new domain and a new mailbox going straight to full volume is the pattern those guidelines are written against.
So a four week pilot on new infrastructure spends most of itself climbing. Whatever reply rate you see in weeks one and two is a measurement of your deliverability ramp, not of your message. If you then average the whole period, you have deliberately contaminated the number you are paying to learn.
You also cannot see the deliverability data at that volume. Google's own Postmaster Tools documentation states that "Data might be missing if the total number of messages for a given day is too low. This is to protect users' privacy." A small pilot therefore sits in a gap where you can neither read your reputation signals nor trust your reply rate. Our cold email deliverability guide covers the setup that has to be right before any of this is worth measuring.
Which is why the pilot needs a stated ramp and a stated measurement window. Weeks one and two are the ramp. Week three is where list quality shows up, in bounces and in silence from addresses that should have replied. Weeks four onward are the measurement period, and they are the only weeks whose numbers go into the decision. Write that split down before you start, because after the fact everyone will want to include or exclude the early weeks depending on which way the numbers went.
The volume thresholds also mean the pilot is not the same compliance surface as the campaign. Google's bulk sender requirements apply to senders of "more than 5,000 messages per day to Gmail accounts", so a pilot below that line is not exercising the requirements your scaled programme will have to meet. Plan for them in the pilot even though they do not yet bite.
The unsubscribe and identification duties apply from the first email, though. The UK regulator's electronic mail marketing guidance states that "You can send unsolicited electronic mail marketing to corporate subscribers without consent or a soft opt-in", and in the same guidance that "You must not disguise or hide your identity in messages to either type of subscriber. You must provide a valid contact address for recipients to opt out or unsubscribe." A pilot is not a period during which those stop applying. This is not legal advice, and the position differs by jurisdiction.
![[SCREENSHOT NEEDED: Google Postmaster Tools showing a sparse or empty chart, illustrating the low-volume data gap]](https://framerusercontent.com/images/xWPcY6nhFv3FTXyOY90ti5dXd50.webp)
Agree the decision rule before the first email
Write down the number that means yes. Not "we will see how it goes". A specific rate, on a specific metric, measured over a specific window. If you cannot name it before you start, you will name it afterwards, and you will name it to match whatever happened.
Write down the number that means no, separately. These are not the same threshold with a different sign. There is usually a middle band where the honest answer is "extend and re-measure", and deciding in advance that the band exists prevents an inconclusive result from being argued into a conclusive one.
Name who owns the list. A pilot that produces a targeting list, verified contacts and a working sequence has produced an asset. Whether that asset stays with you if you do not continue is a contract question, not a goodwill question, and it is much easier to ask now.
Name what happens to the domains and mailboxes. If the pilot ran on infrastructure the agency owns, then the warmed sending reputation you paid several weeks to build does not come with you. If it ran on yours, you keep it and you also keep the risk.
Agree the reporting cadence and the raw data. Weekly is right. Ask for sends, deliveries, opens if they are tracked at all, replies split into positive, neutral and negative, and meetings booked, as counts rather than percentages. Percentages on small denominators are the mechanism by which a pilot gets oversold.
Agree who writes the copy and who approves it. The most common reason a pilot underperforms is three rounds of internal approval sanding the message down to something nobody could object to, which is also something nobody replies to.
Fix the offer for the duration. Changing the message halfway through a pilot means you have run two underpowered tests instead of one adequately powered one, which is strictly worse than either.
The list matters more than the copy at this size
At pilot volumes, targeting error dominates everything else. A thousand well-chosen contacts and a mediocre email will beat five thousand loosely chosen ones and a good email, because the reply rate you are measuring is a property of the pairing, not of the message alone.
So build the list before you write anything. If the pilot is a test of whether this segment responds, the segment has to be defined tightly enough that the result generalises to the segment you would scale into. A pilot list assembled from whoever was easiest to find tells you about ease of finding, and nothing else. Our ICP framework is the version of this we use.
And verify it, because bounces come out of your denominator and your reputation at the same time. A ten percent bounce rate on a 500-send pilot removes fifty sends from an already thin sample and damages the sending reputation you will need for the real campaign.
One segment, not four. Splitting a small pilot across multiple industries or personas produces four samples that are each too small to read, and then a comparison between them that is entirely noise. If you want to compare segments, that is a separate and much larger exercise.
What a pilot genuinely cannot tell you
Whether the channel works for your business over a full sales cycle. A pilot measures the top of the funnel. If your cycle is six months, the pilot cannot tell you about close rates, and any projection to revenue is arithmetic performed on an assumption.
Whether this agency is good. It tells you whether this agency, on this list, with this offer, in this window, produced this many replies. Operational quality shows up over months, in how they handle the weeks that go badly. Our piece on the first 90 days with an outbound agency covers what that period actually looks like.
Whether the offer is right. A pilot tests one offer. A poor result is at least as likely to mean the offer was wrong as the channel was, and those two conclusions lead to opposite decisions.
What it will cost at scale. Per-meeting cost in a pilot is calculated on a handful of meetings and is close to meaningless as a unit economic. Treat it as a range, widely.
What we do not publish here
Benchmark reply rates, meeting rates or cost per meeting. Ours come from a specific set of clients, offers and markets, and publishing them as though they were general would be exactly the kind of number this article argues against planning around.
A recommended pilot budget. It follows from the volume you need, which follows from the difference you need to detect, which is specific to your decision.
A recommended pilot length in weeks. It depends on how quickly your infrastructure can ramp and how many sends per day the list supports without burning it.
Any claim about how inbox providers decide placement. Neither Google nor Microsoft publishes its filtering mechanics, and everything in circulation on the subject is inference.
Legal advice. The sender requirements referenced here are technical, not legal, and the rules on contacting people without prior consent differ by jurisdiction. Take advice on your own markets.
FAQ
How many emails should a lead generation pilot send?
Enough to detect the difference that changes your decision. Against a two percent baseline, distinguishing two from four percent takes roughly 483 sends, two from three percent roughly 1,747, and two from two and a half roughly 6,587, using the standard sample size formula for a proportion at 95 percent confidence and 80 percent power. Decide which of those questions you are asking first.
How long should a lead generation pilot run?
Long enough that the measurement window sits after the sending ramp, not across it. Google's guidance is to start at low volume and increase slowly, so on new infrastructure the first two weeks are warm-up and should be excluded from the decision by agreement, not by argument afterwards.
What is a good reply rate for a pilot?
We do not publish one, and you should be careful with anyone who does without saying which market, offer and list it came from. The more useful question is what rate would make you sign, because that is the number the pilot has to be sized to detect.
Should a pilot test more than one segment?
No, not at pilot volumes. Splitting a small sample across segments produces several samples too small to read individually and a comparison between them that is noise. Pick the segment you would scale into and test that one properly.
Who should own the domains a pilot runs on?
Decide it in writing before the pilot starts. Sending reputation takes weeks to build and does not transfer, so if the pilot runs on infrastructure you do not own, you start again from zero if you continue with someone else.
What should you do if the pilot result is inconclusive?
Extend and re-measure rather than interpret, which is why the middle band belongs in the agreement from the start. An inconclusive result is a sample size problem, and the only thing that fixes a sample size problem is more sample.
Bottom line
A pilot is not a smaller version of the campaign. It is a measurement exercise with a required size, and if you run it below that size you have bought an anecdote at campaign prices. Work out which difference in reply rate would change your decision, look up how much sending that difference needs, and either commit to that volume or change the question. Then put the ramp period, the measurement window, the yes number, the no number, the middle band, the list ownership and the domain ownership in writing before the first email leaves. None of that costs anything, and all of it prevents the meeting where six replies get argued about for an hour.
Want a pilot sized to answer the question rather than to look affordable? Book a call with GROU. We run lead generation and outbound inside B2B revenue engines across verticals.
We are GROU, a B2B pipeline agency that runs lead generation, outbound, and LinkedIn content for clients across manufacturing, fintech, iGaming, software, and professional services. The sample size figures in this article are computed by us from the published NIST formula and shown with their inputs so you can check them. Nothing here is legal advice.
Pipeline OS Newsletter
Build qualified pipeline
Get weekly tactics to generate demand, improve lead quality, and book more meetings.






Trusted by industry leaders
Trusted by industry leaders
Trusted by industry leaders
Ready to build qualified pipeline?
Ready to build qualified pipeline?
Ready to build qualified pipeline?
Book a call to see if we're the right fit, or take the 2-minute quiz to get a clear starting point.
Book a call to see if we're the right fit, or take the 2-minute quiz to get a clear starting point.
Book a call to see if we're the right fit, or take the 2-minute quiz to get a clear starting point.
Copyright © 2026 – All Right Reserved
Copyright © 2026 – All Right Reserved
Copyright © 2026 – All Right Reserved



![Every comparison of cold email tools lines up the sticker prices and calls it a ranking. That is the one thing you should not do here, because the tools are not selling the same unit. Two of them charge per seat. Three charge per workspace with unlimited users. One does not price on emails at all. And across three independent vendors, the entry tier costs between five and twelve times more per email sent than the tier immediately above it. [INSERT HERO, hero-best-lemlist-alternatives.svg] Alt: Best Lemlist alternatives in 2026, compared on published prices normalised by email volume and by seat structure. TL;DR Lemlist lists an Email plan at $69 a month for 50,000 emails with unlimited users, and a Multichannel plan at $109 per user per month. That per user wording is the single most important thing on the page, because a team of five on Multichannel is $545 a month while every other tool here includes unlimited users at the same price. On volume, the entry tiers across the category are dramatically poor value: Instantly's Growth plan works out at roughly $9.40 per thousand emails, Smartlead's Base at $6.50 and Saleshandy's Starter at $6.00, against $1.38 for Lemlist's Email plan, $0.78 for Instantly Hypergrowth and $0.66 for Saleshandy Outreach Pro. Stepping up one tier typically multiplies your sending allowance by fifteen to twenty-five times for roughly two to three times the price. Woodpecker sits outside the comparison entirely, charging $7.00 per 100 contacted prospects rather than per email or per seat. So the honest question is not which tool is cheapest, it is how many people need logins and how many emails you actually send. The three things that decide this [INSERT CHART 1, best-lemlist-alternatives-chart-1-models.svg] Alt: How five cold email platforms price in 2026, comparing the billing unit, seat treatment and sending allowance. Seats. Lemlist's pricing page lists the Email plan with "Unlimited users" and the Multichannel plan at "$109" per user per month with "5 Senders /User". Instantly, Smartlead, Saleshandy and Woodpecker all advertise unlimited email accounts, and Woodpecker states unlimited team members free. Volume. Every tool caps monthly sends except Lemlist's Multichannel and Enterprise tiers, which state "Unlimited emails & messages/mo". The billing unit itself. Woodpecker charges for contacted prospects, not emails. If your sequences are long, that is dramatically in your favour. If they are short and your list is enormous, it is not. Everything else is a feature argument, and feature arguments in this category are decided by a two week trial rather than by an article. Lemlist, so you know what you are leaving Email plan at $69 a month. Includes "50,000 emails/mo", "Unlimited users" and "Unlimited Contacts", falling to "$55/month" on annual billing with a stated 20% discount, or 10% quarterly. Multichannel at $109 per user a month. Falls to "$87/month" annually. Includes "Unlimited emails & messages/mo" and "5 Senders /User". Enterprise is custom with five or more senders per user. A 14 day free trial with no card, and a credit system priced at "$10" for "1k credits", where a credit buys email verification at 5 credits per email and phone numbers at 20 credits each. Which makes the Email plan quietly one of the better deals here, at $1.38 per thousand emails with no per-seat cost, and the Multichannel plan the one to model carefully before you commit a team to it. [SCREENSHOT NEEDED: Lemlist, the pricing page showing the Email and Multichannel plans with the per user wording visible] Instantly Growth at $47 a month. Instantly's pricing page lists "Unlimited Email Accounts", "Unlimited Email Warmup", "1000 Uploaded Contacts" and "5000 Emails Monthly". Hypergrowth at $97 a month. Same unlimited accounts and warmup, with "25 000 Uploaded Contacts" and "125 000 Emails Monthly". Lightspeed at $358 a month, with "500 000 Emails Monthly" and "100 000 Uploaded Contacts". Annual billing takes 10% off, at $37.60, $77.60 and $286.30 a month respectively. Note what happens between the first two tiers. The price roughly doubles and the sending allowance goes up twenty-five times. If you are on Growth and sending anywhere near the cap, you are paying the worst rate in this entire article. [SCREENSHOT NEEDED: Instantly, the pricing page showing the Growth and Hypergrowth allowances side by side] Smartlead Smartlead's pricing page lists Base at $39 a month, with "2,000 contacts", "6,000 Email sends" and "2,000 Verified Emails". Pro at $94 a month, with "30,000 contacts", "90,000 Email sends" and "30,000 Verified Emails". Unlimited Smart at $174 and Unlimited Prime at $379, both with unlimited contacts and 150,000 and 500,000 email sends respectively. Annual billing takes 17% off, the largest annual discount in the set, at $32.50, $78.30, $144.50 and $314.60. Unlimited email accounts are included on every tier at no extra cost, and email verification credits are bundled rather than sold separately, which is a real difference from the credit model. [SCREENSHOT NEEDED: Smartlead, the pricing page showing the four tiers with contact and send limits] Saleshandy Saleshandy's pricing page lists Outreach Starter at $36 a month monthly, or $25 a month on annual billing, with 6,000 emails a month, 2,000 active prospects and unlimited email accounts. Outreach Pro at $99 monthly, or $69 annually, with 150,000 emails a month and 30,000 active prospects. Outreach Scale at $199 monthly or $139 annually, with 240,000 emails and 60,000 prospects, adding whitelabel and SSO. Outreach Scale Plus at $299 monthly or $209 annually, with 300,000 emails and 100,000 prospects, adding a dedicated success manager. Which makes Outreach Pro the cheapest email allowance in this article at roughly $0.66 per thousand emails on monthly billing, cheaper per email than plans costing three times as much. [SCREENSHOT NEEDED: Saleshandy, the pricing page showing the monthly and annual toggle on the Outreach tiers] Woodpecker, which prices differently on purpose "$7.00 per 100 Contacted prospects". Woodpecker's pricing page uses a usage-based model rather than named tiers, with annual billing stated to save 33%. Unlimited team members and unlimited email accounts are free, along with catch-all email verification. The base calculator position includes 16,000 emails a month, 4,000 stored prospects, 4 warm-ups and 100 Lead Finder credits. Add-ons are itemised, including LinkedIn outreach at "$29 /monthly per LinkedIn account connected", extra warm-ups at "$5 /monthly per email account", email addresses at "$6 /monthly" for Google or Microsoft and "$4 /monthly" for Maildoso or Mailforge, dedicated servers at "$59 /monthly per server" and an agency panel at "$27 /monthly" per active client. Model this one on prospects, not emails. A five step sequence to 1,000 people is 1,000 contacted prospects and up to 5,000 emails, which is $70 here. The same activity is inside the entry tier almost everywhere else. Run your own numbers, because the answer swings hard on sequence length. [SCREENSHOT NEEDED: Woodpecker, the pricing calculator showing the per prospect rate and the add-on list] The number nobody publishes: cost per thousand emails [INSERT CHART 2, best-lemlist-alternatives-chart-2-per-thousand.svg] Alt: Computed cost per thousand emails across six published cold email plans in 2026, showing the entry tier penalty. This is our arithmetic on their published figures, and here is the working. Divide the monthly list price by the monthly email allowance, then multiply by a thousand. The entry tiers. Instantly Growth is $47 over 5,000 emails, or $9.40 per thousand. Smartlead Base is $39 over 6,000, or $6.50. Saleshandy Outreach Starter is $36 over 6,000, or $6.00. The tier above. Lemlist Email is $69 over 50,000, or $1.38. Instantly Hypergrowth is $97 over 125,000, or $0.78. Saleshandy Outreach Pro is $99 over 150,000, or $0.66. Which is the finding. Across three independent vendors the second tier gives roughly fifteen to twenty-five times the sending allowance for roughly two to three times the price. Instantly goes from 5,000 to 125,000 emails for a price increase of about 2.1 times. Saleshandy goes from 6,000 to 150,000 for about 2.75 times. Smartlead goes from 6,000 to 90,000 for about 2.4 times. The practical read. If you are on an entry tier and using most of it, you are almost certainly better off one tier up, and the saving is not marginal. If you are on an entry tier and using a fraction of it, you are paying for headroom you will never touch. A caveat that matters. These rates assume you use the full allowance, which almost nobody does. Compute yours on your real sending volume rather than on the cap. Which one actually fits [INSERT CHART 3, best-lemlist-alternatives-chart-3-fit.svg] Alt: Which cold email platform suits which team in 2026, mapped by number of seats needed against monthly sending volume. One person, low volume. Almost any of them, and the entry tiers exist for exactly this. Pick on interface and move on. One person, real volume. The step-up tiers, and this is where the per thousand arithmetic pays for the twenty minutes it takes. A team, real volume. Check the seat model first. Lemlist Multichannel is the only one here that multiplies by headcount, and for five people that is $545 a month against $97 or $99 elsewhere. Long sequences, modest lists. Woodpecker's per prospect model is worth modelling properly, because a long sequence costs the same there and more everywhere else. And if the problem is deliverability rather than software, the tool is not the variable. Our deliverability guide covers what actually moves inbox placement, and our infrastructure roundup covers the layer underneath the sending tool. What we do not publish here Any deliverability or reply rate comparison between these tools. We have not run a controlled test with matched lists, offers and domains, and every public figure of that kind comes from one of the vendors. An overall ranking. The unit differs by vendor, so a single ordering would be misleading by construction. Negotiated or annual-only pricing beyond what each vendor publishes. Every figure here is the published list price. Feature-by-feature tables. They go stale within a quarter and the two week trials are free. Any claim about which tool is safest for your domains. That depends on your infrastructure and your sending behaviour, not on the vendor. FAQ What is the cheapest Lemlist alternative? On headline price, Saleshandy Outreach Starter at $25 a month billed annually and Smartlead Base at $32.50 annually. On cost per email sent, Saleshandy Outreach Pro at roughly $0.66 per thousand and Instantly Hypergrowth at roughly $0.78. Those are different questions and they have different answers. Is Lemlist expensive? The Email plan at $69 a month for 50,000 emails with unlimited users is competitive, working out at about $1.38 per thousand emails with no per-seat cost. The Multichannel plan at $109 per user a month is where it becomes expensive for teams, because it is the only plan in this comparison that multiplies with headcount. Which cold email tool is best for agencies? Look at the workspace and client features rather than the send price. Smartlead offers a clients and workspace feature from the Pro plan, Saleshandy adds whitelabel and SSO from Outreach Scale, and Woodpecker sells an agency panel at $27 a month per active client. Those are the lines that matter at agency scale. How much should cold email software cost per month? For one person sending real volume, roughly $70 to $100 a month buys 50,000 to 150,000 emails across these vendors. Below that you are on an entry tier paying five to twelve times more per email. Above it you are buying headroom you should check you need. Does Woodpecker work out cheaper? It depends entirely on sequence length. At $7.00 per 100 contacted prospects, a long sequence to a modest list is cheap because you pay per person rather than per email. A short sequence to a very large list is not. Model your own numbers before deciding. Should you switch tools to save money? Only after computing your real cost per thousand emails on your actual volume, and only after checking the seat model. The most common saving available is not a switch at all, it is moving one tier up with your existing vendor. Bottom line Do not read the sticker prices as a ranking. Work out two numbers first: how many people need a login, and how many emails you actually send in a month. If you need seats, Lemlist Multichannel is the only plan here that charges by headcount and it should be modelled against the unlimited-user alternatives before you commit. If you send real volume, compute cost per thousand emails on your own figures, because the entry tiers across this category run five to twelve times the rate of the tier above and stepping up usually buys fifteen to twenty-five times the allowance for double the price. And if your sequences are long and your lists are modest, Woodpecker's per prospect model deserves a proper calculation rather than a glance. Everything else in this category is decided by a free trial. Want the outbound run rather than the tool chosen? Book a call with GROU. We run outbound and lead generation inside B2B revenue engines across verticals. We are GROU, a B2B pipeline agency that runs lead generation, outbound, and LinkedIn content for clients across manufacturing, fintech, iGaming, software, and professional services. Some links in this article are affiliate links, including Lemlist, Instantly and Woodpecker. Every price quoted is the published list price taken from each vendor's own pricing page and verified in August 2026, and the cost per thousand figures are our own arithmetic on those numbers. Prices change, so check before you buy.](https://framerusercontent.com/images/oP9oy999nFzcIm3HqB5SD9X3ZIs.jpg?width=1600&height=900)


