NEW WEBINAR: Learn how to fill your B2B webinar seatsHow to fill your B2B webinar seatsSave your seat
×
NEW WEBINAR: Learn how to fill your B2B webinar seatsHow to fill your B2B webinar seatsSave your seat
×
NEW WEBINAR: Learn how to fill your B2B webinar seatsHow to fill your B2B webinar seatsSave your seat
×

How to run a lead generation pilot 2026

How to run a lead generation pilot 2026

How to run a lead generation pilot 2026

How to run a lead generation pilot 2026

How to run a lead generation pilot 2026

How to run a lead generation pilot 2026

Author

Aljaz Peklaj

modern flat 2D digital illustration, solid vivid sky blue background colour hex 7BC5D6 highly saturated soft sky blue, a single cream white paper airplane at slight 3/4 angle in the centre flying to the upper right, the body of the paper airplane is shaped like a small cream white pipe section instead of a regular folded paper body, a single small mustard yellow hex F5B841 orb trailing behind the pipe airplane from the rear showing the pilot is already producing early pipeline, three small cream white motion lines behind the airplane showing it is in flight and testing, the paper airplane IS the lightweight pilot and the pipe body IS the lead gen system being test-flown, a small cobalt blue hex 203EEB dot at the airplane nose, a small burnt orange hex E85A32 dot at the wing tip, a small coral red hex FF6B6B dot at the trailing orb, centered composition with slight 3/4 perspective, generous negative space, 3 main elements, deep navy outlines hex 0B1E4A medium weight on all objects, flat colour fills with simple cel shading using one darker tone per shape, smooth clean surfaces, soft drop shadow beneath airplane in slightly darker sky blue, modern SaaS marketing illustration, in the style of Timo Kuilder and Dawid Ryski, premium b2b aesthetic, playful cartoon proportions, no photorealism, no 3D rendering --ar 3:2 --style raw --s 150 --v 7 --no brushed texture, grain, abstract shapes
Share this article
Table of content
0 min read

Most pilots are designed so that they cannot fail. Which is the same thing as designing them so that they cannot inform.

Three hundred emails over four weeks, and then a meeting about whether it worked. At a two percent reply rate that pilot produces six replies. At three percent it produces nine. Nobody in that meeting can tell the difference between a good campaign and a bad one, because three replies is not a difference, it is noise. So the decision gets made on tone of voice and how the calls felt, which is exactly the decision you were trying to avoid making.

TL;DR

A lead generation pilot is a measurement exercise, and most are sized too small to measure anything. Using the standard sample size formula published by NIST, telling a two percent reply rate apart from a four percent one takes roughly 483 sends. Telling two percent from three percent takes about 1,747. Telling two percent from two and a half takes about 6,587. A 300-email pilot can only detect enormous differences, so it will almost always come back inconclusive and get read as a failure. Length has the same problem in reverse: Google's own sender guidance tells you to "Start with a low sending volume to engaged users, and slowly increase the volume over time", so the first weeks of any new sending setup measure your warm-up rather than your offer. Size the pilot to the difference you actually need to detect, run it long enough to clear the ramp, and agree the decision rule in writing before the first email goes out.

The pilot that cannot fail is the pilot that cannot inform

A pilot is a measurement, not a trial run. Everyone treats it as a low-commitment way to see whether an agency or a channel is any good. It is really an attempt to estimate one number, your reply or meeting rate, precisely enough to make a spending decision worth five or six figures.

Which means it has a required size, and the size is not negotiable by wanting it smaller. If you want to know whether your reply rate is two percent or four percent, there is an amount of sending below which you cannot know. Running less than that does not give you a weaker answer. It gives you a number that is indistinguishable from chance.

The failure mode is not a bad result. It is an ambiguous one. An ambiguous pilot gets interpreted, and interpretation is where the person who wanted to buy sees promise and the person who did not sees waste. Both are reading the same six replies.

So the real design question is: what difference do I need to detect? Not "how many emails can I afford". Work backwards from the decision. If you would sign a retainer at a three percent reply rate and walk away at two, then three versus two is the difference you must be able to see, and that sets everything else.

Size it against the difference you need to see

How many sends a lead generation pilot needs in 2026 to distinguish one reply rate from another with statistical confidence.

The formula is public and it is not complicated. NIST's Engineering Statistics Handbook publishes the sample size needed to test whether a proportion differs from an assumed value by a detectable amount. For a two-sided test at the conventional 95 percent confidence and 80 percent power, against a baseline of two percent, it produces the following.

To detect two percent versus six percent, about 141 sends. A tripling is easy to see. If your pilot only needs to answer "is this catastrophically broken", a small pilot is fine.

To detect two percent versus four percent, about 483 sends. A doubling. This is the most common real threshold, and it is already well above what most pilots run.

To detect two percent versus three percent, about 1,747 sends. A fifty percent improvement. This is the difference that usually decides whether a channel is worth funding, and it needs roughly six times the volume of a typical pilot.

To detect two percent versus two and a half, about 6,587 sends. At this point you are no longer running a pilot. You are running the campaign.

Read the ladder rather than the individual numbers. Halving the difference you want to detect roughly quadruples the sending you need. That relationship is why "let's start small and see" quietly fails: small pilots can only find effects so large that you would have noticed them without measuring.

And if 1,747 sends is not affordable, that is useful information too. It means you cannot answer the question you asked, so you should either change the question to one a smaller pilot can answer, or accept that the first phase is an operational test rather than a performance test, and say so out loud. Our piece on cold email benchmarks covers what rates people actually report and why you should not plan against them.

a campaign dashboard showing a low send volume alongside a reply rate percentage, to illustrate how a small denominator produces an unstable percentage

The first weeks measure your warm-up, not your offer

What each week of a lead generation pilot actually measures in 2026, from the sending ramp through to the offer itself.

Google tells you to ramp, in its own sender guidelines. The published advice is to "Start with a low sending volume to engaged users, and slowly increase the volume over time", and to "Send email at a consistent rate. Avoid sending email in bursts." A new domain and a new mailbox going straight to full volume is the pattern those guidelines are written against.

So a four week pilot on new infrastructure spends most of itself climbing. Whatever reply rate you see in weeks one and two is a measurement of your deliverability ramp, not of your message. If you then average the whole period, you have deliberately contaminated the number you are paying to learn.

You also cannot see the deliverability data at that volume. Google's own Postmaster Tools documentation states that "Data might be missing if the total number of messages for a given day is too low. This is to protect users' privacy." A small pilot therefore sits in a gap where you can neither read your reputation signals nor trust your reply rate. Our cold email deliverability guide covers the setup that has to be right before any of this is worth measuring.

Which is why the pilot needs a stated ramp and a stated measurement window. Weeks one and two are the ramp. Week three is where list quality shows up, in bounces and in silence from addresses that should have replied. Weeks four onward are the measurement period, and they are the only weeks whose numbers go into the decision. Write that split down before you start, because after the fact everyone will want to include or exclude the early weeks depending on which way the numbers went.

The volume thresholds also mean the pilot is not the same compliance surface as the campaign. Google's bulk sender requirements apply to senders of "more than 5,000 messages per day to Gmail accounts", so a pilot below that line is not exercising the requirements your scaled programme will have to meet. Plan for them in the pilot even though they do not yet bite.

The unsubscribe and identification duties apply from the first email, though. The UK regulator's electronic mail marketing guidance states that "You can send unsolicited electronic mail marketing to corporate subscribers without consent or a soft opt-in", and in the same guidance that "You must not disguise or hide your identity in messages to either type of subscriber. You must provide a valid contact address for recipients to opt out or unsubscribe." A pilot is not a period during which those stop applying. This is not legal advice, and the position differs by jurisdiction.

[SCREENSHOT NEEDED: Google Postmaster Tools showing a sparse or empty chart, illustrating the low-volume data gap]

Agree the decision rule before the first email

What to agree in writing before a lead generation pilot starts in 2026, and what each term protects against.

Write down the number that means yes. Not "we will see how it goes". A specific rate, on a specific metric, measured over a specific window. If you cannot name it before you start, you will name it afterwards, and you will name it to match whatever happened.

Write down the number that means no, separately. These are not the same threshold with a different sign. There is usually a middle band where the honest answer is "extend and re-measure", and deciding in advance that the band exists prevents an inconclusive result from being argued into a conclusive one.

Name who owns the list. A pilot that produces a targeting list, verified contacts and a working sequence has produced an asset. Whether that asset stays with you if you do not continue is a contract question, not a goodwill question, and it is much easier to ask now.

Name what happens to the domains and mailboxes. If the pilot ran on infrastructure the agency owns, then the warmed sending reputation you paid several weeks to build does not come with you. If it ran on yours, you keep it and you also keep the risk.

Agree the reporting cadence and the raw data. Weekly is right. Ask for sends, deliveries, opens if they are tracked at all, replies split into positive, neutral and negative, and meetings booked, as counts rather than percentages. Percentages on small denominators are the mechanism by which a pilot gets oversold.

Agree who writes the copy and who approves it. The most common reason a pilot underperforms is three rounds of internal approval sanding the message down to something nobody could object to, which is also something nobody replies to.

Fix the offer for the duration. Changing the message halfway through a pilot means you have run two underpowered tests instead of one adequately powered one, which is strictly worse than either.

a one-page pilot agreement or scope document showing the success criteria section

The list matters more than the copy at this size

At pilot volumes, targeting error dominates everything else. A thousand well-chosen contacts and a mediocre email will beat five thousand loosely chosen ones and a good email, because the reply rate you are measuring is a property of the pairing, not of the message alone.

So build the list before you write anything. If the pilot is a test of whether this segment responds, the segment has to be defined tightly enough that the result generalises to the segment you would scale into. A pilot list assembled from whoever was easiest to find tells you about ease of finding, and nothing else. Our ICP framework is the version of this we use.

And verify it, because bounces come out of your denominator and your reputation at the same time. A ten percent bounce rate on a 500-send pilot removes fifty sends from an already thin sample and damages the sending reputation you will need for the real campaign.

One segment, not four. Splitting a small pilot across multiple industries or personas produces four samples that are each too small to read, and then a comparison between them that is entirely noise. If you want to compare segments, that is a separate and much larger exercise.

What a pilot genuinely cannot tell you

Whether the channel works for your business over a full sales cycle. A pilot measures the top of the funnel. If your cycle is six months, the pilot cannot tell you about close rates, and any projection to revenue is arithmetic performed on an assumption.

Whether this agency is good. It tells you whether this agency, on this list, with this offer, in this window, produced this many replies. Operational quality shows up over months, in how they handle the weeks that go badly. Our piece on the first 90 days with an outbound agency covers what that period actually looks like.

Whether the offer is right. A pilot tests one offer. A poor result is at least as likely to mean the offer was wrong as the channel was, and those two conclusions lead to opposite decisions.

What it will cost at scale. Per-meeting cost in a pilot is calculated on a handful of meetings and is close to meaningless as a unit economic. Treat it as a range, widely.

What we do not publish here

Benchmark reply rates, meeting rates or cost per meeting. Ours come from a specific set of clients, offers and markets, and publishing them as though they were general would be exactly the kind of number this article argues against planning around.

A recommended pilot budget. It follows from the volume you need, which follows from the difference you need to detect, which is specific to your decision.

A recommended pilot length in weeks. It depends on how quickly your infrastructure can ramp and how many sends per day the list supports without burning it.

Any claim about how inbox providers decide placement. Neither Google nor Microsoft publishes its filtering mechanics, and everything in circulation on the subject is inference.

Legal advice. The sender requirements referenced here are technical, not legal, and the rules on contacting people without prior consent differ by jurisdiction. Take advice on your own markets.

FAQ

How many emails should a lead generation pilot send?

Enough to detect the difference that changes your decision. Against a two percent baseline, distinguishing two from four percent takes roughly 483 sends, two from three percent roughly 1,747, and two from two and a half roughly 6,587, using the standard sample size formula for a proportion at 95 percent confidence and 80 percent power. Decide which of those questions you are asking first.

How long should a lead generation pilot run?

Long enough that the measurement window sits after the sending ramp, not across it. Google's guidance is to start at low volume and increase slowly, so on new infrastructure the first two weeks are warm-up and should be excluded from the decision by agreement, not by argument afterwards.

What is a good reply rate for a pilot?

We do not publish one, and you should be careful with anyone who does without saying which market, offer and list it came from. The more useful question is what rate would make you sign, because that is the number the pilot has to be sized to detect.

Should a pilot test more than one segment?

No, not at pilot volumes. Splitting a small sample across segments produces several samples too small to read individually and a comparison between them that is noise. Pick the segment you would scale into and test that one properly.

Who should own the domains a pilot runs on?

Decide it in writing before the pilot starts. Sending reputation takes weeks to build and does not transfer, so if the pilot runs on infrastructure you do not own, you start again from zero if you continue with someone else.

What should you do if the pilot result is inconclusive?

Extend and re-measure rather than interpret, which is why the middle band belongs in the agreement from the start. An inconclusive result is a sample size problem, and the only thing that fixes a sample size problem is more sample.

Bottom line

A pilot is not a smaller version of the campaign. It is a measurement exercise with a required size, and if you run it below that size you have bought an anecdote at campaign prices. Work out which difference in reply rate would change your decision, look up how much sending that difference needs, and either commit to that volume or change the question. Then put the ramp period, the measurement window, the yes number, the no number, the middle band, the list ownership and the domain ownership in writing before the first email leaves. None of that costs anything, and all of it prevents the meeting where six replies get argued about for an hour.

Want a pilot sized to answer the question rather than to look affordable? Book a call with GROU. We run lead generation and outbound inside B2B revenue engines across verticals.

We are GROU, a B2B pipeline agency that runs lead generation, outbound, and LinkedIn content for clients across manufacturing, fintech, iGaming, software, and professional services. The sample size figures in this article are computed by us from the published NIST formula and shown with their inputs so you can check them. Nothing here is legal advice.

Most pilots are designed so that they cannot fail. Which is the same thing as designing them so that they cannot inform.

Three hundred emails over four weeks, and then a meeting about whether it worked. At a two percent reply rate that pilot produces six replies. At three percent it produces nine. Nobody in that meeting can tell the difference between a good campaign and a bad one, because three replies is not a difference, it is noise. So the decision gets made on tone of voice and how the calls felt, which is exactly the decision you were trying to avoid making.

TL;DR

A lead generation pilot is a measurement exercise, and most are sized too small to measure anything. Using the standard sample size formula published by NIST, telling a two percent reply rate apart from a four percent one takes roughly 483 sends. Telling two percent from three percent takes about 1,747. Telling two percent from two and a half takes about 6,587. A 300-email pilot can only detect enormous differences, so it will almost always come back inconclusive and get read as a failure. Length has the same problem in reverse: Google's own sender guidance tells you to "Start with a low sending volume to engaged users, and slowly increase the volume over time", so the first weeks of any new sending setup measure your warm-up rather than your offer. Size the pilot to the difference you actually need to detect, run it long enough to clear the ramp, and agree the decision rule in writing before the first email goes out.

The pilot that cannot fail is the pilot that cannot inform

A pilot is a measurement, not a trial run. Everyone treats it as a low-commitment way to see whether an agency or a channel is any good. It is really an attempt to estimate one number, your reply or meeting rate, precisely enough to make a spending decision worth five or six figures.

Which means it has a required size, and the size is not negotiable by wanting it smaller. If you want to know whether your reply rate is two percent or four percent, there is an amount of sending below which you cannot know. Running less than that does not give you a weaker answer. It gives you a number that is indistinguishable from chance.

The failure mode is not a bad result. It is an ambiguous one. An ambiguous pilot gets interpreted, and interpretation is where the person who wanted to buy sees promise and the person who did not sees waste. Both are reading the same six replies.

So the real design question is: what difference do I need to detect? Not "how many emails can I afford". Work backwards from the decision. If you would sign a retainer at a three percent reply rate and walk away at two, then three versus two is the difference you must be able to see, and that sets everything else.

Size it against the difference you need to see

How many sends a lead generation pilot needs in 2026 to distinguish one reply rate from another with statistical confidence.

The formula is public and it is not complicated. NIST's Engineering Statistics Handbook publishes the sample size needed to test whether a proportion differs from an assumed value by a detectable amount. For a two-sided test at the conventional 95 percent confidence and 80 percent power, against a baseline of two percent, it produces the following.

To detect two percent versus six percent, about 141 sends. A tripling is easy to see. If your pilot only needs to answer "is this catastrophically broken", a small pilot is fine.

To detect two percent versus four percent, about 483 sends. A doubling. This is the most common real threshold, and it is already well above what most pilots run.

To detect two percent versus three percent, about 1,747 sends. A fifty percent improvement. This is the difference that usually decides whether a channel is worth funding, and it needs roughly six times the volume of a typical pilot.

To detect two percent versus two and a half, about 6,587 sends. At this point you are no longer running a pilot. You are running the campaign.

Read the ladder rather than the individual numbers. Halving the difference you want to detect roughly quadruples the sending you need. That relationship is why "let's start small and see" quietly fails: small pilots can only find effects so large that you would have noticed them without measuring.

And if 1,747 sends is not affordable, that is useful information too. It means you cannot answer the question you asked, so you should either change the question to one a smaller pilot can answer, or accept that the first phase is an operational test rather than a performance test, and say so out loud. Our piece on cold email benchmarks covers what rates people actually report and why you should not plan against them.

a campaign dashboard showing a low send volume alongside a reply rate percentage, to illustrate how a small denominator produces an unstable percentage

The first weeks measure your warm-up, not your offer

What each week of a lead generation pilot actually measures in 2026, from the sending ramp through to the offer itself.

Google tells you to ramp, in its own sender guidelines. The published advice is to "Start with a low sending volume to engaged users, and slowly increase the volume over time", and to "Send email at a consistent rate. Avoid sending email in bursts." A new domain and a new mailbox going straight to full volume is the pattern those guidelines are written against.

So a four week pilot on new infrastructure spends most of itself climbing. Whatever reply rate you see in weeks one and two is a measurement of your deliverability ramp, not of your message. If you then average the whole period, you have deliberately contaminated the number you are paying to learn.

You also cannot see the deliverability data at that volume. Google's own Postmaster Tools documentation states that "Data might be missing if the total number of messages for a given day is too low. This is to protect users' privacy." A small pilot therefore sits in a gap where you can neither read your reputation signals nor trust your reply rate. Our cold email deliverability guide covers the setup that has to be right before any of this is worth measuring.

Which is why the pilot needs a stated ramp and a stated measurement window. Weeks one and two are the ramp. Week three is where list quality shows up, in bounces and in silence from addresses that should have replied. Weeks four onward are the measurement period, and they are the only weeks whose numbers go into the decision. Write that split down before you start, because after the fact everyone will want to include or exclude the early weeks depending on which way the numbers went.

The volume thresholds also mean the pilot is not the same compliance surface as the campaign. Google's bulk sender requirements apply to senders of "more than 5,000 messages per day to Gmail accounts", so a pilot below that line is not exercising the requirements your scaled programme will have to meet. Plan for them in the pilot even though they do not yet bite.

The unsubscribe and identification duties apply from the first email, though. The UK regulator's electronic mail marketing guidance states that "You can send unsolicited electronic mail marketing to corporate subscribers without consent or a soft opt-in", and in the same guidance that "You must not disguise or hide your identity in messages to either type of subscriber. You must provide a valid contact address for recipients to opt out or unsubscribe." A pilot is not a period during which those stop applying. This is not legal advice, and the position differs by jurisdiction.

[SCREENSHOT NEEDED: Google Postmaster Tools showing a sparse or empty chart, illustrating the low-volume data gap]

Agree the decision rule before the first email

What to agree in writing before a lead generation pilot starts in 2026, and what each term protects against.

Write down the number that means yes. Not "we will see how it goes". A specific rate, on a specific metric, measured over a specific window. If you cannot name it before you start, you will name it afterwards, and you will name it to match whatever happened.

Write down the number that means no, separately. These are not the same threshold with a different sign. There is usually a middle band where the honest answer is "extend and re-measure", and deciding in advance that the band exists prevents an inconclusive result from being argued into a conclusive one.

Name who owns the list. A pilot that produces a targeting list, verified contacts and a working sequence has produced an asset. Whether that asset stays with you if you do not continue is a contract question, not a goodwill question, and it is much easier to ask now.

Name what happens to the domains and mailboxes. If the pilot ran on infrastructure the agency owns, then the warmed sending reputation you paid several weeks to build does not come with you. If it ran on yours, you keep it and you also keep the risk.

Agree the reporting cadence and the raw data. Weekly is right. Ask for sends, deliveries, opens if they are tracked at all, replies split into positive, neutral and negative, and meetings booked, as counts rather than percentages. Percentages on small denominators are the mechanism by which a pilot gets oversold.

Agree who writes the copy and who approves it. The most common reason a pilot underperforms is three rounds of internal approval sanding the message down to something nobody could object to, which is also something nobody replies to.

Fix the offer for the duration. Changing the message halfway through a pilot means you have run two underpowered tests instead of one adequately powered one, which is strictly worse than either.

a one-page pilot agreement or scope document showing the success criteria section

The list matters more than the copy at this size

At pilot volumes, targeting error dominates everything else. A thousand well-chosen contacts and a mediocre email will beat five thousand loosely chosen ones and a good email, because the reply rate you are measuring is a property of the pairing, not of the message alone.

So build the list before you write anything. If the pilot is a test of whether this segment responds, the segment has to be defined tightly enough that the result generalises to the segment you would scale into. A pilot list assembled from whoever was easiest to find tells you about ease of finding, and nothing else. Our ICP framework is the version of this we use.

And verify it, because bounces come out of your denominator and your reputation at the same time. A ten percent bounce rate on a 500-send pilot removes fifty sends from an already thin sample and damages the sending reputation you will need for the real campaign.

One segment, not four. Splitting a small pilot across multiple industries or personas produces four samples that are each too small to read, and then a comparison between them that is entirely noise. If you want to compare segments, that is a separate and much larger exercise.

What a pilot genuinely cannot tell you

Whether the channel works for your business over a full sales cycle. A pilot measures the top of the funnel. If your cycle is six months, the pilot cannot tell you about close rates, and any projection to revenue is arithmetic performed on an assumption.

Whether this agency is good. It tells you whether this agency, on this list, with this offer, in this window, produced this many replies. Operational quality shows up over months, in how they handle the weeks that go badly. Our piece on the first 90 days with an outbound agency covers what that period actually looks like.

Whether the offer is right. A pilot tests one offer. A poor result is at least as likely to mean the offer was wrong as the channel was, and those two conclusions lead to opposite decisions.

What it will cost at scale. Per-meeting cost in a pilot is calculated on a handful of meetings and is close to meaningless as a unit economic. Treat it as a range, widely.

What we do not publish here

Benchmark reply rates, meeting rates or cost per meeting. Ours come from a specific set of clients, offers and markets, and publishing them as though they were general would be exactly the kind of number this article argues against planning around.

A recommended pilot budget. It follows from the volume you need, which follows from the difference you need to detect, which is specific to your decision.

A recommended pilot length in weeks. It depends on how quickly your infrastructure can ramp and how many sends per day the list supports without burning it.

Any claim about how inbox providers decide placement. Neither Google nor Microsoft publishes its filtering mechanics, and everything in circulation on the subject is inference.

Legal advice. The sender requirements referenced here are technical, not legal, and the rules on contacting people without prior consent differ by jurisdiction. Take advice on your own markets.

FAQ

How many emails should a lead generation pilot send?

Enough to detect the difference that changes your decision. Against a two percent baseline, distinguishing two from four percent takes roughly 483 sends, two from three percent roughly 1,747, and two from two and a half roughly 6,587, using the standard sample size formula for a proportion at 95 percent confidence and 80 percent power. Decide which of those questions you are asking first.

How long should a lead generation pilot run?

Long enough that the measurement window sits after the sending ramp, not across it. Google's guidance is to start at low volume and increase slowly, so on new infrastructure the first two weeks are warm-up and should be excluded from the decision by agreement, not by argument afterwards.

What is a good reply rate for a pilot?

We do not publish one, and you should be careful with anyone who does without saying which market, offer and list it came from. The more useful question is what rate would make you sign, because that is the number the pilot has to be sized to detect.

Should a pilot test more than one segment?

No, not at pilot volumes. Splitting a small sample across segments produces several samples too small to read individually and a comparison between them that is noise. Pick the segment you would scale into and test that one properly.

Who should own the domains a pilot runs on?

Decide it in writing before the pilot starts. Sending reputation takes weeks to build and does not transfer, so if the pilot runs on infrastructure you do not own, you start again from zero if you continue with someone else.

What should you do if the pilot result is inconclusive?

Extend and re-measure rather than interpret, which is why the middle band belongs in the agreement from the start. An inconclusive result is a sample size problem, and the only thing that fixes a sample size problem is more sample.

Bottom line

A pilot is not a smaller version of the campaign. It is a measurement exercise with a required size, and if you run it below that size you have bought an anecdote at campaign prices. Work out which difference in reply rate would change your decision, look up how much sending that difference needs, and either commit to that volume or change the question. Then put the ramp period, the measurement window, the yes number, the no number, the middle band, the list ownership and the domain ownership in writing before the first email leaves. None of that costs anything, and all of it prevents the meeting where six replies get argued about for an hour.

Want a pilot sized to answer the question rather than to look affordable? Book a call with GROU. We run lead generation and outbound inside B2B revenue engines across verticals.

We are GROU, a B2B pipeline agency that runs lead generation, outbound, and LinkedIn content for clients across manufacturing, fintech, iGaming, software, and professional services. The sample size figures in this article are computed by us from the published NIST formula and shown with their inputs so you can check them. Nothing here is legal advice.

Most pilots are designed so that they cannot fail. Which is the same thing as designing them so that they cannot inform.

Three hundred emails over four weeks, and then a meeting about whether it worked. At a two percent reply rate that pilot produces six replies. At three percent it produces nine. Nobody in that meeting can tell the difference between a good campaign and a bad one, because three replies is not a difference, it is noise. So the decision gets made on tone of voice and how the calls felt, which is exactly the decision you were trying to avoid making.

TL;DR

A lead generation pilot is a measurement exercise, and most are sized too small to measure anything. Using the standard sample size formula published by NIST, telling a two percent reply rate apart from a four percent one takes roughly 483 sends. Telling two percent from three percent takes about 1,747. Telling two percent from two and a half takes about 6,587. A 300-email pilot can only detect enormous differences, so it will almost always come back inconclusive and get read as a failure. Length has the same problem in reverse: Google's own sender guidance tells you to "Start with a low sending volume to engaged users, and slowly increase the volume over time", so the first weeks of any new sending setup measure your warm-up rather than your offer. Size the pilot to the difference you actually need to detect, run it long enough to clear the ramp, and agree the decision rule in writing before the first email goes out.

The pilot that cannot fail is the pilot that cannot inform

A pilot is a measurement, not a trial run. Everyone treats it as a low-commitment way to see whether an agency or a channel is any good. It is really an attempt to estimate one number, your reply or meeting rate, precisely enough to make a spending decision worth five or six figures.

Which means it has a required size, and the size is not negotiable by wanting it smaller. If you want to know whether your reply rate is two percent or four percent, there is an amount of sending below which you cannot know. Running less than that does not give you a weaker answer. It gives you a number that is indistinguishable from chance.

The failure mode is not a bad result. It is an ambiguous one. An ambiguous pilot gets interpreted, and interpretation is where the person who wanted to buy sees promise and the person who did not sees waste. Both are reading the same six replies.

So the real design question is: what difference do I need to detect? Not "how many emails can I afford". Work backwards from the decision. If you would sign a retainer at a three percent reply rate and walk away at two, then three versus two is the difference you must be able to see, and that sets everything else.

Size it against the difference you need to see

How many sends a lead generation pilot needs in 2026 to distinguish one reply rate from another with statistical confidence.

The formula is public and it is not complicated. NIST's Engineering Statistics Handbook publishes the sample size needed to test whether a proportion differs from an assumed value by a detectable amount. For a two-sided test at the conventional 95 percent confidence and 80 percent power, against a baseline of two percent, it produces the following.

To detect two percent versus six percent, about 141 sends. A tripling is easy to see. If your pilot only needs to answer "is this catastrophically broken", a small pilot is fine.

To detect two percent versus four percent, about 483 sends. A doubling. This is the most common real threshold, and it is already well above what most pilots run.

To detect two percent versus three percent, about 1,747 sends. A fifty percent improvement. This is the difference that usually decides whether a channel is worth funding, and it needs roughly six times the volume of a typical pilot.

To detect two percent versus two and a half, about 6,587 sends. At this point you are no longer running a pilot. You are running the campaign.

Read the ladder rather than the individual numbers. Halving the difference you want to detect roughly quadruples the sending you need. That relationship is why "let's start small and see" quietly fails: small pilots can only find effects so large that you would have noticed them without measuring.

And if 1,747 sends is not affordable, that is useful information too. It means you cannot answer the question you asked, so you should either change the question to one a smaller pilot can answer, or accept that the first phase is an operational test rather than a performance test, and say so out loud. Our piece on cold email benchmarks covers what rates people actually report and why you should not plan against them.

a campaign dashboard showing a low send volume alongside a reply rate percentage, to illustrate how a small denominator produces an unstable percentage

The first weeks measure your warm-up, not your offer

What each week of a lead generation pilot actually measures in 2026, from the sending ramp through to the offer itself.

Google tells you to ramp, in its own sender guidelines. The published advice is to "Start with a low sending volume to engaged users, and slowly increase the volume over time", and to "Send email at a consistent rate. Avoid sending email in bursts." A new domain and a new mailbox going straight to full volume is the pattern those guidelines are written against.

So a four week pilot on new infrastructure spends most of itself climbing. Whatever reply rate you see in weeks one and two is a measurement of your deliverability ramp, not of your message. If you then average the whole period, you have deliberately contaminated the number you are paying to learn.

You also cannot see the deliverability data at that volume. Google's own Postmaster Tools documentation states that "Data might be missing if the total number of messages for a given day is too low. This is to protect users' privacy." A small pilot therefore sits in a gap where you can neither read your reputation signals nor trust your reply rate. Our cold email deliverability guide covers the setup that has to be right before any of this is worth measuring.

Which is why the pilot needs a stated ramp and a stated measurement window. Weeks one and two are the ramp. Week three is where list quality shows up, in bounces and in silence from addresses that should have replied. Weeks four onward are the measurement period, and they are the only weeks whose numbers go into the decision. Write that split down before you start, because after the fact everyone will want to include or exclude the early weeks depending on which way the numbers went.

The volume thresholds also mean the pilot is not the same compliance surface as the campaign. Google's bulk sender requirements apply to senders of "more than 5,000 messages per day to Gmail accounts", so a pilot below that line is not exercising the requirements your scaled programme will have to meet. Plan for them in the pilot even though they do not yet bite.

The unsubscribe and identification duties apply from the first email, though. The UK regulator's electronic mail marketing guidance states that "You can send unsolicited electronic mail marketing to corporate subscribers without consent or a soft opt-in", and in the same guidance that "You must not disguise or hide your identity in messages to either type of subscriber. You must provide a valid contact address for recipients to opt out or unsubscribe." A pilot is not a period during which those stop applying. This is not legal advice, and the position differs by jurisdiction.

[SCREENSHOT NEEDED: Google Postmaster Tools showing a sparse or empty chart, illustrating the low-volume data gap]

Agree the decision rule before the first email

What to agree in writing before a lead generation pilot starts in 2026, and what each term protects against.

Write down the number that means yes. Not "we will see how it goes". A specific rate, on a specific metric, measured over a specific window. If you cannot name it before you start, you will name it afterwards, and you will name it to match whatever happened.

Write down the number that means no, separately. These are not the same threshold with a different sign. There is usually a middle band where the honest answer is "extend and re-measure", and deciding in advance that the band exists prevents an inconclusive result from being argued into a conclusive one.

Name who owns the list. A pilot that produces a targeting list, verified contacts and a working sequence has produced an asset. Whether that asset stays with you if you do not continue is a contract question, not a goodwill question, and it is much easier to ask now.

Name what happens to the domains and mailboxes. If the pilot ran on infrastructure the agency owns, then the warmed sending reputation you paid several weeks to build does not come with you. If it ran on yours, you keep it and you also keep the risk.

Agree the reporting cadence and the raw data. Weekly is right. Ask for sends, deliveries, opens if they are tracked at all, replies split into positive, neutral and negative, and meetings booked, as counts rather than percentages. Percentages on small denominators are the mechanism by which a pilot gets oversold.

Agree who writes the copy and who approves it. The most common reason a pilot underperforms is three rounds of internal approval sanding the message down to something nobody could object to, which is also something nobody replies to.

Fix the offer for the duration. Changing the message halfway through a pilot means you have run two underpowered tests instead of one adequately powered one, which is strictly worse than either.

a one-page pilot agreement or scope document showing the success criteria section

The list matters more than the copy at this size

At pilot volumes, targeting error dominates everything else. A thousand well-chosen contacts and a mediocre email will beat five thousand loosely chosen ones and a good email, because the reply rate you are measuring is a property of the pairing, not of the message alone.

So build the list before you write anything. If the pilot is a test of whether this segment responds, the segment has to be defined tightly enough that the result generalises to the segment you would scale into. A pilot list assembled from whoever was easiest to find tells you about ease of finding, and nothing else. Our ICP framework is the version of this we use.

And verify it, because bounces come out of your denominator and your reputation at the same time. A ten percent bounce rate on a 500-send pilot removes fifty sends from an already thin sample and damages the sending reputation you will need for the real campaign.

One segment, not four. Splitting a small pilot across multiple industries or personas produces four samples that are each too small to read, and then a comparison between them that is entirely noise. If you want to compare segments, that is a separate and much larger exercise.

What a pilot genuinely cannot tell you

Whether the channel works for your business over a full sales cycle. A pilot measures the top of the funnel. If your cycle is six months, the pilot cannot tell you about close rates, and any projection to revenue is arithmetic performed on an assumption.

Whether this agency is good. It tells you whether this agency, on this list, with this offer, in this window, produced this many replies. Operational quality shows up over months, in how they handle the weeks that go badly. Our piece on the first 90 days with an outbound agency covers what that period actually looks like.

Whether the offer is right. A pilot tests one offer. A poor result is at least as likely to mean the offer was wrong as the channel was, and those two conclusions lead to opposite decisions.

What it will cost at scale. Per-meeting cost in a pilot is calculated on a handful of meetings and is close to meaningless as a unit economic. Treat it as a range, widely.

What we do not publish here

Benchmark reply rates, meeting rates or cost per meeting. Ours come from a specific set of clients, offers and markets, and publishing them as though they were general would be exactly the kind of number this article argues against planning around.

A recommended pilot budget. It follows from the volume you need, which follows from the difference you need to detect, which is specific to your decision.

A recommended pilot length in weeks. It depends on how quickly your infrastructure can ramp and how many sends per day the list supports without burning it.

Any claim about how inbox providers decide placement. Neither Google nor Microsoft publishes its filtering mechanics, and everything in circulation on the subject is inference.

Legal advice. The sender requirements referenced here are technical, not legal, and the rules on contacting people without prior consent differ by jurisdiction. Take advice on your own markets.

FAQ

How many emails should a lead generation pilot send?

Enough to detect the difference that changes your decision. Against a two percent baseline, distinguishing two from four percent takes roughly 483 sends, two from three percent roughly 1,747, and two from two and a half roughly 6,587, using the standard sample size formula for a proportion at 95 percent confidence and 80 percent power. Decide which of those questions you are asking first.

How long should a lead generation pilot run?

Long enough that the measurement window sits after the sending ramp, not across it. Google's guidance is to start at low volume and increase slowly, so on new infrastructure the first two weeks are warm-up and should be excluded from the decision by agreement, not by argument afterwards.

What is a good reply rate for a pilot?

We do not publish one, and you should be careful with anyone who does without saying which market, offer and list it came from. The more useful question is what rate would make you sign, because that is the number the pilot has to be sized to detect.

Should a pilot test more than one segment?

No, not at pilot volumes. Splitting a small sample across segments produces several samples too small to read individually and a comparison between them that is noise. Pick the segment you would scale into and test that one properly.

Who should own the domains a pilot runs on?

Decide it in writing before the pilot starts. Sending reputation takes weeks to build and does not transfer, so if the pilot runs on infrastructure you do not own, you start again from zero if you continue with someone else.

What should you do if the pilot result is inconclusive?

Extend and re-measure rather than interpret, which is why the middle band belongs in the agreement from the start. An inconclusive result is a sample size problem, and the only thing that fixes a sample size problem is more sample.

Bottom line

A pilot is not a smaller version of the campaign. It is a measurement exercise with a required size, and if you run it below that size you have bought an anecdote at campaign prices. Work out which difference in reply rate would change your decision, look up how much sending that difference needs, and either commit to that volume or change the question. Then put the ramp period, the measurement window, the yes number, the no number, the middle band, the list ownership and the domain ownership in writing before the first email leaves. None of that costs anything, and all of it prevents the meeting where six replies get argued about for an hour.

Want a pilot sized to answer the question rather than to look affordable? Book a call with GROU. We run lead generation and outbound inside B2B revenue engines across verticals.

We are GROU, a B2B pipeline agency that runs lead generation, outbound, and LinkedIn content for clients across manufacturing, fintech, iGaming, software, and professional services. The sample size figures in this article are computed by us from the published NIST formula and shown with their inputs so you can check them. Nothing here is legal advice.

Trusted by industry leaders

Trusted by industry leaders

Trusted by industry leaders

Ready to build qualified pipeline?

Ready to build qualified pipeline?

Ready to build qualified pipeline?

Book a call to see if we're the right fit, or take the 2-minute quiz to get a clear starting point.

Book a call to see if we're the right fit, or take the 2-minute quiz to get a clear starting point.

Book a call to see if we're the right fit, or take the 2-minute quiz to get a clear starting point.