Outreach experiments

How to A/B test LinkedIn outreach messages without picking a false winner

By Moez Zhioua12 min read
Two differently marked message cards beside randomly divided prospect cards and an open experiment notebook, illustrating a controlled outreach comparison

Version A gets 10 positive replies from 100 people. Version B gets 16 from 100. Calling that a 60% relative lift is correct arithmetic. Calling B the proven winner is a different claim, and those counts alone do not justify it.

This is where an outreach test can become expensive. You spend weeks refining the sentence that appeared to win, when a different mix of prospects, unequal follow-up time or ordinary variation might explain the result.

Before changing the campaign, you need to explain who received each version and whether the apparent advantage could reasonably be variation. The test plan below makes that possible, including when the answer is 'we don't know yet.'

01

Write the decision before writing version B

Start with a question you could act on. For example: 'For operations managers at recently expanded service companies, does asking about handoff ownership produce more relevant conversations than asking about reporting work?' Define the eligible audience and the evidence that qualifies a company before dividing the list.

Choose a primary outcome that matches the offer. If you are offering a practical worksheet, a positive reply might mean a recipient asks to see it or discusses the problem it addresses. A polite acknowledgment is not the same outcome. Neither is an angry response, even though both increase total reply count.

Write a short classification rubric before the replies arrive. Keep total human responses, positive responses, meetings and opt-outs as separate measures. Review ambiguous replies without knowing the assigned version where practical. Don't redefine 'positive' after seeing which version benefits.

Specify whether you are testing an invitation note, the first message after acceptance or an entire sequence. For an invitation test, acceptance is an outcome of the treatment. Comparing later replies only among people who accepted can select different kinds of people into each group. Preserve an assigned-population measure as well as the stage-level diagnostics.

For a first-message-only test, give both groups a defined interval without follow-ups. If follow-ups are included, keep their policy consistent and call the result a comparison of that message within the sequence. A reply after the third touch does not identify which sentence caused it.

02

Make the difference between the messages explicit

A narrow change makes the result easier to interpret. Keep the sender, offer, audience rules and follow-up policy constant, then change the part you want to evaluate. That might be the opening question, the amount of context or the next step you propose.

Consider this original, fictional example for companies whose public information confirms a recent second-location opening. Both messages begin: 'Hi Morgan, I saw your announcement about the second location.' Version A then asks, 'Who owns the handoff when a job moves between locations?' Version B asks, 'Where do you record the handoff when a job moves between locations?' Both end: 'I have a short handoff checklist if it would help.' The tested difference is an ownership question versus a record-keeping question, not a new product offer.

Those lines are only appropriate when the announcement is real, the person plausibly owns the work and the promised checklist exists. A test is not permission to invent a shared connection, customer result or operational problem.

You can also compare two substantially different messages. Call that a message-package test. If B changes the opener, offer and request, its result cannot establish that the opener alone helped. Likewise, personalized outreach can be a defined policy: record which research is required and how it becomes relevant context. It does not require sending identical strings to everyone.

Save the exact templates and personalization rules with a version identifier. If someone improves B halfway through, the earlier and later recipients no longer received the same treatment. Stop and document a necessary correction; start a new version rather than silently blending the records.

03

Split the audience randomly, then freeze the assignment

Start with one deduplicated pool of eligible prospects, excluding people you should not contact. Assign each person to one version before sending. Don't give your strongest leads to the personalized version and compare them with everyone else. Don't send A to the first ten search results and B to the next ten: result order may encode relevance or other differences.

For a small spreadsheet-managed pool, generate a random number for every eligible row, immediately paste those numbers as values, sort by them and allocate the planned share to A and B. Preserve the resulting identity-to-version mapping. A live random formula that recalculates when you edit the sheet can move people between versions.

If multiple senders or materially different segments are involved, plan the split within those groups and retain the group labels. Avoid assigning all of one sender's contacts to A and another sender's contacts to B when the question concerns copy. Keep the analysis consistent with the allocation design; small subgroups do not create reliable separate winners.

Interleave the sending schedule so both versions have comparable opportunities across the same period. Sending every A on Monday and every B on Friday adds timing to the comparison. Freeze the same person's assignment across retries and follow-ups; they should not receive both versions as if they were independent prospects.

When several recipients work at the same company, they may discuss your messages. Account-level assignment can reduce that interference, but it changes the analysis: people inside an account are not automatically independent observations. Microsoft's research on organizational experiments explains why clustered assignment needs different uncertainty calculations. For a modest list, choosing one appropriate contact per account may make the initial question simpler; don't manufacture extra independence by counting coworkers separately.

A platform's variation feature is convenient only if its allocation, persistence and reporting fit this plan. Separate campaigns can also be usable when those conditions are controlled. Neither setup is valid merely because a dashboard labels it an experiment.

04

Plan the sample and the response clock

There is no universally sufficient 50, 100 or 200 sends per version. The required sample depends on the baseline outcome rate, the smallest improvement worth detecting, the chosen false-positive threshold, desired statistical power and allocation between groups. It also depends on whether the observations are independent. A rare meeting outcome can require a different plan from a common acknowledgment.

Use a calculation designed for the actual comparison, such as two independent proportions when that is your design. Record the inputs and method. A tool asking for a standardized effect size may need a transformation, not the raw percentage-point difference. A one-sample proportion formula is not interchangeable with a two-arm calculation.

If the required eligible population is larger than you can reasonably reach, reduce the ambition of the claim. Run a pilot to check message clarity, classification and sending reliability. You can make a provisional operating choice from limited evidence, but don't present it as a statistically established winner or keep contacting unsuitable people to fill a sample.

Define enrollment and observation separately. Set how many people you intend to assign, the last enrollment date and what happens if you miss the planned sample. Also set an equal outcome window for each person. For example, you might observe positive replies for 14 days from each assigned person's scheduled first-contact opportunity. Fourteen days is an illustrative convention here, not a proven optimum.

Record both the scheduled opportunity and actual first send. A long sending delay can leave less time for the message to work under that policy measure; it is a delivery problem to report, not a timestamp to rewrite afterward. Finalize the comparison only when every included person's defined window has closed. Track later replies separately. If meetings are another outcome, give them their own predeclared window and maturity check.

05

Keep the assigned population and the delivery record

For an experiment about the outreach policy, retain outcomes against the originally assigned eligible people. Someone whose attempt fails remains part of that assignment record. So does someone who replies, books a meeting or opts out. Removing converted people from the denominator would change the population according to the result.

Keep a separate delivery view showing confirmed sends, failures, unattempted records and unknown states by version. A descriptive reply rate among confirmed recipients can help diagnose execution, but it answers a different question from outcomes per assigned person. Don't switch between them because one produces a more attractive lift.

Unknown is not zero. If you cannot determine whether messages or replies were recorded, disclose the missing outcomes and investigate. Preserve the assignment rows, but don't present their outcome comparison as complete. Check native threads where appropriate before retrying an ambiguous send; a software error does not necessarily prove nothing was sent.

Compare the observed allocation with the plan. A small count difference is not automatically evidence of a broken split. A substantial, statistically unexpected mismatch can indicate assignment, logging or filtering trouble. Microsoft's sample-ratio-mismatch guidance is useful here: investigate the mechanism before trusting the result.

Keep a brief event log alongside the summary: identity key, account, sender, version, assignment time, scheduled opportunity, actual sends, outcome evidence, opt-out status and any correction. Store only the prospect information needed for the work and limit access to it.

06

Work through the 10-versus-16 result

Return to the fictional example. Two hundred eligible, independent prospects are randomly assigned, 100 to each version. All receive a confirmed send and complete the same observation window. There are no missing outcomes. This makes the assigned and confirmed-recipient denominators equal for this example; that will not always happen in practice.

The categories below use each person's first human response. Positive, other human and no human response are mutually exclusive. Opt-outs are an additional contact-control flag within the other-human group, not extra people. These figures illustrate interpretation, not real OutreachGenie results or a recommended sample size.

Fictional teaching example. Response categories sum to 100 in each group; opt-out flags overlap those categories. The sample does not establish a winning version.
Outcome at the fixed cutoffVersion AVersion B
Assigned people with confirmed sends100100
Positive first human responses1016
Other first human responses1210
No human response7874
All human responders22 / 100 = 22%26 / 100 = 26%
Positive response rate10 / 100 = 10%16 / 100 = 16%
Opt-out flags within other responses23

07

Separate the observed lift from the conclusion

B's positive response rate is 6 percentage points higher: 16% minus 10%. Relative to A's 10%, that is a 60% increase. Report both, together with 10 of 100 versus 16 of 100. A large relative percentage can describe a small number of additional responses.

Under a two-sided, independent two-proportion z-test with pooled variance and no continuity correction, these fictional counts give z about 1.262 and p about 0.207. The calculation uses uncertainty from both groups, as NIST's formula specifies. It does not meet a preselected 0.05 threshold. Sparse counts or dependent observations require a method appropriate to those conditions, not automatic reuse of this example.

That result is inconclusive about a difference under this test. It does not prove the messages perform equally. Nor does p = 0.207 mean a 20.7% chance that B is wrong, or a 79.3% chance that B is better. The American Statistical Association explicitly distinguishes p-values from probabilities that a hypothesis is true.

For the same fictional B-minus-A comparison, a 95% Newcombe interval built from Wilson proportion intervals runs from about -3.5 to +15.5 percentage points. It includes a modest decline as well as a substantial improvement. That range, calculated for independent binomial samples without continuity correction, makes the uncertainty more useful than the headline lift alone. It is not a 95% probability that the fixed underlying difference lies inside this particular interval.

Compare that uncertainty with the improvement that would matter commercially. Even strong evidence of a tiny effect may not justify substantially more research time per message. With limited evidence, you may still choose the less costly version provisionally, document why and revisit it when useful new evidence arrives.

Read the actual responses and check the opt-outs before adopting either version. Here, B has one more opt-out flag, but those small counts do not establish a reliable difference in harm. Every opt-out still needs to be honored. Statistical uncertainty about a rate is not uncertainty about an individual's request to stop.

08

Monitor harm without testing until something wins

Check delivery failures, complaints, incorrect personalization and account warnings while the test runs. Pause harmful or broken outreach immediately. Keeping an experiment unchanged is not a reason to continue sending a false claim or ignore a request to stop.

For statistical decisions, follow the stopping plan. Recomputing an ordinary fixed-sample p-value every day and stopping the first time it crosses 0.05 can inflate false positives. Extending an unfavorable or inconclusive result until it becomes favorable creates the same kind of selection problem. Valid sequential methods exist, but they must actually account for repeated looks.

Microsoft's experimentation guidance separates early operational monitoring from trustworthy statistical conclusions. Apply that distinction here: investigate problems now, while reserving the winner decision for the planned analysis unless you have a suitable sequential design.

If you inspect many versions, metrics or audience slices, one may look promising by chance. Decide the primary comparison beforehand. Label an unexpected subgroup result as exploratory and seek a fresh, properly planned test if it matters. Don't quietly replace the original question with whichever one produced the smallest p-value.

09

Save the decision, including an inconclusive one

Keep a compact experiment record with the hypothesis, eligible audience, frozen versions, assignment method, sample plan, response rubric, windows, stopping rules, results and limitations. Add the operating decision and the reason. 'We kept A because the result was inconclusive and B required more research' is a useful record. 'B won by 60%' without the counts is not.

In OutreachGenie, prospect records, notes and campaign context can help keep the work organized. Maintain a separate experiment ledger when the available report does not supply the assignment and outcome definitions you need. This guide does not assume the product automatically randomizes variants, classifies positive replies or calculates statistical significance.

After adopting a version, continue checking delivery and conversation quality. Treat a new audience, offer or sending policy as a changed context. The result belongs to the population and conditions you tested; it is not a permanent ranking of two sentences.

Common questions

Questions that come up in practice

How many LinkedIn prospects do I need for an A/B test?

There is no universal minimum that establishes a reliable winner. Plan around the baseline outcome rate, the smallest useful change, power, error threshold, allocation and independence assumptions. Use a calculation for your actual design. If the eligible pool is too small, run a clearly labeled pilot rather than treating an arbitrary send count as proof.

How long should a LinkedIn message test run?

Allow the planned enrollment and equal response window to finish. A two-week calendar run does not give every recipient two weeks to respond. Set a window per assigned contact opportunity, record sending delays and wait for the final person's window to close. The appropriate duration depends on the outcome and your operating context.

Can I A/B test messages manually?

Yes. Use a frozen eligible list, random assignment, saved message versions, a comparable sending schedule and a ledger of outcomes. A spreadsheet can hold the plan, but live random formulas must be frozen so assignments do not change. Manual sending does not remove the need for a suitable analysis or respectful contact practices.

Should I change only one part of the message?

Change one component when you want to understand that component's effect. Comparing two whole messages is also possible, but the result applies to the package, not an individual sentence. In either case, allocation and exposure must support the comparison; one changed variable alone does not establish causation.

What if neither message wins?

Report the result as inconclusive unless a suitable analysis supports a more specific conclusion. Do not call it proof of equality. Keep the clearer or less costly version as a provisional operating choice if appropriate, fix any execution problems, and plan a new test only when the likely information is worth the work.

Is LinkedIn Ads A/B testing the same as testing outreach messages?

No. Paid campaigns have their own delivery, targeting and reporting systems. An ad-rotation setting or audience-overlap discussion does not establish a valid direct-message test. Define the ordinary outreach audience, assignment, message policy and person-level outcomes separately.

Research used for this guide

Try the workflow

Open the workspace and try it with a small, reviewed list.