How to A/B test LinkedIn posts when you cannot really A/B test

The Growtempo Team13 min read

You cannot truly A/B test an organic LinkedIn post. There is no mechanism that splits your audience and shows half of them one version and half another, so any two posts you compare differ not only in what you changed but in who was online, what else was in the feed, who commented first, and which slice of your network LinkedIn tested it on. The honest substitute is a paired-post routine: publish variants in matched slots over several weeks, change exactly one structural thing, and only believe a difference that is large and repeats.

The short version

  • Organic LinkedIn posts have no split-traffic mechanism. Every comparison is between two different moments, not two versions of one moment.
  • The workable substitute is paired posts in matched slots, repeated eight to ten times, with the order alternated so the weekday does not track the variant.
  • You are testing a class of posts, never two versions of the same post, because you cannot publish the same idea twice.
  • Test things you expect to change results by a lot. Small effects stay invisible at the volume an individual publishes.
  • Even 12,988 posts cannot separate a 206-character hook from a 205-character one. Your forty posts will separate less.

Why can you not A/B test LinkedIn posts properly?

A real A/B test has one requirement that does all the work: the two groups have to be comparable in every respect except the thing you changed. That is what randomly splitting traffic buys you. Everything else about the world is held constant by design, so a difference in outcome can be attributed to the variant.

An organic LinkedIn post gets none of that. Publish version A on Tuesday and version B on Thursday and here is a partial list of what else differed:

  • Who was online. Your audience is not a fixed panel. A different subset opened the app, and the ones who did were in a different mood, at a different point in their week.
  • The initial audience LinkedIn chose. A post goes first to some slice of your network, and what happens in that slice governs everything after. Two posts get two different slices.
  • Who commented first. Early engagement shapes onward distribution, so a single well-connected person replying in the first twenty minutes can change the whole trajectory. That is covered in what actually happens in the golden hour.
  • Everything else in the feed. You are competing with whatever else your audience follows published that morning, which you neither control nor observe.
  • Your own recent posting. What you published three days earlier affects who sees this one.

There is a subtler problem underneath all of that. You cannot publish the same post twice with one word changed, because your audience would see both. So the two things you are comparing are never two versions of one post. They are two different posts that share a structural property. You are testing a class, not a variant, and that is a weaker claim than the one A/B testing usually makes.

None of this means measurement is pointless. It means the claims you can support are coarser than the ones people usually make. “Video posts get roughly double the comments for me” is supportable. “This hook increased engagement 12%” is not, and never was.

What can you test on LinkedIn, and what can you not?

The dividing line is not what you are curious about. It is how large an effect the variable plausibly has, and whether you can hold everything else steady while you change it.

VariableWorth testing?Why
Post format (text, image, video, document)Yes, firstThe largest gap in our cohort. Video is 26.3% of the top decile and 10.1% of the bottom half. Easy to hold everything else constant while you change it.
Opening style (story, question, declaration)Yes, secondFirst-person openers run 19.6% in the top decile against 10.1% in the bottom half. Cheap to vary and it is the part readers judge first.
Topic or content pillarYes, but slowlyUsually the biggest effect on a personal account, and the slowest to read, because changing topics also changes who follows you.
Hashtag countWeaklyFour or more hashtags appear in 41.9% of bottom-half posts against 25.9% of the top decile, which is suggestive. Your own sample will struggle to confirm it.
Post lengthNoHook length does not separate our top decile from our bottom half at all: 206 characters against 205. If it is invisible there, it is invisible in your archive.
Posting time within a dayNoThe effect is small and the confounds are enormous. You would need a year of posts in each slot with nothing else changing.
Call to action wordingNoA small effect on a noisy outcome. Choose one on judgement and keep it. See how to end a LinkedIn post.
Emoji useCareful yesA very large association in our data, 24.3% against 2.7%, but heavily confounded with format and author. Worth trying, worth doubting.

If you are running paid campaigns, that is a different product with its own controls, and nothing in this article applies to it. Everything here is about organic posting from a personal profile.

What is a paired post test and how do you run one?

A paired post test is the closest honest approximation to an experiment available to an individual. The idea is to make the two groups as similar as you can on everything except the variable, and then repeat the pairing enough times that the differences you could not control average out.

  1. Pick one variable and two levels. Text post against image post. First-person opening against declarative opening. One variable. If you change two, you learn nothing about either.
  2. Fix everything else you can. Same posting slot, same rough length, same content pillar, same call to action, same use of hashtags. Write both posts in the same sitting so your own state of mind is roughly constant.
  3. Publish them in matched slots. Same weekday and time in consecutive weeks is usually the cleanest available pairing.
  4. Alternate the order. If variant A always goes first, the first week of every pair is doing the work and you have measured week position instead of the variant. Flip the order every pair.
  5. Repeat eight to ten times. One pair proves nothing at all. Ten pairs pointing the same direction is worth acting on, and it takes about five months at two posts a week, which is the real cost of this exercise.
  6. Compare group medians, not pairs. Counting how many pairs variant A won throws away the size of the differences. Take the median rate per 1,000 followers for each group and look at both that and the pair-by-pair record.

Then run the check that makes or breaks the whole thing: drop the single best post from each group and recompute. If your conclusion flips, you found one post. That test is the backbone of analysing your LinkedIn post performance and it catches most false findings before they cost you a quarter.

How do you hold the topic constant while varying the opener?

This is the most useful single test available, because the opening line is what a reader judges before deciding whether to expand the post, and it is the cheapest thing in the whole post to change.

The method is to write from one content pillar for a run of posts, and rotate the opening style deliberately rather than by mood. Say your pillar is hiring. Over ten posts you might alternate between opening with a specific incident from your own week, and opening with a flat claim about how hiring works. Same subject, same audience, same slot, two openings.

A worked example of what varying the opener looks like, with both versions invented here for illustration rather than taken from the data: a first-person opening might begin “I turned down a candidate last week for a reason I am still not sure was fair.” The declarative version of the same idea begins “Most hiring rejections are decided on something nobody writes in the feedback.” Same argument underneath. Completely different promise to the reader.

Two cautions specific to opener tests. Your own writing gets better over the run, so put the styles in alternating order rather than doing five of one then five of the other. And keep the openings genuinely different in kind, not in wording. Testing two versions of a first-person opening will produce nothing, because the variable you actually care about did not change. The taxonomy of openings and the numbers behind each is in what 12,988 posts say about hooks and how to start a LinkedIn post.

When employees start leaving your team, and managers complain about employees not having loyalty, other managers or companies "poaching" them, or making last-minute desperate counters with things they should …see more

1,347 reactions · 73 comments · 15,213 followers · From our dataset, roughly 108 engagements per 1,000 followers. There is no way to know what the same argument would have earned with a first-person opening, because that version was never published. This missing comparison is the whole problem.

How many posts do you need before you believe a difference?

The honest answer is that it depends on how much your own posts vary, and for most people the number is larger than the number of posts they will publish this year. That is not a satisfying answer, so here is the reasoning behind it in plain terms.

Start with volume. Someone posting four times a week publishes a bit over two hundred posts a year. Split between two variants that is roughly a hundred each, and only if you change nothing else for twelve months, which nobody does. In practice a test that runs a quarter gives you perhaps a dozen posts per side.

Now add the spread. In our cohort, the top-decile cut sits at 5.95 engagements per 1,000 followers while the median is 0.40. The distance between a normal post and a good one is larger than an order of magnitude. Individual accounts show the same shape: most posts cluster low and a few go far above. When the variation inside each group is that wide, a modest difference between the groups is buried underneath it. You are trying to hear a conversation in a room where everyone is shouting at random intervals.

Put those two facts together and the conclusion is uncomfortable but simple. With a dozen posts per side and that much spread, you can see differences of the “roughly double” kind. You cannot see differences of the “about fifteen percent better” kind, and no amount of careful spreadsheet work will conjure them into view. Most creators do not post often enough to detect small effects. That is a fact about arithmetic, not a criticism of anyone's discipline.

We are deliberately not giving you a number of posts per variant, a confidence threshold, or a formula. Any such number would need an estimate of your personal variation that neither you nor we have, and publishing one would be inventing precision. What we can give you are three checks that do not require inventing anything:

  • The drop-one test. Remove the best post from each side. Does the answer survive?
  • The split-in-half test. Split your test period into the first half and the second half. Does the same variant win in both? If not, you are looking at something that changed over time, not at the variant.
  • The describe-it-out-loud test.Can you state the difference in a sentence a colleague would find obviously large? “My video posts get about twice the comments” passes. “Declarative openings did 8% better” does not, because 8% is well inside the noise.

What do 12,988 posts say about which differences are detectable at all?

Our own dataset is a useful calibration instrument, because it is enormous compared with any individual archive. If a variable fails to separate the best posts from the worst across 12,988 posts from 65 creators, an individual account has no chance with it.

Hook featureTop 10%Bottom 50%Verdict for your own testing
Contains an emoji24.3%2.7%Huge gap. Detectable, though confounded with format and author.
Native video26.3%10.1%Large. The best first test for most accounts.
Four or more hashtags25.9%41.9%Clear, and in the opposite direction to popular belief.
First-person opening19.6%10.1%Roughly double. Worth a deliberate run of paired posts.
Says “you” anywhere in the hook47.5%32.7%Solid. Cheap to apply without a formal test.
Opens with a question word12.4%10.6%Small even at this sample size. Do not expect to see it in your own data.
Contains a number32.1%32.8%No separation. Stop testing this.
Median hook length206 characters205 charactersNo separation. A one-character difference across 12,988 posts.

The last two rows are the most useful ones in this article. Numbers in hooks and hook length are two of the most confidently advised variables in the niche, and across a very large sample they do not distinguish the best posts from the worst at all. If you have been running informal tests on either, you have been reading noise carefully.

One important qualification: these are correlations across a cohort of established creators, not causal results from an experiment, and they describe the visible hook rather than complete posts. A large gap here means the variable is worth your attention. It does not mean adding an emoji to your next post will do anything.

How do you keep the day of the week out of your results?

The most common way a homemade LinkedIn test destroys itself is by accidentally testing the calendar. If variant A always goes out on Tuesday and variant B on Thursday, then the result is about Tuesday and Thursday, and the variant is along for the ride.

Three fixes, all cheap:

  • Same slot, consecutive weeks. Both variants get the same weekday and time, one week apart. The pairing is across weeks rather than within one.
  • Alternate which variant leads. Pair one runs A then B. Pair two runs B then A. Over ten pairs, whatever is systematically true about first weeks cancels out.
  • Keep the surrounding schedule steady. If you post three times in test week one and once in test week two, the posts are competing with different amounts of your own recent activity.

On the underlying question of whether the day even matters much, the published research is summarised in the best time to post on LinkedIn, and the short version is that it is a second-order lever compared with what your first line says.

What should you test first?

In order, and with the reasoning attached, because the order is doing most of the work.

  1. Format. Largest observed gap, easiest to change, and the change is binary so there is nothing to interpret. Run text against video, or text against document, for two months. See text-only against image posts for what our cohort shows before you start.
  2. Opening style. Second largest, and the one that transfers to every post you write afterwards even outside the test.
  3. Content pillar. Often the biggest real effect and the slowest to read, because switching topics changes your audience as well as your engagement. Give it two quarters, not two months.
  4. Everything else. Hashtags, length, time, emoji, call to action. These are worth a decision, not a test. Decide once on the balance of evidence, apply it consistently, and spend the attention you saved on the first three.

How do you log a paired post test?

One row per post, two rows per pair, and a column that records the thing you will otherwise forget: what else was happening that week. The log below is deliberately shorter than a full performance tracking sheet, because a test you have to maintain elaborately is a test you will abandon in week four.

ColumnExample entryWhy it is here
Pair number4Lets you drop a contaminated pair cleanly rather than one orphaned post
VariantA (text) or B (video)The one thing you changed
Order within pairFirst or secondProves you alternated, and lets you check whether first weeks win systematically
SlotTuesday 08:30Confirms the pairing actually matched, rather than drifting by two days
PillarHiringKeeps the subject matter constant, which is the confound people forget
Followers at publish4,180The denominator. Without it, later pairs look better than earlier ones by default.
Reactions and unique commenters62 and 9Two components that respond to different things, so record them separately
Rate per 1,00023.4The comparable number across the whole test
ContaminationReshared by a 200k accountThe column that stops one lucky post ending the test early

At the end of the run you need four numbers: the median rate for each variant, the same two medians with each group's best post removed, and the pair-by-pair record. If the medians and the record disagree, the record is usually telling you that one post carried the whole result. That is a finding about that post.

What makes a LinkedIn test worthless?

  • Changing two things at once. A video with a new opening style is not a format test. You will attribute the result to whichever change you were more excited about.
  • Stopping when you get the answer you wanted. If you would have kept going had the first three pairs gone the other way, the test was decorative.
  • Testing during an unusual period. A product launch, a conference, a holiday fortnight, or a week when you were quoted somewhere. Note it and discard those pairs.
  • Counting reactions only. Reactions and comments respond to different things. A variant that lifts reactions and flattens comments is not obviously better. Use a weighted measure and look at both components.
  • Comparing across a period when your audience grew a lot. Raw counts drift upward as followers accumulate. Normalise per 1,000 followers, or you are measuring your own growth.
  • Running the test for three weeks. The single most common failure. Three weeks is six posts and six posts is a coin flip.

When should you stop testing and just write?

Sooner than most people expect. The honest position is that the majority of what makes writing good is not testable at the volume an individual publishes. Whether a specific story is worth telling, whether an argument is actually yours, whether the ending lands: none of that decomposes into a variable you can hold constant across twenty posts.

So run one structured test at a time, on the format or the opening, and let everything else be decided by reading. Read your comments. Notice which posts produce replies that say something and which produce agreement. Notice which ones get you messages. That feedback arrives faster than any measurement and it is about the part of the work that actually varies.

The measurement is worth doing because it stops you believing things that are false. It is not going to tell you what to write. Those are different jobs and it is worth keeping them separate.

How we handle this

Our product writes a daily LinkedIn post in your voice and holds it for 24 hours so you can edit it or kill it before it goes out. Because it knows the format, opening style and topic of everything it drafts, the inputs for a paired comparison get recorded without anyone maintaining a spreadsheet. We deliberately do not describe that as A/B testing, because it is not, and any tool that offers you organic A/B testing on LinkedIn is describing something the platform does not provide.

About this data

Numbers come from our analysis of a public dataset of 34,012 LinkedIn influencer posts. We scored 12,988 English posts from 65 creators by engagement rate (reactions + 4× comments, divided by the author's followers) and compared the top 10% against the bottom half. The dataset captures each post's text up to LinkedIn's “see more” fold, which is exactly what a reader sees before deciding to engage. These are correlations, not guarantees. Full methodology and caveats are in the full study.

All the figures above come from the same cohort, and the caveats matter for a piece about measurement more than for most. It is 12,988 English posts from 65 creators, all with at least 1,000 followers and at least 25 reactions per post, so it describes established accounts. The relationships are correlational rather than causal, which is precisely the limitation this article is about. Engagement counts are lifetime-cumulative at collection time and post ages vary. The text findings describe the visible hook above LinkedIn's fold, not complete post bodies. Full methodology is in our study of 34,000 LinkedIn posts.

The short version on A/B testing LinkedIn posts

You cannot A/B test LinkedIn posts in the strict sense, because organic posts have no split-traffic mechanism and no two posts ever meet the same audience under the same conditions. What you can run is a paired post routine: one variable, matched slots, alternating order, eight to ten pairs, compared on group medians and checked by dropping your best post from each side. Test format first and opening style second, because those are the only variables with gaps large enough for an individual archive to see. Everything smaller is a decision, not a test.

Frequently asked questions

Can you A/B test LinkedIn posts?

Not in the strict sense. Organic LinkedIn posts have no split-traffic mechanism, so two variants are never shown to comparable groups at the same moment. What you can do is run paired posts on matched days over several weeks, changing one structural thing at a time, and only trust differences that are large and repeat.

How do you test a LinkedIn hook without a real A/B test?

Hold everything else fixed and vary only the opening style across a run of posts. Same format, similar length, same posting slot, same topic area. Alternate which style goes first from week to week so the day of the week does not line up with the variant. Then compare group medians, not individual pairs.

How many LinkedIn posts do you need to test something?

More than most people publish. Someone posting four times a week produces roughly two hundred posts a year, and splitting them between two variants leaves about a hundred each only if nothing else changes all year. Small differences stay invisible at that volume, so test things you expect to matter a lot.

Why do my two similar LinkedIn posts get completely different results?

Because everything except the post changed too. Different people were online, a different set of first commenters shaped early distribution, the news cycle moved, and LinkedIn showed it to a different initial slice of your network. That variation is usually larger than the difference you were trying to measure.

What is worth testing on LinkedIn?

Format and opening style, in that order. In our cohort of 12,988 posts, native video appears in 26.3% of top-decile posts against 10.1% of the bottom half, and first-person openers in 19.6% against 10.1%. Those are large gaps. Posting time and hashtag counts produce differences too small for an individual account to detect.

Want posts that already follow this data?

Growtempo writes and publishes a LinkedIn post in your voice every day, with these findings built into how it writes. You approve each one before it goes live.

Get Started