FIELD NOTES

The post-Andromeda Meta ads framework that produces 2% CTR winners in 60 ad variants

After the Andromeda update Meta's CTR distribution flattened — winners are rarer and creative volume matters more than ever. Here's the 20-ads-per-batch framework we use at TrainYourAgent to systematically produce 2%+ CTR winners on cold B2B traffic in 2026.

Andromeda — Meta's late-2024 ranking system overhaul — broke the creative-testing playbook every paid media buyer used in 2022 and 2023. The old way (test 3 ads, scale the winner, iterate at 5x spend) doesn't work anymore. Andromeda's enrichment model picks winners off micro-signals in the first 200 impressions, and once it decides an ad isn't a winner, no amount of budget recovers it.

The framework that actually works in 2026 is volume + variation + brutal kill-rates. At TrainYourAgent, we use it to produce 2%+ CTR cold-traffic winners for B2B service businesses out of every batch of 20 ads we ship. Here's the system.

What Andromeda actually changed

Three things, in order of impact on creative testing:

  1. Enrichment-stage signals are weighted higher than feedback-stage signals. Translation: Meta uses micro-engagement data (scroll-stop, dwell, hover) from the first ~200 impressions to predict CTR before enough clicks accumulate to be statistically significant. If your ad doesn't stop the thumb in 0.6 seconds, it's dead.
  2. CTR distribution flattened. Pre-Andromeda, a top 5% creative did 4.2x the CTR of a median creative. Post-Andromeda, that ratio is 2.6x. The median got better (more ads "work") but the ceiling got lower (fewer ads are crushing it).
  3. Variant similarity penalty is real. Ships 5 ads with the same hook but different colors? Meta now treats them as one ad for delivery and you waste 80% of your test budget. Variation must be conceptual, not cosmetic.

The framework below is designed around those three facts.

The 20-ads-per-batch framework

Every two weeks we ship a batch of 20 ads to one campaign for one offer. Twenty is not arbitrary — it's the smallest number that produces a statistically robust winner under Andromeda's enrichment thresholds. We've tested 10 (too noisy), 30 (too expensive), 50 (diminishing returns). Twenty works.

Inside the 20 there's a strict structure:

  • 5 hooks × 4 visual treatments = 20 ads

The 5 hooks must hit 5 different message angles. Not 5 wordings of the same angle. The treatments are 4 different visual concepts (UGC selfie video, founder-led talking head, screenshot of a customer DM, animated mockup). No swapping headlines on the same image. That's the variant similarity trap.

The 5 hook angles

These are the 5 hook archetypes we use for B2B service offers. They map to 5 different stages of buyer awareness (Schwartz's framework, post-Andromeda-adjusted):

  1. Pain agitation — Name the specific problem with specific dollar numbers. "Missed calls cost the average HVAC company $3,200/mo. Here's the math."
  2. Counter-narrative — Attack the dominant belief in the market. "Stop hiring more SDRs. The math hasn't worked since 2023."
  3. Proof-first — Lead with the result, then explain. "We took a Boise dental practice from 38 to 71 booked appointments per week. The system is 4 components."
  4. Curiosity gap — Set up a question the buyer must click to answer. "The reason your $4,000/mo agency-built voice agent has a 31% pickup rate (and how to fix it for $200)."
  5. Specificity-as-trust — Hyper-specific operational detail that signals you actually do the work. "Our Twilio webhook returns a TwiML response in 180ms. Here's the architecture."

Each angle gets one hook. Five hooks total. Then each hook gets all 4 visual treatments. Done — 20 ads.

The 4 visual treatments

Treatments are NOT design variations. They're conceptual variations:

  1. Founder-led talking head — 9-second vertical video, founder on a phone camera, no editing, eye contact in the first frame. Highest trust signal in B2B.
  2. UGC-style selfie video — A customer (or actor playing one) recording in their work environment. "I'm an HVAC dispatcher. Here's what changed." Looks like a TikTok.
  3. Screenshot proof — A real customer DM, Slack message, or dashboard screenshot. Static image. Lowest production cost, highest CTR for proof-first hooks.
  4. Animated workflow mockup — 6-second Loom-style screen capture showing the actual product in motion. Best for curiosity-gap and counter-narrative hooks.

Kill rules (this is where most buyers fail)

Andromeda decides in 200 impressions. You decide at 1,000. Past that, you're funding losers.

Kill at 1,000 impressions if:

  • CTR < 0.8% on cold traffic
  • CPM is 2x the campaign average (Meta is taxing the ad)
  • 3s-video view rate < 12%

Kill at 5,000 impressions if:

  • CTR < 1.2% AND no purchase / qualified-lead event
  • Frequency > 2.1 in under 72 hours (audience saturation)

Promote at 5,000 impressions if:

  • CTR > 1.6%
  • Hook rate (3s-view ÷ impressions for video) > 28%
  • Quality ranking "above average"

The brutal part: out of 20 ads, you typically kill 14 before they hit 1,000 impressions, kill 4 more by 5,000, and end up scaling 2. That's the math. If you're not killing 70%+ of every batch, you're not testing aggressively enough.

What 2%+ CTR winners actually look like in 2026

Across our last 12 batches (240 ads tested), the 9 ads that broke 2% CTR cold had three things in common:

  1. The first 0.4 seconds carried a number. Not a word, a number. "$3,200/mo," "47 calls/day," "180ms." Numbers stop scrolls. Words don't.
  2. The hook contradicted the audience's current behavior. Not "here's a better way" — "stop doing this." Negation outperforms affirmation 1.8x in cold B2B.
  3. No agency-looking polish. Selfie video, slight grain, ambient room noise. The over-produced corporate-explainer ads all underperformed by 30–50% in the same batch.

The campaign structure that supports this

Andromeda also broke the "1 campaign, many ad sets, manual budget per ad set" structure. The 2026 structure for B2B service offers is:

  • 1 campaign per offer. Advantage+ Shopping or Advantage+ Sales objective.
  • 1 ad set per campaign. Broad targeting, no detailed interests, no lookalikes (use Advantage+ audiences and let Meta find them).
  • All 20 ads in the one ad set. This is non-negotiable — Andromeda needs the comparison signal inside a single auction context to pick winners.
  • Budget: ~$50/day per ad during testing phase, so $1,000/day for the 20-ad batch.
  • Scale phase: duplicate winners into a separate scale campaign at 3x budget. Don't scale by editing the test campaign — you'll reset the learning.

What this costs in 2026

For a B2B service offer with a $200–$500 lead value:

  • Creative production for 20 ads (UGC + founder vid + screenshots + Loom): ~$1,200–$2,800 if outsourced to a UGC creator network. Under $400 if you produce internally.
  • Media spend per batch: $14,000 over 2 weeks ($1,000/day) to fully test the 20.
  • Expected outcome: 2 winners that scale, 14 dead within 72 hours, 4 marginal.
  • CAC at scale (after 2 batches): typically 40–55% lower than pre-Andromeda CAC on the same offer, once you stop testing the wrong things.

The compounding effect

The framework gets better over time because it generates structured learning. After 5 batches (100 ads), you have a 100-row dataset of (hook angle × visual treatment × CTR) for your exact offer and audience. By batch 10, you can predict winners before they ship — and that's when you stop running 20 and start running 8 high-confidence ads per batch instead. We're at batch 14 internally and our hit rate on 2%+ CTR ads went from 9/240 (3.75%) to 5/80 (6.25%) over the last 3 batches.

If you want to see the actual ad files (10 winning examples with their hook angles, treatments, and final CTR/CPL), book a 15-minute walkthrough. I'll screen-share the ads manager.


Related reading:

Related cornerstones:

Filed under