Likely Beta Sign in

Methodology

How we compare ad tools

We sell one of the tools we compare. So the only thing that makes our comparisons worth reading is a scorecard we fixed before we looked at anyone — published here in full, including the four dimensions where Likely scores zero.

Read this and you can check our bias instead of trusting it. The dimensions were fixed before we looked at anyone, the rubric is public, and every cell cites a page you can open yourself.

Methodology Published 15 August 2026 Last updated 15 August 2026 Re-checked quarterly
On this page
  1. Why vendor comparisons are usually worthless
  2. The eight dimensions
  3. What each score means
  4. Why there is no total score
  5. How "best for" pages are allowed to rank
  6. The evidence rules
  7. Where Likely scores zero
  8. How to correct us

Why vendor comparisons are usually worthless

Whoever picks the axes picks the winner. A vendor writing a comparison chooses the eight rows where they say yes and the competitor says no, and the table does the rest. Nothing in it is technically false.

You have read these. Fourteen green ticks in one column, three in the others. The rows are real features. The scoring is accurate. And the page is still useless, because the selection of rows is the argument, and it was made after the answer was known.

There is no way to fix that with more honesty inside the table. The fix has to be structural: fix the dimensions before looking at the tools, publish them, score everyone on the same card, and refuse to produce a winner. That is what this page is.

The eight dimensions

Eight dimensions, grouped under the three questions a buyer is actually choosing between. They were chosen to describe the category, not to describe Likely — which is why Likely loses an entire group of them.

#DimensionThe question it answers
A · Research — what can it tell you about the market?
1Access modelWhat do you have to hand over before it works?
2CoverageAre you seeing a selection, or all of it?
3StructureIs the data labelled consistently enough to count?
B · Decision — does it help you choose what to make?
4Gap detectionCan it show you what you are not doing?
5PrioritisationDoes it rank by what things cost, or only by how big they are?
6Evidence trailCan you check a claim against the creative behind it?
C · Operations — does it fit how you actually work?
7Workflow surfaceCan your team live in it day to day?
8CommercialsCan you find out what it costs, and does the price punish growth?

Group C is where we lose. Dimensions 7 and 8 describe daily workflow and pricing transparency, and they are in the card precisely because they are the ones a reader is most likely to actually care about on a Tuesday. Leaving them out would have been the easy way to win.

What each score means

Every dimension is scored 0 to 3. Not 0 to 10 — a ten-point scale on judgements like these is false precision, and it invites a decimal place nobody can defend.

The general shape is the same everywhere: 0 absent · 1 partial or manual · 2 present · 3 the best implementation in the category. A 3 is not a compliment, it is a statement about the field as it stands today, and it can move when somebody builds something better.

1 · Access model

ScoreMeans
0Nothing works without connecting an ad account.
1The main value needs a connection; some browsing works without.
2Core research works unconnected; a connection unlocks extras.
3Nothing requires account access at all. No OAuth, no partner request.

2 · Coverage

ScoreMeans
0Only what you saved yourself.
1A searchable library. Selection-based, no denominator.
2A large library with filters. Category volume can be approximated.
3A complete census of a defined set of brands, so shares and gaps are computable rather than estimated.

Library size is deliberately not what this measures. A hundred million ads with no denominator cannot tell you what share of a category does anything. Two hundred ads that are all of a category can.

3 · Structure

ScoreMeans
0No labelling. Folders and free text.
1Free-form tags the user applies by hand.
2Automatic labelling, vocabulary not published.
3A closed vocabulary, published, and stable enough that this quarter compares with last.

4 · Gap detection

ScoreMeans
0Nothing. It shows what exists, not what is missing.
1You can filter and work it out yourself.
2It highlights under-used attributes.
3An explicit matrix with empty cells marked against a real denominator.

5 · Prioritisation

ScoreMeans
0No ranking of any kind.
1Sorts by volume, recency or engagement.
2Ranks by an opportunity or impact measure.
3Ranks on impact, confidence and what it costs to produce.

6 · Evidence trail

ScoreMeans
0Numbers only. You take them on trust.
1Some drill-down, some dead ends.
2Most metrics open the underlying creative.
3Every number opens the ads behind it.

7 · Workflow surface

ScoreMeans
0Analysis only. Nothing to work inside.
1Export, and not much else.
2Two or more of: extension, boards, briefs, asset storage, integrations.
3The full production loop — capture, organise, brief, store, integrate.

8 · Commercials

ScoreMeans
0No public pricing. Demo-led, or not selling yet.
1Public pricing, but gated on your ad spend.
2Public flat pricing by seats or usage.
3Public flat pricing, a real free tier, no spend gate.

Spend-gating is scored down rather than merely noted, and that is a judgement worth defending: a price that rises with your ad spend charges you more for the same work the better you do. It may still be the right tool. It is a real cost of ownership, and a comparison that hides it in a footnote is not helping.

Why there is no total score

We never add the eight numbers up. The moment you weight dimensions you have chosen the winner, and any weighting we published would be one we happened to be good at.

So every comparison ends in a profile, not a verdict: eight scores, a one-line "buy this if", and a one-line "do not buy this if". Two tools can both be excellent and score nothing alike, because they are answering different questions.

If you want a single number, the honest way to get one is to weight the eight yourself against what your team is actually short of. That is a five-minute job and the answer will be better than ours.

How "best for" pages are allowed to rank

Ranking is allowed, but only within a named job. The dimensions used get printed above the list, before you read a single entry.

Which in practice means:

  • The job is named first. "Best for finding concepts you have never tested" — not "best ad tool".
  • The dimensions used are listed before the list. If a page ranks on 2, 3 and 4, it says so above the first entry.
  • The dimensions not used are listed too. A tool that ranks last for one job usually wins another, and the page has to say which.
  • The underlying scores do not change between pages. One card, many weightings. If a score moves, it moved for everyone, on the re-check date.

The thumb is still on the scale — it has to be, that is what ranking is. It is just printed on the page instead of hidden in the row selection.

The evidence rules

Every score cites the vendor's own page and the date we read it. Anything we could not verify is marked "not verified" and is never silently scored as a zero.

  1. First-party sources only for feature claims. A vendor's product, feature, pricing or documentation pages. Not review sites, not listicles, not their competitors' comparison pages — including ours.
  2. A gap in our research is not a mark against a tool. If we cannot confirm something, the cell reads "not verified". Scoring an unknown as a zero is how you get a table that flatters whoever wrote it.
  3. Dated, because products move. Every comparison carries the date its claims were checked. Anything older than a quarter should be treated as stale.
  4. We have not run paid accounts on the competitors. These pages compare what vendors publish, not what their products feel like on the fortieth day. Where that distinction matters, we say so on the page.
  5. No inferred performance claims. We do not report what is "working" for anyone — competitor spend and results are private, for every tool on every list, including ours.
Conflict of interest, stated plainly. Likely is one of the tools on these lists. We built the scorecard, we do the scoring, and we benefit when you pick us. That is not neutral and no amount of process makes it neutral. What the process can do is make the bias checkable: the dimensions are fixed and published, the rubric is public, every cell cites a source you can open, and our own zeros are printed next to everyone else's.

Where Likely scores zero

Four dimensions, and they are not small ones. If any of these is what you are shopping for, one of the other tools on our lists is the better buy and our own comparison pages say so.

DimensionLikelyWhy
7 · Workflow surface 0–1 No Chrome extension, no boards, no asset management. You cannot run a creative team's day inside Likely.
8 · Commercials 0 Beta. Paid plans are priced individually and there is no public price list, which is exactly what we score other tools down for.
Own-account analytics Absent Not a scored dimension because it is a different category — but Likely never sees your ad account, so it can tell you nothing about your own results.
Platform breadth Meta only Others are live on TikTok today. Queued is not shipped, and we do not score roadmaps.

This section exists because a methodology page that never costs its author anything is marketing wearing a lab coat. If you find a dimension where we have quietly graded ourselves generously, that is a bug — tell us and we will publish the correction.

How to correct us

Email [email protected] with the page, the cell, and a link to the vendor page that shows we are wrong. We will fix it and update the page's "last checked" date.

That applies most of all if you work at one of the tools we scored. You know your own product better than we do, and a comparison that a competitor has read and not disputed is worth considerably more than one they have never seen.

Scores are re-checked quarterly. When a re-check changes something, the page says what changed and when — we do not silently edit a table and leave the date alone.