Methodology
How we compare ad tools
We sell one of the tools we compare. So the only thing that makes our comparisons worth reading is a scorecard we fixed before we looked at anyone — published here in full, including the four dimensions where Likely scores zero.
Read this and you can check our bias instead of trusting it. The dimensions were fixed before we looked at anyone, the rubric is public, and every cell cites a page you can open yourself.
On this page
Why vendor comparisons are usually worthless
Whoever picks the axes picks the winner. A vendor writing a comparison chooses the eight rows where they say yes and the competitor says no, and the table does the rest. Nothing in it is technically false.
You have read these. Fourteen green ticks in one column, three in the others. The rows are real features. The scoring is accurate. And the page is still useless, because the selection of rows is the argument, and it was made after the answer was known.
There is no way to fix that with more honesty inside the table. The fix has to be structural: fix the dimensions before looking at the tools, publish them, score everyone on the same card, and refuse to produce a winner. That is what this page is.
The eight dimensions
Eight dimensions, grouped under the three questions a buyer is actually choosing between. They were chosen to describe the category, not to describe Likely — which is why Likely loses an entire group of them.
| # | Dimension | The question it answers |
|---|---|---|
| A · Research — what can it tell you about the market? | ||
| 1 | Access model | What do you have to hand over before it works? |
| 2 | Coverage | Are you seeing a selection, or all of it? |
| 3 | Structure | Is the data labelled consistently enough to count? |
| B · Decision — does it help you choose what to make? | ||
| 4 | Gap detection | Can it show you what you are not doing? |
| 5 | Prioritisation | Does it rank by what things cost, or only by how big they are? |
| 6 | Evidence trail | Can you check a claim against the creative behind it? |
| C · Operations — does it fit how you actually work? | ||
| 7 | Workflow surface | Can your team live in it day to day? |
| 8 | Commercials | Can you find out what it costs, and does the price punish growth? |
Group C is where we lose. Dimensions 7 and 8 describe daily workflow and pricing transparency, and they are in the card precisely because they are the ones a reader is most likely to actually care about on a Tuesday. Leaving them out would have been the easy way to win.
What each score means
Every dimension is scored 0 to 3. Not 0 to 10 — a ten-point scale on judgements like these is false precision, and it invites a decimal place nobody can defend.
The general shape is the same everywhere: 0 absent · 1 partial or manual · 2 present · 3 the best implementation in the category. A 3 is not a compliment, it is a statement about the field as it stands today, and it can move when somebody builds something better.
1 · Access model
| Score | Means |
|---|---|
| 0 | Nothing works without connecting an ad account. |
| 1 | The main value needs a connection; some browsing works without. |
| 2 | Core research works unconnected; a connection unlocks extras. |
| 3 | Nothing requires account access at all. No OAuth, no partner request. |
2 · Coverage
| Score | Means |
|---|---|
| 0 | Only what you saved yourself. |
| 1 | A searchable library. Selection-based, no denominator. |
| 2 | A large library with filters. Category volume can be approximated. |
| 3 | A complete census of a defined set of brands, so shares and gaps are computable rather than estimated. |
Library size is deliberately not what this measures. A hundred million ads with no denominator cannot tell you what share of a category does anything. Two hundred ads that are all of a category can.
3 · Structure
| Score | Means |
|---|---|
| 0 | No labelling. Folders and free text. |
| 1 | Free-form tags the user applies by hand. |
| 2 | Automatic labelling, vocabulary not published. |
| 3 | A closed vocabulary, published, and stable enough that this quarter compares with last. |
4 · Gap detection
| Score | Means |
|---|---|
| 0 | Nothing. It shows what exists, not what is missing. |
| 1 | You can filter and work it out yourself. |
| 2 | It highlights under-used attributes. |
| 3 | An explicit matrix with empty cells marked against a real denominator. |
5 · Prioritisation
| Score | Means |
|---|---|
| 0 | No ranking of any kind. |
| 1 | Sorts by volume, recency or engagement. |
| 2 | Ranks by an opportunity or impact measure. |
| 3 | Ranks on impact, confidence and what it costs to produce. |
6 · Evidence trail
| Score | Means |
|---|---|
| 0 | Numbers only. You take them on trust. |
| 1 | Some drill-down, some dead ends. |
| 2 | Most metrics open the underlying creative. |
| 3 | Every number opens the ads behind it. |
7 · Workflow surface
| Score | Means |
|---|---|
| 0 | Analysis only. Nothing to work inside. |
| 1 | Export, and not much else. |
| 2 | Two or more of: extension, boards, briefs, asset storage, integrations. |
| 3 | The full production loop — capture, organise, brief, store, integrate. |
8 · Commercials
| Score | Means |
|---|---|
| 0 | No public pricing. Demo-led, or not selling yet. |
| 1 | Public pricing, but gated on your ad spend. |
| 2 | Public flat pricing by seats or usage. |
| 3 | Public flat pricing, a real free tier, no spend gate. |
Spend-gating is scored down rather than merely noted, and that is a judgement worth defending: a price that rises with your ad spend charges you more for the same work the better you do. It may still be the right tool. It is a real cost of ownership, and a comparison that hides it in a footnote is not helping.
Why there is no total score
We never add the eight numbers up. The moment you weight dimensions you have chosen the winner, and any weighting we published would be one we happened to be good at.
So every comparison ends in a profile, not a verdict: eight scores, a one-line "buy this if", and a one-line "do not buy this if". Two tools can both be excellent and score nothing alike, because they are answering different questions.
If you want a single number, the honest way to get one is to weight the eight yourself against what your team is actually short of. That is a five-minute job and the answer will be better than ours.
How "best for" pages are allowed to rank
Ranking is allowed, but only within a named job. The dimensions used get printed above the list, before you read a single entry.
Which in practice means:
- The job is named first. "Best for finding concepts you have never tested" — not "best ad tool".
- The dimensions used are listed before the list. If a page ranks on 2, 3 and 4, it says so above the first entry.
- The dimensions not used are listed too. A tool that ranks last for one job usually wins another, and the page has to say which.
- The underlying scores do not change between pages. One card, many weightings. If a score moves, it moved for everyone, on the re-check date.
The thumb is still on the scale — it has to be, that is what ranking is. It is just printed on the page instead of hidden in the row selection.
The evidence rules
Every score cites the vendor's own page and the date we read it. Anything we could not verify is marked "not verified" and is never silently scored as a zero.
- First-party sources only for feature claims. A vendor's product, feature, pricing or documentation pages. Not review sites, not listicles, not their competitors' comparison pages — including ours.
- A gap in our research is not a mark against a tool. If we cannot confirm something, the cell reads "not verified". Scoring an unknown as a zero is how you get a table that flatters whoever wrote it.
- Dated, because products move. Every comparison carries the date its claims were checked. Anything older than a quarter should be treated as stale.
- We have not run paid accounts on the competitors. These pages compare what vendors publish, not what their products feel like on the fortieth day. Where that distinction matters, we say so on the page.
- No inferred performance claims. We do not report what is "working" for anyone — competitor spend and results are private, for every tool on every list, including ours.
Where Likely scores zero
Four dimensions, and they are not small ones. If any of these is what you are shopping for, one of the other tools on our lists is the better buy and our own comparison pages say so.
| Dimension | Likely | Why |
|---|---|---|
| 7 · Workflow surface | 0–1 | No Chrome extension, no boards, no asset management. You cannot run a creative team's day inside Likely. |
| 8 · Commercials | 0 | Beta. Paid plans are priced individually and there is no public price list, which is exactly what we score other tools down for. |
| Own-account analytics | Absent | Not a scored dimension because it is a different category — but Likely never sees your ad account, so it can tell you nothing about your own results. |
| Platform breadth | Meta only | Others are live on TikTok today. Queued is not shipped, and we do not score roadmaps. |
This section exists because a methodology page that never costs its author anything is marketing wearing a lab coat. If you find a dimension where we have quietly graded ourselves generously, that is a bug — tell us and we will publish the correction.
How to correct us
Email [email protected] with the page, the cell, and a link to the vendor page that shows we are wrong. We will fix it and update the page's "last checked" date.
That applies most of all if you work at one of the tools we scored. You know your own product better than we do, and a comparison that a competitor has read and not disputed is worth considerably more than one they have never seen.
Scores are re-checked quarterly. When a re-check changes something, the page says what changed and when — we do not silently edit a table and leave the date alone.
The comparisons themselves