Measurement & metrics
Where AI visibility tools actually differ. Ten products, point by point.
There is no single best AI visibility tool. Across ten products and fourteen points the differences gather in four places: what gets measured and how, whether findings turn into work, which markets and languages, and what it takes to start. Nothing is scored. This is written by the company behind one of the ten.
First: this is written by one of the ten
This article is written by the company that builds Suparanku, which is one of the tools compared here. We would rather say that up front.
A comparison that includes your own product is not neutral just because it says so. So this article ranks nothing and scores nothing. Each tool has its own comparison page, written one product at a time; this article is the way in to those. It sets out only what differs, point by point — including the points where the other tool is stronger.
How to use this article
It is long. You do not have to read it in order. Go to the points that matter where you are.
- What gets measured, and how — engines, how it measures, cadence and history, crawler signals, diagnosing your site
- Turning what you find into work — recommendations, producing content, sites you don’t own, proving it worked
- Which markets and languages — Japanese and multilingual, whether it works for an agency
- What it takes to start — cost, access from your own tools
- The two points Suparanku never wins — start here if you want the weaknesses first
- Ten tools: which suits what — one line each
- Three steps for choosing
The section that follows, on method, underpins every point above. Read it first.
”Which one is best” has no answer
The question we get most often is which tool is best. But the products in this category are not built for the same job. Some are for seeing how AI answers introduce your company; some are for finding what is wrong on your site; some are for an agency watching many clients at once. The words “AI visibility tool” cover all three, and they are not the same thing.
So we replace the question with a better one: which of these points matters in your situation? Laid side by side, the ten comparisons differ in four places.
- What gets measured, and how — engines, how it measures, cadence and history, what an answer tells you, crawler signals, diagnosing your site
- Turning what you find into work — what to do next, producing content, sites you don’t own, proving it worked
- Which markets and languages — Japanese and multilingual, whether it works for an agency
- What it takes to start — cost, access from your own tools
These four are ordered by which part they cover of managing the brand information AI passes on, AI Representation Management (ARM). GEO, LLMO and AEO are the names of tactics used inside it. The section “Why the category divides along these four” explains this.
How the comparison was made: there are no scores
The method changes how you should read the rest, so it comes first.
- Every verdict is one of five phrasings. Only Suparanku does this · Suparanku is stronger here · Comparable, different approaches · The other tool is stronger here · Only the other tool does this. We invent no scores, and nothing is summed into a ranking. Where the facts on a point do not force a verdict, it is “comparable”.
- Everything said about another tool comes from primary sources — pricing pages, official documentation, changelogs. Reviews are used only in the form “users report that…”.
- “We could not find this documented” and “the product does not do this” are different statements. Where we could not confirm something, we write that we could not find it documented. We do not write that the feature is missing. What we can check is whether it is published, not whether it exists.
- Plan gates are stated too. In this category the difference is less often the feature than which plan it starts on. We state the other tool’s gates and our own the same way. Where our free plan cannot do something, we say it cannot.
- The verification date is on the page. The pages were checked on 27 August and 1 September 2026. These products change monthly; if time has passed, check the vendor’s own pages as well.
- We re-read everything before publishing. On 1 September we went back through the finished comparisons against each vendor’s own pages. Seven of eight carried something that needed correcting. Some errors ran in our favour, some against us. All were fixed before publication — and without that pass they would have gone out as written.
One more thing about comparison itself. The same metric name means different arithmetic in different tools. Every product uses the words “share”, “visibility”, “sentiment”. But what goes in the denominator, how many questions there are and where they come from, and how tone is judged all differ by product, and almost no product publishes its method. Numbers produced by two different tools are not comparable side by side. That is one reason this article gives no scores. Figures in a row look like a comparison, but what they measure is not the same thing.
If you find something inaccurate, tell us and we will check it and correct it.
What gets measured, and how
Start here. Which AI, asked what, how often — that is what determines the data you end up holding.
Which AI engines. Suparanku measures ChatGPT, Claude, Gemini and Google AI Overviews. How many you get depends on the plan. The free plan gives two, ChatGPT and Google AI Overviews, with no choice. Starter lets you pick three of the four. All four start at Business. And you cannot add Perplexity or Copilot by paying more: we do not cover them. Of the ten comparisons, five are “comparable”, and DolphinX AIO and Semrush are “Suparanku is stronger”. And on Profound, Scrunch AI and AKARUMI we write “the other tool is stronger”. On engine breadth, those three beat us.
How each measurement is taken. Ask the same question twice and the answer changes. In SparkToro’s research, the chance of the same brand list coming back twice in a row was under 1 in 100, and the chance of the order matching too was about 1 in 1,000. A single answer cannot be treated as the result.
Suparanku asks three times at every scheduled measurement and averages, and publishes that method. That is what six of nine comparisons come out “Suparanku is stronger” on. This does not mean three is statistically sufficient. Treating the figures as a distribution and reporting confidence intervals is held to need dozens of runs per question. Three is the floor for not deciding on one answer.
Nor does it mean the others ask once. None of them publishes how many times it asks, so all we can say is that it is not published. What separates us here is not the number of runs but publishing it.
Cadence and history. On all six comparisons where this is a row, the other tool is stronger. Suparanku scans weekly. That is on the paid plans; the free plan scans once. Some products measure daily. We also lose on how far back the history goes. On Ahrefs Brand Radar we write “only the other tool does this”. If continuous, fine-grained observation is your first priority, read this row first.
Crawler signals. On all five comparisons where this is a row, the other tool is stronger. This is reading your CDN logs to see when AI crawlers visited your site. Profound documents nine integrations.
Diagnosing your site. We win all eight comparisons, though not with the same verdict: seven are “Suparanku is stronger” and AKARUMI is “only Suparanku does this”. Across the fourteen points, this is where the margin is widest.
As for how much you can read out of a single answer — the sources cited, the position, the sentiment, the competitors alongside you — four are “Suparanku is stronger” and five are “comparable”.
Turning what you find into work
Measurement tells you where you stand. This group is about whether anything follows from it.
Sites you don’t own. AI does not only quote your own site. On Webflow AEO and ミエルカGEO we write “only Suparanku does this”; on Profound, “the other tool is stronger”; on AKARUMI, “Suparanku is stronger”; the remaining five are “comparable”.
The other three rows in this group, what to do next, producing content and proving it worked, are mostly “comparable”. Everyone is investing here, so the gaps have closed. There are three exceptions: on Ahrefs Brand Radar two rows are “only Suparanku does this”, and on Profound’s content production the other tool is stronger. One thing to be clear about: what Suparanku produces is recommendations and content briefs. Writing and publishing is done by your team or your agency.
Which markets and languages
Japanese and multilingual. On all seven overseas tools we write “Suparanku is stronger”. But on two of the three Japanese tools (ミエルカGEO, DolphinX AIO) this row is not in the comparison at all. Japanese handling is not where we differ from them. On the third, AKARUMI, it is in the comparison and it comes out “comparable” — they are a Tokyo company too, with Japanese support and Japanese data residency. If you are considering an overseas tool, check how it handles Japanese answers.
For agencies. Five of eight are “comparable”, two are “Suparanku is stronger”, and on Semrush the other tool is stronger. This row compares two different things: the per-client cost and how built-out the partner programme is, against whether there is a deliverable with your own name on it to hand the client. Suparanku’s client report is co-branded, “Suparanku × your logo”. Full white-label, with our name gone, does not exist. Semrush has it.
What it takes to start
Cost. Seven of ten are “Suparanku is stronger”, three “comparable”. Access from your own tools: one “only Suparanku does this”, three “Suparanku is stronger”, six “comparable”.
One correction worth making: an MCP server for connecting AI agents is not ours alone. Ahrefs, for one, ships one on every paid plan.
We quote no prices in this article. Pricing in this category changes monthly, and what a plan includes changes with it. An article full of figures is out of date the month after it is written. Each vendor’s figures are on the individual comparison pages, with the date they were checked.
Seen whole, our losses gather in two points
Counting fourteen points across ten companies produces a clear shape.
There are two points we do not win a single row on. Cadence and history, where the other tool is stronger on all six, and crawler signals, where the other tool is stronger on all five. If you need daily measurement, a long history, or AI-crawler activity read from your own server logs, Suparanku is not the tool.
There is now one point we win every row on. Diagnosing your site, all eight. Japanese and multilingual is no longer one of them: we still win it on all seven overseas tools, but against AKARUMI it is “comparable”, so the clean sweep is gone. When the other product is a domestic one, Japanese support is not a difference.
The remaining eleven points are mostly “comparable”. So the products in this category are not divided by which is better overall. They are divided by which points you need.
Why the category divides along these four
One piece of background first. Managing the brand information AI passes on is what we call AI Representation Management (ARM): a system for managing how AI assistants introduce your business, in which the following six come round and start again.
- Get your business information in order
- Measure what AI answers
- Produce recommendations
- Publish content
- Record the URL you published
- Measure again
This is the closed loop of representation management.
GEO, LLMO and AEO are names for the work inside that loop
All three name much the same thing from different angles, and in practice they overlap.
- GEO (Generative Engine Optimization) — getting your company named accurately, favourably and early inside the answers ChatGPT and Gemini generate.
- LLMO (Large Language Model Optimization) — the term that took hold in the Japanese marketing industry; near-synonymous with GEO. Outside Japan, GEO is the one people know.
- AEO (Answer Engine Optimization) — writing so your information gets picked as the answer to a question, in a form that is easy to extract.
The names differ; the foundations do not. And each of them names the work, not when to do it, on what evidence, or in what order. Deciding that, publishing, and re-measuring the effect as one operation is what ARM is. When you are choosing a tool, that distinction is the difference between the products.
The four groups are four questions
The four groups in this article follow the order of the loop. Restated, they are four questions.
- What gets measured, and how — how accurately can you know what AI says about you?
- Turning what you find into work — once you know, does it tell you what to do next?
- Which markets and languages — can you run that in your market and your language?
- What it takes to start — can you keep doing it next month, and next year?
And this is where the products divide. But it is not a clean split between “products that stop at measurement” and “products that close the loop”. Many cover part of the loop, and in this article’s own comparisons “what to do next”, “producing content” and “proving it worked” are mostly “comparable”. On execution, some go further than we do: Profound writes the content and publishes it, while Suparanku stops at the recommendation.
Suparanku’s position is narrower than that. The whole loop sits in one product at one price, the measurement method is published, and it works in Japanese. We claim nothing beyond that.
Ten tools: which one suits what
Each comparison page ends with a card saying when the other tool is the better fit. One line from each.
- Webflow AEO — your site is on Webflow and you are already on Team or Enterprise. Fixes go live on the next publish.
- Scrunch AI — you need engine breadth today. It covers Perplexity, Microsoft Copilot and Meta AI. Suparanku covers none of them.
- Ahrefs Brand Radar — you want to look up any brand instantly, with history back to 2024. Nothing in Suparanku looks backwards.
- Profound — you want engine breadth, and the content written and published for you. Suparanku stops at the recommendation.
- Otterly AI — budget is the binding constraint. Daily tracking costs less there than any Suparanku plan.
- Peec AI — you need daily measurement, crawler data from your own server logs, or a REST API. Suparanku offers none of the three.
- Semrush — you already use Semrush. Or you are an agency that needs true white-label.
- ミエルカGEO — you want production and operations from the same company, not just measurement. Or a listed company’s track record matters to your internal sign-off. Or you need a qualified invoice (インボイス): we are a tax-exempt business and cannot issue one.
- AKARUMI — you need Perplexity, Copilot or Google AI Mode measured. You need daily measurement. You want AI-crawler access logs read directly. We offer none of the three.
- DolphinX AIO — you want hands-on support building the capability in house, from the same company. Or SEO and AIO under one vendor.
The evidence for each is on its own comparison page. All ten are listed point by point on the comparison index.
Three steps for choosing
1. Pick the two or three points of the fourteen that matter to you. No product is best on all of them. Read the four groups above and choose the points that mean something in your situation. A company that needs daily measurement and a company that needs good Japanese answers will not arrive at the same tool.
2. Check those points, and only those, against each vendor’s own pages. This article and the comparison pages carry their verification dates. Where time has passed, pricing and plan gates are the most likely to have moved.
3. Run it on your own data. Reading brochures will not tell you what AI says about your brand. Suparanku’s free audit runs one real scan on your domain and emails you the result. No card required. Want to see how AI introduces your company? Start with the free AI brand audit.
Choosing a tool and doing the work are two different things. Once the measurement tells you what to fix, someone has to fix it — your own team, or someone outside. If it is someone outside, our certified partners are the agencies that already run Suparanku day to day.
Sources and verification dates
- This article is a summary of the ten comparison pages. The full source list for every claim about another tool is on those pages, with the date it was checked. Nothing appears here that is not written there.
- The counts in this article are of the verdicts on those pages as of 3 September 2026. Every figure like “five of ten are comparable” is counted that way. If a page’s verdict changes, the figures here change with it.
- The pages were verified on 27 August and 1 September 2026.
- If you find an error, tell us. We will check it, correct it, and update the verification date.