Skip to content

Keyword clustering: what it is, how to measure it and when two keywords don't belong together

Author: Matteo Pellegrini

Keyword clustering is the work of grouping several queries into a single set, to be covered by a single page, because search engines already treat them as the same question.

The check you need to make the call takes two minutes and needs no paid tools: open the first page of Google UK for the two keywords and count how many URLs appear in both. If they are the majority, it is one question and one page. If there are one or two, they are two different pages, however much the terms look like synonyms to you.

Google already clusters, but on the page side

The word "cluster" in Google's documentation is not about keywords, it is about URLs. The Google Search Central page on canonicalisation explains that "if Google finds multiple pages that seem to be the same or the primary content very similar, it clusters them together", and then picks as canonical the one that is "objectively the most complete and useful for search users". The others stay in the index but are crawled less often.

Put plainly: if you publish two pages on the same question, you no longer decide which one is shown. Keyword clustering is the mirror operation done beforehand, on the demand side, so that you are not left managing duplicate content six months later.

The largest sample available confirms that the unit of measurement is the page, not the individual keyword. Ahrefs analysed 3 million searches and found that a page in first position also ranks in the top 10 for around 1,000 other keywords, with a median of around 400. The sample is international and the study dates from 2017: treat it as an order of magnitude, not a figure to quote to the decimal point.

Counting shared URLs: three pairs measured on Google UK

On 26 September 2026 I took three pairs of English keywords and compared the first-page organic URLs, desktop, English, United Kingdom, using DataForSEO SERP data. The denominator changes from pair to pair because SERP features such as AI Overviews, product boxes and People also ask take up positions and leave fewer organic results.

Keyword pairOrganic URLs in commonWhat follows
"seo meaning" and "what does seo stand for"4 of 7grey area, the writer decides
"running shoes" and "running trainers"0 of 9two SERPs, two pages
"seo meaning" and "seo definition"0 of 7two questions, two pages
Overlap of first-page organic URLs, Google UK, desktop, measured on 26 September 2026 (data source: DataForSEO).

None of the three gave the easy case, where the same results come back and only the order changes. That case exists: in Italy, the same test on two word orders of "what is SEO" returned 8 identical URLs out of 8. When that happens, writing two pages means competing with yourself.

The third pair is the one that surprises people who work from lists. Anyone will tell you that "seo meaning" and "seo definition" ask for the same thing, and Keyword Planner agrees: it reports exactly the same UK volume for both, 6,600 searches a month, with identical monthly figures, because it groups close variants together. The organic results do not agree. "seo meaning" brings up Google's SEO Starter Guide, Wikipedia, the Cambridge Dictionary and Semrush; "seo definition" brings up glossary pages from agencies and tools, most of which define other SEO terms. No URLs in common out of seven. The meaning is the same; the search intent is not.

The second pair is the one every automated tool gets wrong. "Running shoes" and "running trainers" are full synonyms in British English, and they share no organic URL at all. Adidas and Zalando appear in both, but with different pages; "running shoes" also brings up Runnea's buying guides, eBay and Reebok, while "running trainers" brings up an AI Overview, Reddit, Amazon, House of Fraser and Wowcher. A semantic tool puts them in the same group without hesitating. The first pair, instead, is a genuine grey area: Wikipedia, Search Engine Land, the Digital Marketing Institute and Michigan Tech appear in both, the rest differ. With four URLs out of seven there is no right answer on paper: it depends on how many pages you can really keep up to date. And it is worth saying, since it is a limit of the method: a single measurement captures one day, and a grey-area pair can shift after the next algorithm update.

Hard, soft and intermediate: where to close the group

The SERP-based method and its three criteria go back to 2015 and the Russian SEO Alexey Chekushin, who also built the first automated tool to do it, as the Wikipedia entry recounts. The names are still used in tools today.

CriterionWhen two keywords end up in the same groupPractical effect
Hardevery keyword in the group shares the same URLs with all the otherssmall groups, almost never wrong
Softevery keyword shares URLs with the group's most searched keywordlarge groups, including queries that have nothing to do with each other
Intermediateevery keyword has at least one link with another in the groupa compromise, needs a manual pass

Almost all the bloated clusters we end up reviewing come from soft clustering left to run unchecked: the most searched keyword acts as a magnet and pulls in everything that looks even vaguely like it.

Semantic clustering and what it misses

The other approach groups by meaning, with language models or vector comparison, without looking at the search results. It is fast and handles lists of tens of thousands of rows, but used on its own it lumps together "seo meaning" and "seo definition", or "running shoes" and "running trainers", which the measurement above separates. Even Semrush, in its guide to keyword clustering, describes its own Keyword Strategy Builder as a tool that groups "based on search intent and SERP similarity": the second criterion is there because the first is not enough on its own. Semantics does the rough cut; the SERP decides.

From cluster to page

Once the group is closed, the main keyword is not automatically the one with the highest volume: it is the one whose SERP looks most like the page you are able to write. The others become subheadings and variants within the text, placed where they clarify something. Repeating them mechanically only produces keyword stuffing, which Google has recognised for twenty years. The group's long-tail keywords fit well in H2s and FAQs, where they answer a specific question. Pages in the same cluster reference each other with internal links in context, and that is what holds a group of articles together, far more than the ritual pillar page.

The mistakes we find most often in other people's lists

  • Clusters built on the US SERP. The tool runs on google.com with US results and the content plan is for the UK. The two first pages often look nothing alike, so the grouping is skewed from the start. Keyword research should be set up on the market you publish for.
  • Clusters sorted by volume. Volume tells you how many times a string is typed, not who answers it. "seo meaning" and "seo definition" have the same 6,600 monthly searches in the UK and zero URLs in common.
  • Clusters never rechecked. A grouping from two years ago is an old photograph. After an algorithm update or the arrival of an AI Overview, a cluster may have split in two.
  • Clusters that ignore pages already live. Before opening a new group, look at what you have already published: the URL Inspection tool in Search Console shows which URL Google has chosen as canonical for an indexed page, and when the chosen canonical is not the one you expect, two of your pieces are already answering the same question.

Something the industry rarely says: clustering is more useful for removing rows from a content plan than for adding them. It is presented as an expansion technique, and in day-to-day work its most useful output is the list of pages you have decided not to write. The term itself is insider jargon: in the UK "keyword clustering" gets around 70 searches a month (DataForSEO, September 2026). It is something you do, not a topic to build traffic on. If you need someone to do it on your content plan before it turns into a cannibalisation problem, it is part of our SEO service.

Keyword clustering FAQs

How many keywords can go in a cluster?

There is no fixed number: it depends on how many first-page URLs the keywords share. With the hard criterion, where every query in the group must have the same results, clusters stay small, often two to five keywords. A group of forty keywords is almost always soft clustering left to run without a checking pass.

Is SERP-based or semantic clustering better?

The two criteria work at different stages. Semantic clustering quickly roughs out very long lists by grouping on meaning; SERP-based clustering checks whether Google really treats the queries as the same question. On a list of thousands of rows you use both, in that order. On a list of fifty keywords the second is enough.

Does keyword clustering apply to AI Overviews too?

The unit that gets cited is still the page. On the query seo meaning, in the measurement of 26 September 2026 on Google UK, the AI Overview cited Wikipedia and the Semrush guide among its sources, two URLs that also appeared in the organic results. A page that covers a group of queries well is still the way to get cited.

How can I tell if I already have two pages in the same cluster?

In the Search Console Performance report, filter by the query and see how many pages get impressions: if there are two and they alternate over time, they are answering the same question. The URL Inspection tool confirms which of the two Google considers canonical.

Matteo Pellegrini

Matteo Pellegrini

I’m a Business Developer, and at Visilay I focus on developing data-driven SEO, Google Ads, and CRO strategies. I love historical museums, have been practicing Karate for as long as I can remember, and on weekends I enjoy exploring Italian villages in search of authentic local food.