What this workflow does
Keyword clustering turns a list of queries into page decisions. Tidy spreadsheet groups are only an intermediate result. The useful output tells the team when one page can satisfy related searches, when readers need separate pages and when the site should create no page at all.
Text similarity is useful for finding candidates. It is weak evidence on its own. The final decision should reflect the user’s task, the page type needed to complete it and, for disputed clusters, the kinds of results search engines currently return.
When to use it
Use this workflow before building an information architecture or editorial roadmap. It also helps with content consolidation, cannibalization reviews and the remapping of an existing site after a product or audience change.
Do not run it as a page factory. Google’s people-first guidance asks whether content serves an existing audience and leaves the reader with enough information to achieve a goal. Google’s spam policies also prohibit scaled pages created mainly to manipulate rankings. A cluster justifies a page only when the page has a useful job and the site can do that job well.
Prepare the data
Keep a raw import that you never edit. Build the working sheet from a copy and retain these fields where available:
| Field | Why it matters |
|---|---|
| Original query | Preserves the language people used |
| Market and language | Prevents accidental merging across different result environments |
| Source and export date | Makes changing metrics traceable |
| Demand and difficulty fields | Supports prioritization, with source-specific caveats |
| Current ranking URL | Shows where the site already has relevance or overlap |
| Device or location | Explains result differences for sensitive queries |
| Business and audience fit | Stops raw demand from becoming the only decision rule |
Remove exact duplicates, obvious encoding errors and empty rows. Normalize case and whitespace for analysis, but keep the original query. Do not stem away modifiers that change the task, such as “for beginners,” a city, a year, a product model or “versus.”
Classify the user task
Write a plain-language task for each query before choosing a cluster. “Learn how to choose a canonical URL” is more useful than a broad “informational” label because it reveals what the page must help the reader do.
Record the likely page type as a hypothesis:
- explanation or guide;
- step-by-step procedure;
- comparison;
- category or collection;
- product, service or tool;
- support, policy or reference page;
- navigational destination.
A query can have mixed intent. Give it a primary task, record the secondary task and lower confidence instead of forcing certainty.
Build provisional clusters
Start with shared entities, qualifiers and task language. Use text embeddings or lexical similarity to suggest neighbors if those tools are available, but keep the threshold visible and editable.
Name each provisional cluster after the reader’s job, not the highest-volume phrase. Choose a representative query that expresses the task clearly. Keep near-duplicates together unless location, product, audience or format changes what a satisfactory answer looks like.
At this stage, over-grouping is acceptable because the next step tests the difficult merges. Preserve a reason for every split so reviewers can challenge it.
Test ambiguous clusters with search results
Sample live results when two queries could plausibly share a page. Record the country, language, device assumption and date because results change.
For each query pair, compare:
- repeated ranking URLs and domains;
- dominant page types;
- whether the results solve the same task;
- important differences in audience, locality or freshness;
- the site’s existing pages that already appear.
Result overlap is evidence, not a command. A result page may be volatile, personalized or dominated by a site type you cannot credibly reproduce. Sample more than once for high-impact decisions and keep uncertain pairs in review.
Decide whether to merge or split
Merge queries when one coherent page can satisfy their main tasks without awkward repetition or hidden sections. Split them when the reader needs a different decision, dataset, tool, transaction or local answer.
Use these questions:
- Would the same page purpose satisfy both queries?
- Would the same primary evidence answer both?
- Are the expected page types compatible?
- Can the page title and main heading describe the combined task naturally?
- Would an internal link between two focused pages serve the reader better than one oversized page?
Do not split only because the wording differs. Do not merge only because the phrases share nouns.
Map clusters to page actions
Every approved cluster needs one of four actions:
| Action | Use when |
|---|---|
| Map to an existing page | The page already satisfies the task and needs no material change |
| Improve an existing page | The page is the right destination but lacks useful coverage or evidence |
| Consolidate pages | Several pages compete for the same task and one stronger destination is feasible |
| Propose a new page | The task is useful, distinct and not served by the current site |
A fifth outcome, “no page,” is valid. Use it when the query is off-topic, too ambiguous, unsupported by the site’s expertise or too thin to justify a useful destination.
For consolidation, document redirects, internal-link changes and the content that must survive. Do not delete a ranking page solely because a clustering model groups it with another URL.
Write the cluster table
The working output should include:
| Column | Contents |
|---|---|
| Cluster ID and name | Stable ID plus a task-based label |
| Representative query | Clearest expression of the task |
| Member queries | Original terms as well as normalized forms |
| Intent and page type | Primary task, secondary task and expected format |
| Evidence | Result overlap sample, current URLs and reviewer notes |
| Page action | Existing, improve, consolidate, new or no page |
| Target URL | Current or proposed lowercase, stable URL |
| Confidence | High, medium, low or manual review |
| Risks | Mixed intent, localization, freshness or weak source data |
Keep a separate review queue with the exact question that blocks a decision. “Mixed intent” alone does not tell the next reviewer what to check.
Guardrails from Google’s guidance
- Build pages for reader tasks, not for keyword permutations.
- Use important audience language naturally in the title, main heading, body and link text where it helps comprehension.
- Avoid repeated blocks, location swaps and other scaled patterns that add little original value.
- Give every proposed page a distinct purpose and contribution.
- Use descriptive internal links between related tasks rather than forcing all terms onto one page.
- Do not use an arbitrary word count. Google explicitly lists writing to a supposed preferred word count as a warning sign.
- Do not treat a cluster, volume metric or current result pattern as a ranking guarantee.
Quality checks
- Raw queries, markets, sources and metric dates remain traceable.
- Every cluster has a stated reader task and page type.
- Ambiguous merges have search-result evidence or a review flag.
- Existing ranking URLs are reviewed before new pages are proposed.
- Page actions allow for “no page” and consolidation as well as new content.
- Proposed titles read naturally and do not repeat keyword variants.
- The output contains no doorway-page or scaled-content pattern.
- A second reviewer can reproduce the reasoning behind important merges and splits.
Limitations
Keyword exports are estimates and often merge variants differently. Search results change by time, market and context. Clustering models also favor surface similarity unless the reviewer supplies intent and page-purpose evidence.
Keep confidence labels and dates with the output. Revisit high-impact clusters when the market, offering or existing site changes.
