How to run card sorting without losing the why

How to run card sorting that maps real mental models: the method, the failure modes, and what reasoning probes catch that drag-and-drop tools miss.

Rizvi Haider··20 min read·Updated June 9, 2026

Most card sorting reports end the same way: a dendrogram, six labeled groups, a confident recommendation about top-level navigation, and a question the team will only ask after the build ships. Why did participants put billing under settings instead of account? The clusters are correct. The reasoning never made it out of the drag-and-drop tool. The team discovers the gap two months later, when support tickets start naming the wrong menu.

The method is right. The problem is how to run card sorting in a way that the reasoning survives alongside the sort. Done in the shape information architects have been writing about since the late 1990s, card sorting is the cheapest way a product team can test whether a navigation, a settings page, a documentation index, or a feature taxonomy reflects how the audience actually thinks. Done in the shape most teams run it, it returns a structure without the explanation of why it is that structure, and the team designs the navigation from a dataset that cannot tell them what their participants meant.

This is a working playbook on card sorting: what the method is, the failure modes that turn it into a chart without a story, the six steps that work in 2026, and how multi-modality reasoning probes turn a sort into a mental model the team can ship from.

What card sorting is

Card sorting is a research method, run on a defined set of content items, in which participants group cards into categories the way that makes sense to them. The output is a similarity matrix showing which items participants treated as belonging together and a labeled taxonomy that reflects the participants' mental model, not the team's. The method is one of the oldest in information architecture work and remains the cheapest single instrument for testing whether a navigation, a settings page, a documentation index, or a feature taxonomy matches how the audience actually thinks. The cleanest short reference is Nielsen Norman Group's definition of card sorting, and the deepest practitioner text is still Donna Spencer's Card Sorting: Designing Usable Categories.

Card sorting sits upstream of usability testing and downstream of content inventory. Content inventory tells you what items exist. Card sorting tells you how the audience would group them. Usability testing tells you whether the navigation built from those groups actually lets a real user finish a task. A team that skips card sorting designs the menu from internal logic and discovers in usability testing that nobody else thinks that way.

Why most card sorts don't return a usable IA

Three failure modes show up across most card sorts that returned a clean dataset and shipped a misread navigation. Each one is structural, not effort-related, and they tend to appear together.

The first is a sort without a why. A card sort run inside a drag-and-drop tool produces a clean dataset of which items went where. It produces no record of the participant's reasoning. The dataset can be analyzed, but the analysis is constrained to "what did participants do" and silent on "why did they do it". A surprising cluster (three items grouped together that the team would never have grouped) is the most useful single output of a card sort, and the dataset alone cannot explain it. The fix is to capture the reasoning beside the sort. Most tools do not by default.

The second is a card set written by the team, not derived from the work. If the card text comes from the existing menu labels, the test is anchored to the team's mental model rather than the audience's. The result is a sort that confirms the existing navigation with statistical confidence, because the participants had to use the team's words to express their groupings. The fix is to draw card text from the words the audience already uses for the items: support tickets, sales transcripts, search logs, and the phrases that appeared in earlier discovery interviews.

The third is false convergence on small samples. Card sorting tools often present a similarity matrix that looks confident at fifteen participants. The confidence is partly real (card sorting reaches stable cluster patterns relatively quickly compared to other qualitative methods) and partly a tool artifact (the visualization averages over disagreement). A useful card sort reports the disagreement, not just the consensus. Two clear participant subgroups with different mental models are more actionable than a single averaged dendrogram that obscures both.

How to run card sorting, step by step

Six steps. The order is opinionated: steps one through three are where most card sorts fail before any participant arrives. Steps four through six are where the data either lands or evaporates.

01 · Decide which question the sort is answering

Card sorting can answer three different questions, and the question changes the design of everything else. A general-navigation test asks: how would the audience group these items at the top level of the product. A taxonomy test asks: which items belong together inside a single category, and what should the category be called. A re-organization test asks: does a candidate new structure read more naturally to the audience than the current one.

Pick one. A card sort that tries to answer all three at once returns a dataset the team cannot act on, because the analysis question is not the same for the three sub-tests. The general-navigation test wants an open sort and a dendrogram. The taxonomy test wants a focused similarity matrix inside one category. The re-organization test wants a closed sort against a candidate structure. Choosing the question first is what tells you which type of sort to run.

02 · Build the card set from the work, not the menu

The card set is the artifact participants actually react to. Twenty to forty cards is the working range. Below twenty, the test is too coarse to reveal mental models. Above forty, participants tire and the last ten cards are sorted by exhaustion rather than reasoning.

Card text comes from the audience, not the product. Search logs (what do they type when looking for a thing). Support tickets (what do they call the thing when they cannot find it). Discovery interview transcripts (which verbs do they use when describing the work). Sales calls (what is the deal-killer when the navigation hides a feature). The discussion guide playbook covers how to source language from interviews; the same language sourcing applies to card text.

Avoid abbreviations, internal acronyms, and feature names the audience would not recognize. If a card reads "ACL settings" and the audience would describe the same thing as "who can see what", the card is testing the audience's familiarity with your jargon, not their mental model.

03 · Choose open, closed, or hybrid

The three types of card sort answer different questions and have different costs.

  • Open card sort. Participants group cards into categories they name themselves. Best for greenfield navigation work, when the team has no candidate structure and wants to know how the audience would organize the items. Expensive to analyze (the analysis is qualitative across participant-generated labels) but highest signal for early-stage IA.
  • Closed card sort. Participants sort cards into categories the team already defined. Best for validating a candidate structure, when the question is whether the proposed top-level groups feel natural to the audience. Cheaper to analyze (quantitative against fixed buckets) but lower signal: the participant cannot tell you the structure is wrong, only that an item is misplaced inside it.
  • Hybrid card sort. Participants sort into team-defined categories with the option to create new ones, or sort openly with seeded category suggestions. Best when the team has a structure they believe in but suspects gaps. Splits the cost and the signal between the two.

For most product teams running card sorting for the first time, an open sort is the right starting point. The candidate structure the team would have used for a closed sort is almost always anchored to the existing menu, and a closed test of an existing menu tends to confirm the existing menu.

04 · Recruit against the audience that will actually use the navigation

Card sorting on the wrong audience returns a mental model that does not match the audience that ships. Internal stakeholders, former employees, and friends of the team are all the wrong audience. The right audience is people who currently do the work the navigation supports, not people who are interested in the product.

Fifteen to thirty participants per audience segment is the working range. Below fifteen, the similarity matrix is too noisy to read; the clusters appear stable but reflect noise as much as signal. Above thirty, marginal returns drop sharply for a single segment unless the team is deliberately comparing two or more audience segments (in which case run fifteen to thirty per segment). Card sorting reaches stable cluster patterns faster than open-ended qualitative methods, which is part of its appeal.

The screener filters on behavior, not interest. "How often do you organize project files in your work?" is a screener. "Are you interested in testing a new navigation?" is not. The first filters for participants whose current behavior is observable evidence of the task. The second filters for participants willing to be polite about prototypes. The operational side of recruiting is in how to recruit user research participants.

05 · Capture the "why" alongside the sort

This is the step most card sorting tools skip. The drag-and-drop interaction records where each card went; nothing else. The reasoning, which is where the real information architecture decision lives, is left on the floor.

The fix is to ask, on each card or each group, why the participant placed it where they did. The prompt is shaped to the question type. On an open sort, ask why the participant named the group what they did. On a closed sort, ask why they hesitated between two categories. On a re-organization sort, ask which category felt wrong before they moved on.

A useful card sort uses multiple input modes for these probes. Voice for the open reasoning on a surprising cluster: the participant talks for sixty seconds about why "billing" belongs with "account" and not "settings", and the answer arrives as a transcript with the reasoning intact. Text for the shorter, typed clarifier on a confident sort. Rating for the confidence score on the participant's own grouping. Choice for the binary "did this feel natural or forced". The right modality is the one the participant would use without thinking. The longer essay on how the medium changes the answer is in what we hear when we stop asking people to write.

Adaptive follow-up probes earn their keep on the open-reasoning answers. The first explanation is usually a rehearsed summary ("it just felt right"); the second turn, after one good follow-up, is where the actual mental model arrives ("billing is something I do, settings is something I configure, they are different verbs"). Treat probing depth as a per-question setting: medium depth on the open-reasoning questions, shallow on the confidence ratings, and expert depth when the participant volunteers a contradiction. The longer treatment of how to set follow-up depth is in how AI follow-up questions work in user research.

06 · Analyze clusters and language together

The standard output of a card sort is a dendrogram (the hierarchical clustering of which items participants treated as similar) and a similarity matrix (the per-pair co-occurrence rate). Both are quantitative; both are useful; neither is sufficient.

The second output is a vocabulary map. Across the open answers, which words did participants use to describe groups, items, and the boundaries between them? A dendrogram that says "items A, B, C cluster" does not name the cluster. The vocabulary map names it from the participant's words, not the team's. The category label that goes into the navigation should come from this map, not from a whiteboarding session.

Three patterns to look for in the synthesis.

  • Disagreement signal. Two participant subgroups with different mental models is a more actionable finding than a single averaged dendrogram. Surface the subgroups (often along persona, role, or experience-level axes) and design the navigation against the dominant one with explicit fallbacks for the secondary one.
  • Surprising clusters. Three items grouped together that nobody on the team would have grouped is the most useful single output. Read the open reasoning on those items first. The explanation is almost always actionable.
  • Vocabulary mismatch. If the words participants used to name groups do not appear anywhere in the team's current navigation, the existing labels are guessing at the audience's mental model. The fix is direct: rename to the participants' words.

The general synthesis pass is covered in how to analyze user interview transcripts. For card sorting specifically, the unit of analysis is the cluster-language pair, not the participant.

What multi-modality reasoning probes add

Voice is one of four input modes in a well-run card sort (voice, text, choice, rating), and the modality choice depends on the question. The open-reasoning question on a surprising cluster benefits most from voice, because the participant's explanation unfolds as a sequence of partial thoughts and revisions that compresses badly into a text field. A voice answer to "why did you put these three together" returns a transcript two or three times longer than the typed equivalent, with the energy of the answer attached.

The confidence rating on a self-assigned group benefits from rating, not voice. The "did this feel natural or forced" check benefits from choice. The optional clarifier on a confident sort benefits from text. Each modality is a fit for a specific question; forcing every question into voice loses the answer the same way forcing every question into a text box does.

The point of the multi-modality setup is not to make every question a voice question. It is to let the participant pick the input the question needs.

"These three belong together because they're all things I do when I'm cleaning up. The other ones are things I do when I'm setting up. The menu doesn't say that anywhere, but that's how I think about it."

Participant · #4137 · open-reasoning probe on a surprising cluster

The pull-quote above is what the dendrogram alone cannot produce. The cluster is the data. The reasoning is the design decision.

When to run card sorting internally before customers see it

A pattern that under-uses card sorting badly: running it only externally. The same instrument works inside the company, and running it internally first usually saves a round of external testing. Before the card set goes to participants, share the same set with engineering, design, support, sales, and operations.

The result is a synthesized view of every stakeholder's mental model before the test is exposed to customers. Engineering surfaces the items that are bundled in the codebase and silently assumed to be bundled in the product. Support surfaces the items the audience calls something different in tickets. Sales surfaces the items that are bundled in the pitch deck. Each surfaces a gap in the card set, a missing item, or a clarifying question that the external test would otherwise hit cold.

The async version is a study link shared in internal channels. The team gets a synthesized view of every stakeholder's sort and reasoning, in less time than scheduling a workshop would have taken, and the card set that ships to external participants is calibrated against the internal consensus rather than against guesswork.

Card sorting is usually treated as a one-shot study: run it, ship the navigation, close the link. The version that scales is a standing instrument. The same link, with the latest card set, lives in places where signal arrives continuously and the team would otherwise miss the navigation mismatch.

Three placements that work for card sorting specifically.

  • In-product, on a "something missing here?" affordance. A persistent link in the settings page or the help menu that lets a user say, in their own words, where the navigation failed them. The reply is not a sort every time; it is a vocabulary signal that compounds over months.
  • Onboarding and activation moments. First session complete, first export, day-seven retention. A short card-sort variant ("which of these would you have expected to find together?") returns a clean read on whether the first impression of the navigation matched the user's mental model. The data lands continuously, not at the end of a campaign.
  • Documentation and docs-search dead-ends. When a docs search returns no result, the dead-end page can carry a short sort that asks the visitor to place the term they searched for into one of a small set of categories. The data names the gap directly and points to the category label the audience expected.

A useful frame for the practice: a card sort is a standing instrument for collecting language signal, not a campaign with a start and end date.

When card sorting is the wrong tool

Three cases where card sorting returns a number that pretends to be a finding.

No working content inventory. Card sorting tests how participants would group items. If the team has not yet decided what items exist, the test is premature. The right tool is a content inventory or a feature audit first. Skip the inventory and the sort returns a structure for a set of items that may not be the items that ship.

Tasks, not items. Card sorting answers "how would the audience group these things". It does not answer "can the audience find a thing once it is grouped". The second question is tree testing, which is the closed-set version of the same artifact running against a task ("where would you click to do X?"). A team that runs card sorting alone and skips tree testing learns the audience's mental model and not whether the built navigation works against the model. Both are necessary; the order matters; card sorting is upstream.

Highly specialized or novel categories. A card set built from terms the audience has never seen returns a sort based on guesswork. The right tool is interview-led terminology research first, then a card sort once the audience vocabulary is calibrated.

How card sorting fits into a wider research practice

Card sorting is one tool in an information-architecture practice and it pairs with three others at different stages of the build.

  • Content inventory runs before card sorting. The inventory names what exists; the sort tests how it groups.
  • Tree testing runs after card sorting. The sort produces a candidate structure; the tree test asks whether real users can find a thing in it.
  • Usability testing runs after both. The sort produces the structure; the tree test validates the structure; the usability testing playbook checks whether the built navigation actually completes the task.

All three sit inside the wider practice covered in the voice user research guide. The shorthand: inventory finds the items, card sorting groups them, tree testing validates the groups, usability testing confirms the build.

FAQ

What is card sorting in user research?

Card sorting is a research method used to surface the audience's mental model of a set of content items. Participants are given a defined set of items (the cards) and asked to group them into categories that make sense to them. The output is a similarity matrix showing which items participants treated as belonging together and, when reasoning probes are added, a vocabulary map of how the audience describes the groups. The method is one of the cheapest ways to test whether a navigation, taxonomy, or documentation index matches how the audience actually thinks.

What is the difference between open, closed, and hybrid card sorting?

An open card sort lets participants name their own categories and is best for greenfield navigation work, when the team wants to know how the audience would organize the items from scratch. A closed card sort uses team-defined categories and is best for validating a candidate structure. A hybrid card sort offers team-defined categories with the option to create new ones, and is best when the team has a structure they believe in but suspects gaps. For most product teams running card sorting for the first time, an open sort returns the highest signal because it does not anchor the result to the existing menu.

How many participants do you need for a card sort?

Fifteen to thirty participants per audience segment is the working range. Below fifteen, the similarity matrix is too noisy to read; the clusters appear stable but reflect noise as much as signal. Above thirty, marginal returns drop sharply for a single segment unless the team is deliberately comparing two or more audience segments, in which case fifteen to thirty per segment is the target. Card sorting reaches stable cluster patterns faster than open-ended qualitative methods, which is part of its appeal as a cheap IA instrument.

How do you analyze card sorting results?

Two outputs, read together. The dendrogram (the hierarchical clustering of items by co-occurrence) names which items participants treated as similar. The vocabulary map (the words participants used to label their own groups and explain their reasoning) names what the resulting categories should be called. A team that ships the dendrogram alone gets the structure right and the labels wrong. Look specifically for disagreement between participant subgroups, surprising clusters that nobody on the team would have grouped, and vocabulary mismatches between the participants' words and the current navigation.

What is the difference between card sorting and tree testing?

Card sorting answers "how would the audience group these items". Tree testing answers "can the audience find a thing once it is grouped". The two are sequential, not interchangeable: the sort produces a candidate structure, and the tree test validates the structure against task completion. A team that runs card sorting alone learns the audience's mental model and never tests whether the navigation built from it actually works. A team that runs tree testing alone tests a structure that may not match the mental model in the first place. Both are necessary; card sorting comes first.

Can card sorting be done remotely?

Yes, and the remote version is now the default for most product teams. The trade-off is that drag-and-drop tools record the sort but not the reasoning. The fix is to capture the participant's "why" alongside the sort, in whichever modality the question wants: voice for open reasoning, text for short clarifiers, rating for confidence scores, choice for natural-versus-forced checks. A well-designed remote card sort returns the same dendrogram as an in-person session and a richer reasoning track than a moderated whiteboard session usually produces, because the participant is not performing for the moderator.


Card sorting fails when the dataset arrives without the reasoning and the team ships a dendrogram. It works when the card set is built from the audience's language, the sort type matches the question, the reasoning is captured beside the sort, and the synthesis reads cluster patterns and vocabulary together rather than averaging them into a single number. Talkful is built for the second shape: a card sorting study link goes out, participants sort and answer in voice, text, choice, or rating on their own time, the AI interviewer probes the surprising clusters into honest reasoning at the depth the question deserves, and the synthesis engine streams a vocabulary map alongside the cluster output, ready for the team to ship from or for the agents you build with to act on. The wider voice user research guide covers where card sorting sits inside a continuous research practice.